Search NASA⌕ Search

SEARCH · Search NASA

Results for “Dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Ongoing Work: A Prototype Dataset for Low-flying Autonomous Medical UAS Operations

This paper presents ongoing work to create a dataset for low-flying autonomous medical UAS operations, focused on human stance recognition. This is an exploration of the viability of airborne classification for the Drone as a First Responder (DFR) concept in which a UAS arrives at the scene of an incident before emergency response personnel can get there and provides some level of situational awareness for the personnel arriving to the scene. Future incarnations could also see the UAS administer some level of care to injured parties at the scene. The data set, focused on detecting human stance, being developed here is the result of 30 test flights at NASA Langley Research Center in early 2024. In addition to flights where the participant (an anthropomorphic testing device or human) is alone in the viewing area holding a particular stance, two emergency scenes have been fabricated and collected through video - ``bike crash'' and ``difficult camping''. These test flights include four human participants. The contribution of this work upon completion will be a publicly available data set for the development of classification engines focused on human stance, and in the future, even triage.

Uncrewed Aerial Systems↗

BCARS Simulated Phantom Dataset for Evaluation of Processing Pipelines

Broadband coherent anti-Stokes Raman scattering (BCARS) microscopy is a powerful label-free biological imaging technique, but the raw signal requires careful processing. The vibrationally resonant (Raman) fingerprint signal is usually small compared with instrumental noise sources and the nonresonant background (NRB) inherent in the BCARS signal. Fortunately, the NRB exhibits a systematic phase relationship with the coherent Raman response, acting as a heterodyne amplifier for the weak fingerprint signal. Due to this heterodyne effect, the Raman response can be recovered quantitatively and invariantly across different instruments, provided the NRB shape is known. Even with heterodyne amplification, the amplitudes of fingerprint signal components are often comparable to system noise. Singular value decomposition (SVD), which utilizes spatial information, is often employed for additional noise filtering. Consequently, finding optimal processing parameters to properly distinguish the NRB and Raman responses and suppress noise in the complex BCARS signal requires a reference system that realistically represents the spectral and spatial properties of BCARS signals obtained from biological samples. We present a digital tissue phantom that meets these criteria as a tool for testing candidate signal processing pipelines. The digital phantom is generated with simulated hyperspectral Raman images having system-specific noise and background characteristics. Here, we analyze phantom datasets with differing background and signal-to-noise conditions to evaluate their impact on the performance of multiple signal processing pipelines. Specifically, we investigate the application of a Butterworth filter-based routine to directly estimate the NRB from the BCARS signal. Additionally, we evaluate a Lorentzian wavelet transform as an alternative to the Hilbert transform for extracting the Raman spectrum from the BCARS signal. While we demonstrate this phantom for BCARS, it can be used for any spectroscopic Raman imaging approach.

Dixon, Jessica Z. [Georgia Institute of Technology↗

MP-ALOE: an r2SCAN dataset for universal machine learning interatomic potentials

We present MP-ALOE, a dataset of nearly 1 million DFT calculations using the accurate r2SCAN meta-generalized gradient approximation. Covering 89 elements, MP-ALOE was created using active learning and primarily consists of off-equilibrium structures. We benchmark a machine learning interatomic potential trained on MP-ALOE, and evaluate its performance on a series of benchmarks, including predicting the thermochemical properties of equilibrium structures; predicting forces of far-from-equilibrium structures; maintaining physical soundness under static extreme deformations; and molecular dynamic stability under extreme temperatures and pressures. MP-ALOE shows strong performance on all of these benchmarks and is made public for the broader community to utilize.

Kuner, Matthew C↗

High-throughput dataset of impurity adsorption on common catalysts in biomass upgrading applications

Abstract An extensive dataset consisting of adsorption energies of pernicious impurities present in biomass upgrading processes on common catalysts and support materials has been generated. This work aims to inform catalyst and process development for the conversion of biomass-derived feedstocks to fuels and chemicals. A high-throughput workflow was developed to execute density functional theory calculations for a diverse set of atomic (Al, B, Ca, Cl, Fe, K, Mg, Mn, N, Na, P, S, Si, Zn) and molecular (COS, H 2 S, HCl, HCN, K 2 O, KCl, NH 3 ) species on 35 unique surfaces for transition-metal (Ag, Au, Co, Cu, Fe, Ir, Ni, Pd, Pt, Re, Rh, Ru) and metal-oxide (Al 2 O 3 , MgO, anatase-TiO 2 , rutile-TiO 2 , ZnO, ZrO 2 ) catalysts and supports. Approximately 3,000 unique adsorption geometries and corresponding adsorption energies were obtained.

09 BIOMASS FUELS↗

High-resolution climate model datasets for energy infrastructure planning in a renewable-dependent future

Electrification and renewables deployment efforts are amplifying the interdependence of the climate and energy systems. Increases in climate model resolution, which is now approaching that of reanalysis datasets and operational weather forecast models, present a unique opportunity to use future climate projections for energy infrastructure planning. In this Perspective, we review recent developments in high-resolution climate modeling, which have been driven by increased computing power and advanced software tools. We then look ahead to discuss how high-resolution climate data can be used to plan for a renewable-dependent future, and envision a unified climate-energy model framework that captures the two-way feedbacks between these interdependent systems.

climate change↗

Search for Majorana Neutrinos with the Complete KamLAND-Zen Dataset

We present a search for neutrinoless double-beta (0⁢𝜈⁢𝛽⁢𝛽) decay of 136 Xe using the full KamLAND-Zen 800 dataset with 745 kg of enriched xenon, corresponding to an exposure of 2.1 ton yr of 136 Xe. This updated search benefits from a more than twofold increase in exposure, recovery of photo-sensor gain, and reduced background from muon-induced spallation of xenon. Combining with the search in the previous KamLAND-Zen phase, we obtain a lower limit for the 0⁢𝜈⁢𝛽⁢𝛽 decay halflife of 𝑇$^{0⁢𝜈}_{1/2}$ >3.8 ×10 26 yr at 90% CL, a factor of 1.7 improvement over the previous limit. The corresponding upper limits on the effective Majorana neutrino mass are in the range 28–122 meV using phenomenological nuclear matrix element calculations.

Abe, S. [Tohoku University] (ORCID:000000022110513↗

Dark energy survey: Modeling strategy for multiprobe cluster cosmology and validation for the full six-year dataset

Here, we introduce an updated To&Krause2021 model for joint analyses of cluster abundances and large-scale two-point correlations of weak lensing and galaxy and cluster clustering (termed CL+3×2 pt analysis) and validate that this model meets the systematic accuracy requirements of analyses with the statistical precision of the final Dark Energy Survey (DES) Year 6 (Y6) dataset. The validation program consists of two distinct approaches, (i) identification of modeling and parametrization choices and impact studies using simulated analyses with each possible model misspecification and (ii) end-to-end validation using mock catalogs from customized Cardinal simulations that incorporate realistic galaxy populations and DES-Y6-specific galaxy and cluster selection and photometric redshift modeling, which are the key observational systematics. In combination, these validation tests indicate that the model presented here meets the accuracy requirements of DES-Y6 for CL+3×2 pt based on a large list of tests for known systematics. In addition, we also validate that the model is sufficient for several other data combinations: the CL+GC subset of this data vector (excluding galaxy–galaxy lensing and cosmic shear two-point statistics) and the CL+3×2 pt+BAO+SN (combination of CL+3×2 pt with the previously published Y6 DES baryonic acoustic oscillation and Y5 supernovae data).

79 ASTRONOMY AND ASTROPHYSICS↗

Expanding on the BRIAR Dataset: A Comprehensive Whole Body Biometric Recognition Resource at Extreme Distances and Real-World Scenarios (Collections 1-4)

The state-of-the-art in biometric recognition algorithms and operational systems has advanced quickly in recent years providing high accuracy and robustness in more challenging collection environments and consumer applications. However, the technology still suffers greatly when applied to non-conventional settings such as those seen when performing identification at extreme distances or from elevated cameras on buildings or mounted to UAVs. This paper summarizes an extension to the largest dataset currently focused on addressing these operational challenges, and describes its composition as well as methodologies of collection, curation, and annotation.

Cornett, David [ORNL] (ORCID:0000000222910860)↗

Remark on Algorithm 1012: Computing Projections with Large Datasets

In ACM TOMS Algorithm 1012, the DELAUNAYSPARSE software is given for performing Delaunay interpolation in medium to high dimensions. When extrapolating outside the convex hull of the training set, DELAUNAYSPARSE calls the nonnegative least squares solver DWNNLS to compute projections onto the convex hull. However, DWNNLS and many other available sum-of-squares optimization solvers were not intended for usage with many variable problems, which result from the large training sets that are typical in machine learning applications. Thus, a new PROJECT subroutine is given, based on the highly customizable quadratic program solver BQPD. This solution is shown to be as robust as DELAUNAYSPARSE for projection onto both synthetic and real-world datasets, where other available solvers frequently fail. Although it is intended as an update for DELAUNAYSPARSE, due to the difficulty and prevalence of the problem, this solution is likely to be of external interest as well.

97 MATHEMATICS AND COMPUTING↗

Intelligent Sampling of Extreme-Scale Turbulence Datasets for Accurate and Efficient Spatiotemporal Model Training

With the end of Moore’s law and Dennard scaling, efficient training increasingly requires rethinking data volume. Can we train better models with significantly less data via intelligent subsampling? To explore this, we develop SICKLE, a sparse intelligent curation framework for efficient learning, featuring a novel maximum entropy (MaxEnt) sampling approach, scalable training, and energy benchmarking. We compare MaxEnt with random and phase-space sampling on large direct numerical simulation (DNS) datasets of turbulence. Evaluating SICKLE at scale on Frontier, we show that subsampling as a preprocessing step can, in many cases, improve model accuracy and substantially lower energy consumption, with observed reductions of up to 38×.

Brewer, Wes [ORNL] (ORCID:0000000236393956)↗

Dataset for the paper titled "Investigating the relationship between bolide entry angle and apparent direction of infrasound signal arrivals"

This dataset includes outputs generated for the journal publication titled: "Investigating the relationship between bolide entry angle and apparent direction of infrasound signal arrivals". The outputs include .csv files with model-generated synthetic trajectories of asteroids entering Earth at a variety of impact and approach (azimuthal) angles. All outputs are based on hypothetical but realistic scenarios.

Herrera, Natalie [Sandia National Laboratories (SN↗

Idaho National Laboratory Quality Of Service Dataset

The code is designed to run tests to generate and collect data from a Wi-Fi network using OPENWRT or a simulated a 5G network using Open5gs and UERANSIM. The tests simulate the network performing downloads or uploads of various files with a varying number of concurrent users. The tests use tcpdump to collect the network traffic but only stores the summarized data. The summarized datasets will be included.

Krome, Cameron [Idaho National Laboratory (INL), I↗

Contribution to Open Molecules Dataset (coordcomplexsampling)

The Open Molecules Dataset is a project led by external collaborators, focused on the production of a wide diversity of molecular chemistries. This project will include high-throughput electronic structure calculations on metal coordination complexes, biomolecules, and electrolytes. Solvation effects, conformers, chemical reactivity, and spin/charge sampling will be pursued for elements on the periodic table up to, but not including the actinides. We aim to contribute code related to sampling for transition metal complexes, in particular, and expand software capabilities related to other thrusts desired.

Taylor, Michael G. [Los Alamos National Laboratory↗

SUNSet: The Software Understanding for National Security Dataset Repo

SAND2025-00511O SUNSet: The Software Understanding for National Security Dataset Repo serves as a repository for software understanding researchers to conduct systematic research in the field. It provides a platform for storing questions, answers, and scripts related to software programs, supporting research in software understanding. The repository is a nascent effort aimed at exploring the support needed by researchers and documenting how software understanding questions can be addressed using existing tools. The software consists of a simple database and front end interface for easy access and management of the stored information. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Amon, Tod [Sandia National Lab. (SNL-CA), Livermor↗

Spatiotemporally Registered In-Situ and Ex-Situ Datasets for Laser-based Blown Powder Directed Energy Deposition

This dataset is comprised of in situ sensing data collected during laser-based, blown powder directed energy deposition (DED) of Inconel 718 representing eight different printing conditions: (1) nominal, (2) +15% scan speed, (3) +12% laser power, (4) +42% powder feed rate, (5) +100% jerk limit, (6) +10% layer height, (7) +20% hatch spacing, (8) +20% carrier gas flow. All eight DED builds constructed an identical test coupon geometry consisting of geometric features representative of industrial print requirements (e.g., bulk deposition, thin walls, overhangs). In situ data consists of xyz-coordinates (100 Hz) and on-axis melt pool camera video (60 Hz), both of which have been temporally synchronized to spatially map the melt pool camera data. In addition, post-build X-ray computed tomography (XCT) data for each of the eight test geometries have been spatially registered to the recorded xyz-coordinates, allowing for comparisons between melt pool camera data and flaws identified in the XCT data.

additive manufacturing↗

Disk Failure Dataset from the Campaign Storage System

This dataset consists of 1,389 disk (HDD) failure events collected from the Campaign storage system at LANL. The Campaign system supported various compute platforms throughout its lifespan, including Cielo, Fire, Ice, and notably, the Trinity supercomputer. Each recorded event includes its detection timestamp (in ISO 8601 format) and details such as its location within the storage system—rack, enclosure, and drive slot number. The data, spanning from May 4, 2021, to July 25, 2023 (2 years, 2 months, and 22 days), represents failure events from the terminal years of Campaign's operational period, accounting for 26% of its total operational time.

97 MATHEMATICS AND COMPUTING↗