Search NASA⌕ Search

SEARCH · Search NASA

Results for “data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Accelerating Structure–Property Relationship Discovery with Multimodal Machine Learning and Self-Driving Microscopy

Microscopy combined with local spectroscopy is widely used to correlate nanoscale structure with functional properties in materials, but conventional measurements rely heavily on human-selected sampling locations and predefined targets, limiting data set diversity and the potential for discovery. Here, we present a framework that integrates autonomous microscopy with dual-novelty deep kernel learning (DN-DKL) for adaptive data acquisition and a dual variational autoencoder (VAE) for representation learning. DN-DKL actively guides the microscopy toward structurally and spectroscopically novel regions, enabling efficient collection of large spectral data sets. Dual-VAE embeds local structures and spectroscopic responses into a shared latent manifold that serves as a structure–property relationship map. We applied this framework for the investigation of halide perovskite films by using conductive atomic force microscopy. The results reveal distinct hysteresis behaviors that are linked to specific nanoscale structural motifs, including grain boundary junction points that show hysteresis under different bias conditions and asymmetric grain boundaries that suppress the charge transport. This framework establishes a general strategy that leverages the complementary strengths of self-driving microscopy, machine learning, and human expertise to accelerate scientific discovery in functional materials.

atomic force microscopy↗

Extracting Material Property Measurements from Scientific Literature with Limited Annotations

Extracting material property data from scientific text is pivotal for advancing data-driven research in chemistry and materials science; however, the extensive annotation effort required to produce training data for named entity recognition (NER) models for this task often makes it a barrier to extracting specialized data sets. Here, in this work, we present a comparative study of the conventional, supervised NER methodology to alternative few-shot learning architectures and large language model (LLM)-based approaches that mitigate the need to label large training data sets. We find that the best-performing LLM (GPT-4o) not only excels in directly extracting relevant material properties based on limited examples but also enhances supervised learning through data augmentation. We supplement our findings with error and data quality assessments to provide a nuanced understanding of factors that impact property measurement extraction.

36 MATERIALS SCIENCE↗

Evaluation of the Planetary Boundary Layer Height From ERA5 Reanalysis With MOSAiC Observations Over the Arctic Ocean

The planetary boundary layer height (PBLH) is a crucial indicator reflecting the region of the atmosphere characterized by continuous turbulence. Here, we use radiosonde and surface meteorological observations (4–7 times per day, year-round measurements) during the Multidisciplinary drifting Observatory for the Study of Arctic Climate (MOSAiC) expedition to derive the PBLH (PBLH MOSAiC ), and further evaluate the PBLH from the ERA5 reanalysis (PBLH ERA5 ). Comparisons between PBLH MOSAiC and PBLH ERA5 from different perspectives reveal that: (a) The overestimation of PBLH ERA5 when the sea ice concentration is >90% is significant with the centered root mean squared error reaching up to 201 m; (b) The difference between the two products is notably pronounced in cold seasons, while it is comparatively diminished in warm seasons; (c) In neutral boundary layers, differences in PBLH ERA5 are larger compared with stable and convective boundary layers. In addition, the analysis of error sources indicates that the bias of PBLH ERA5 is sensitive to the bias of vertical thermal structure and wind speed profiles in ERA5 data sets in all conditions. Finally, we find a Random Forest model effectively reduces the bias of PBLH ERA5 with the index of agreement reaching up to 0.71 in the test data set, while a multiple linear regression demonstrates comparable performance to the Random Forest model.

54 ENVIRONMENTAL SCIENCES↗

Machine learning inversion from scattering for mechanically driven polymers

A machine learning inversion method is developed for analyzing scattering functions of mechanically driven polymers and extracting the corresponding feature parameters, which include energy parameters and conformation variables. The polymer is modeled as a chain of fixed-length bonds constrained by bending energy, and it is subject to external forces such as stretching and shear. We generate a data set consisting of random combinations of energy parameters, including bending modulus, stretching and shear force, along with Monte Carlo-calculated scattering functions and conformation variables such as end-to-end distance, radius of gyration and off-diagonal component of the gyration tensor. The effects of the energy parameters on the polymer are captured by the scattering function, and principal component analysis ensures the feasibility of the machine learning inversion. Finally, we train a Gaussian process regressor using part of the data set as a training set and validate the trained regressor for inversion using the rest of the data. The regressor successfully extracts the feature parameters.

Gaussian process regressors↗

Turbulent Parameters by airborne measurements over BNF in March 2025

The original data were collected during the AAF Engineering Flights (AEF2025) in the vicinity of the ARM Bankhead National Forest (BNF) Atmospheric Observatory (https://www.arm.gov/capabilities/observatories/bnf ) in northwestern Alabama in March 2025. The ARM Aerial Facility ArcticShark uncrewed aerial system (UAS, https://www.arm.gov/capabilities/observatories/aaf/uas) was based at the public-use airport of Posey Field, Alabama (FAA LID: 1M4, 34.28027778° N, 87.60055556° W, 283m MSL) from March 10 through March 24, 2025. The ArcticShark UAS performed nine flights, including eight research flights over the AMF3 (BNF Main Site) and Supplemental Facilities to measure atmospheric state, turbulence, surface IR temperature and imagery, and aerosol number concentration and size distribution. The current data set presents a collection of turbulent parameters in the atmospheric boundary layer or lower free troposphere based on airborne measurement throughout the field campaign. The primary instruments used to create the current data set were the Aircraft Integrated Meteorological Measurement System (AIMMS-30) and the fine-wire thermocouple probe.

Atmosphere↗

CHELAX-BNF: Turbulent Parameters by airborne measurements

The original data were collected on board the ARM Aerial Facility ArcticShark uncrewed aerial system (UAS; https://www.arm.gov/capabilities/observatories/aaf/uas ) during the “Characterizing HEterogeneous Land-Atmosphere eXchanges at BNF” field campaign (CHEAX-BNF; https://arm.gov/research/campaigns/aaf2025CHELAX-BNF ). The ARM Aerial Facility ArcticShark UAS was based at the public-use airport of Posey Field, AL (FAA LID: 1M4, 34.28027778° N, 87.60055556° W, 283m MSL) from May 28 through June 23, 2025. The ArcticShark UAS performed 5 flights, including 4 research flights over the BNF Main Site (ARM Mobile Facility 3, https://arm.gov/capabilities/observatories/amf ) and Supplemental Facilities to measure atmospheric state, turbulence, surface IR temperature and imagery, aerosol number concentration, and aerosol size distribution. The current data set presents a collection of turbulent parameters in the atmospheric boundary layer or lower free troposphere based on airborne measurement throughout the field campaign. The primary instruments used to create the current data set were the Aircraft Integrated Meteorological Measurement System (AIMMS-30) and the fine-wire thermocouple probe.

Aircraft Integrated Meteorological Measurement Sys↗

Developing Asparagaceae1726: An Asparagaceae‐specific probe set targeting 1726 loci for Hyb‐Seq and phylogenomics in the family

Abstract Premise Target sequence capture (Hyb‐Seq) is a cost‐effective sequencing strategy that employs RNA probes to enrich for specific genomic sequences. By targeting conserved low‐copy orthologs, Hyb‐Seq enables efficient phylogenomic investigations. Here, we present Asparagaceae1726—a Hyb‐Seq probe set targeting 1726 low‐copy nuclear genes for phylogenomics in the angiosperm family Asparagaceae—which will aid the often‐challenging delineation and resolution of evolutionary relationships within Asparagaceae. Methods Here we describe and validate the Asparagaceae1726 probe set (https://github.com/bentzpc/Asparagaceae1726) in six of the seven subfamilies of Asparagaceae. We perform phylogenomic analyses with these 1726 loci and evaluate how inclusion of paralogs and bycatch plastome sequences can enhance phylogenomic inference with target‐enriched data sets. Results We recovered at least 82% of target orthologs from all sampled taxa, and phylogenomic analyses resulted in strong support for all subfamilial relationships. Additionally, topology and branch support were congruent between analyses with and without inclusion of target paralogs, suggesting that paralogs had limited effect on phylogenomic inference. Discussion Asparagaceae1726 is effective across the family and enables the generation of robust data sets for phylogenomics of any Asparagaceae taxon. Asparagaceae1726 establishes a standardized set of loci for phylogenomic analysis in Asparagaceae, which we hope will be widely used for extensible and reproducible investigations of diversification in the family.

Plant Sciences↗

Evaluating multistation phase picking algorithm phase neural operator (PhaseNO) on local seismic networks

Reliable automatic phase picking is important for many seismic applications. With the development of machine learning approaches, many algorithms are proposed, evaluated and applied to different areas. Many of these algorithms are single station based, while recent proposed methods start to combine surrounding stations into consideration in the problem of phase picking. Among these algorithms, the phase neural operator (PhaseNO) shows promising results on regional data sets comparing to existing algorithms. But there are many use cases for the local seismic networks in our community, therefore in this paper we evaluate the performance of PhaseNO on four different local data sets and compare the results to PhaseNet and EQTransformer. We used both individual phase picking metrics as well as association metrics to illustrate the performance of PhaseNO. By manually reviewing the newly detected events, we find that the PhaseNO model outperforms the single station-based approaches in the local-scale use cases due to its consideration of coherent signals from multiple stations. We also explored PhaseNO’s behaviours when only using one station, as well as gradually increasing the number of stations in the seismic network to better understand its behaviour. Overall, using the off-the-shelf machine learning based phase pickers, PhaseNO demonstrated its good performance on local-scale seismic networks.

58 GEOSCIENCES↗

New scaling and nuclear structure aspects in heavy-ion fusion reactions

Three new behaviors have been found in comparisons of fusion cross sections for different collision systems. root (1) Replacing the energy E with a scaling one, E scal = (E-V g )/($\sqrt{2}$W g ), is successful for washing out the Coulomb interaction in the spectra of fusion cross sections, where V g and W g are barrier height and width of the single-Gaussian barrier distribution model. (2) In a representation of σE vs the scaling energy, E scal , all data sets display in parallel. Here, the ratio for sigma E from any two fusion systems over the whole range is a constant value. That behavior is also studied in another representation, in which the data sets display as parallel horizontal lines for any heavy-ion fusion system. (3) The constant ratio value is the ratio of parameter products, $R^2_gW_g$, of the two systems; where R g is the barrier radius obtained in the single-Gaussian barrier distribution model. Moreover, when comparing neighboring collision systems at the same E scal , the ratio of sigma is near a constant value within a few percent over the whole range. Thus a quantitative comparison for the fusion enhancement for neighboring systems is developed. The present finding could be beneficial for predicting unmeasured fusion cross sections.

Jiang, C. L. [Argonne National Laboratory (ANL), A↗

Generative AI for Grid Operations [Slides]

In the last few years, the development and use of generative artificial intelligence (AI) and large-language models (LLMs) have changed the landscape of how AI and machine learning (ML) are being used in power systems. LLMs are built on foundational models based on large data sets that can be trained to provide information rapidly and through simple natural language prompts. Generative AI can then perform human-like tasks using ML models to identify and mimic pattens in the data sets. This presentation explores how generative AI can enhance grid operations by improving forecasts, enabling rapid contingency analyses, and offering real-time operational suggestions. By providing grid operators with valuable insights, generative AI will empower them to manage power systems more effectively.

24 POWER TRANSMISSION AND DISTRIBUTION↗

In-medium changes of nucleon cross sections tested in neutrino-induced reactions

Historically, studied in the context of heavy-ion collisions, the extent to which free nucleon-nucleon cross sections are modified in-medium remains undetermined by these data sets. Therefore, we investigate the impact of N N in-medium modifications on neutrino-nucleus cross section predictions using the GiBUU transport model. We find that including an in-medium lowering of the N N cross section and density dependence on Δ excitation improves agreement with MicroBooNE neutrino-argon scattering data. This is observed for both proton and neutral pion spectra in charged-current muon neutrino and neutral-current single pion production data sets. The impact of collision broadening of the Δ resonance is also investigated. The absence of Δ broadening is slightly favored, but the larger uncertainties on the pion production data prevent definitive conclusions. Published by the American Physical Society 2024

Bogart, B. (ORCID:0000000305588934)↗

Model-independent search for pair production of new bosons decaying into muons in proton-proton collisions at $\sqrt{s}$ = 13 TeV

The results of a model-independent search for the pair production of new bosons within a mass range of 0.21 < m < 60 GeV, are presented. This study utilizes events with a four-muon final state. We use two data sets, comprising 41.5 fb −1 and 59.7 fb −1 of proton-proton collisions at $\sqrt{s}$ = 13 TeV, recorded in 2017 and 2018 by the CMS experiment at the CERN LHC. The study of the 2018 data set includes a search for displaced signatures of a new boson within the proper decay length range of 0 < cτ < 100 mm. Our results are combined with a previous CMS result, based on 35.9 fb −1 of proton-proton collisions at $\sqrt{s}$ = 13 TeV collected in 2016. No significant deviation from the expected background is observed. Results are presented in terms of a model-independent upper limit on the product of cross section, branching fraction, and acceptance. The findings are interpreted across various benchmark models, such as an axion-like particle model, a vector portal model, the next-to-minimal supersymmetric standard model, and a dark supersymmetric scenario, including those predicting a non-negligible proper decay length of the new boson. In all considered scenarios, substantial portions of the parameter space are excluded, expanding upon prior results.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

WUS324: Multiscale Full Waveform Inversion Approaching Convergence Improves Waveform Fits While Imaging Seismic Structure of the Western United States

Abstract We report a new model of radially anisotropic crustal and upper mantle structure of the western United States (WUS324) obtained from full waveform inversion of earthquake data. We ran three multiscale inversion stages beyond model WUS256 (Rodgers et al., 2022, https://doi.org/10.1029/2022jb024549 ) allowing them to approach convergence to fit a larger data set to a shorter minimum period of 16 s. WUS324 is based on 324 total iterations from its starting model, significantly more (16 times) than previous studies. Waveform misfit reductions are 66%–70% for both the inversion data and an independent validation data set providing confidence in the predictive power of the model. WUS324 provides much better fits and reveals shear wavespeed, v S , structure of this large region with more detail than previous waveform tomography models. We show representative images demonstrating the resolution of diverse seismic structure across this highly heterogeneous region including oceanic lithosphere, subducting slabs and continental magmatism.

58 GEOSCIENCES↗

Comparative clumped isotope temperature relationships in freshwater carbonates

Lacustrine, riverine and spring carbonates represent archives of terrestrial climates and their geochemistry has been used to study palaeoenvironments. Clumped isotope thermometry is an emerging tool that has been applied to freshwater carbonates. Limited work has been done to evaluate comparative relationships between clumped isotopes and temperature in different types of modern freshwater carbonates. This study assembles an extensive calibration data set with 135 samples of modern freshwater carbonates from 96 sites and constrains the relationship between independent observations of water temperature and the clumped isotopic composition of carbonates (denoted by Δ 47 ), including new measurements, and recalculates published data in accordance with current community-defined standard values. For temperature reconstruction, the study reports a composite freshwater calibration and material-specific calibrations for biogenic carbonates (freshwater gastropods and bivalves), fine-grained carbonate (e.g. micrites), biologically mediated carbonates (microbialites and tufas) and travertines. Material-specific calibration trends show a convergence of slopes that are in agreement with recently published syntheses, but statistically significant differences in intercepts occur between some materials (e.g. some biogenics, fine-grained carbonates). These differences may arise due to unresolved seasonal biases, kinetic isotope effects and/or varying degrees of biological influence. The impact of different calibrations is shown through application to new data for glacial and deglacial age travertines from Austria and published data sets. While material-specific calibrations may yield more accurate results for biogenic and fine-grained carbonate samples, the use of material-specific and the composite freshwater calibrations generally produces values within 1.0–1.5°C of each other, and typically fall within calibration uncertainty given limitations of precision.

54 ENVIRONMENTAL SCIENCES↗

Forecasting constraints on the high-z IGM thermal state from the Lyman-α forest flux autocorrelation function

ABSTRACT The autocorrelation function of the Lyman-$\alpha$ (Ly $\alpha$) forest flux from high-z quasars probes the small-scale structure of the intergalactic medium (IGM). The thermal state of the IGM, determined by the physics of reionization, sets the small-scale power observed in the Ly $\alpha$ forest. To explore the sensitivity of the autocorrelation function to the IGM’s thermal state, we compute the autocorrelation function from a cosmological hydrodynamical simulation with an instantaneous reionization model and 135 post-processed thermal states. Using mock data sets of 20 quasars, we forecast constraints on $T_0$ and $\gamma$, which characterize the post-processed IGM thermal state, at $5.4 \le z \le 6$. While this model simplifies the IGM’s thermal state, it serves as a key first step in assessing future observational prospects. We also perform an inference test on mocks and re-weight out posterior distributions to guarantee that they exhibit statistically correct behaviour. At $z = 5.4$, we find that an idealized data set constrains $T_0$ to 59 per cent and $\gamma$ to 16 per cent at the 1$\sigma$ equivalent confidence level. To explore more realistic, non-instantaneous reionization scenarios, we analyse four models combining temperature and ultraviolet background (UVB) fluctuations at $z = 5.8$. We find that mock data generated from a model with both temperature and UVB fluctuations can rule out a model with only temperature fluctuations at the $> 1\sigma$ level 73.9 per cent of the time.

Wolfson, Molly↗

DONKEY: A Flexible and Accurate Algorithm for Clustering

We propose an accurate clustering algorithm suitable for the varied and multidimensional data sets that correspond to temporal snapshots from on-the-fly nonadiabatic trajectory-based simulations of photoexcited dynamics. The algorithm approximates the underlying probability density function using variable kernel density estimation, with local maxima corresponding to cluster centers. Each data point is then assigned to one of the maxima by employing a maximization procedure. Finally, clusters artificially separated by minor fluctuations in the probability density are merged. The algorithm does not require parameter tuning, which ensures flexibility and reduces the risk of bias. It is tested on several synthetic data sets, where it consistently outperforms conventional clustering algorithms. As a final example, the algorithm is applied to the excited dynamics of the norbornadiene ⇌ quadricyclane (C 7 H 8 ) molecular photoswitch, demonstrating how distinct reaction pathways can be identified.

algorithms↗

Ice Phase Classification Made Easy with Score-Based Denoising

Accurate identification of ice phases is essential for understanding various physicochemical phenomena. However, such classification for structures simulated with molecular dynamics is complicated by the complex symmetries of ice polymorphs and thermal fluctuations. For this purpose, both traditional order parameters and data-driven machine learning approaches have been employed, but they often rely on expert intuition, specific geometric information, or large training data sets. In this work, we present an unsupervised phase classification framework that combines a score-based denoiser model with a subsequent model-free classification method to accurately identify ice phases. Further, the denoiser model is trained on perturbed synthetic data of ideal reference structures, eliminating the need for large data sets and labeling efforts. The classification step utilizes the smooth overlap of atomic position (SOAP) descriptors as the atomic fingerprint, ensuring Euclidean symmetries and transferability to various structural systems. Our approach achieves a remarkable 100% accuracy in distinguishing ice phases of test trajectories using only seven ideal reference structures of ice phases as model inputs. This demonstrates the generalizability of the score-based denoiser model in facilitating phase identification for complex molecular systems. The proposed classification strategy can be broadly applied to investigate structural evolution and phase identification for a wide range of materials, offering new insights into the fundamental understanding of water and other complex systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗