Search NASA⌕ Search

SEARCH · Search NASA

Results for “validation dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Heterogeneous Multi-Domain Dataset Synthesis to Facilitate Privacy and Risk Assessments in Smart City IoT

The emergence of the Smart Cities paradigm and the rapid expansion and integration of Internet of Things (IoT) technologies within this context have created unprecedented opportunities for high-resolution behavioral analytics, urban optimization, and context-aware services. However, this same proliferation intensifies privacy risks, particularly those arising from cross-modal data linkage across heterogeneous sensing platforms. To address these challenges, this paper introduces a comprehensive, statistically grounded framework for generating synthetic, multimodal IoT datasets tailored to Smart City research. The framework produces behaviorally plausible synthetic data suitable for preliminary privacy risk assessment and as a benchmark for future re-identification studies, as well as for evaluating algorithms in mobility modeling, urban informatics, and privacy-enhancing technologies. As part of our approach, we formalize probabilistic methods for synthesizing three heterogeneous and operationally relevant data streams—cellular mobility traces, payment terminal transaction logs, and Smart Retail nutrition records—capturing the behaviors of a large number of synthetically generated urban residents over a 12-week period. The framework integrates spatially explicit merchant selection using K-Dimensional (KD)-tree nearest-neighbor algorithms, temporally correlated anchor-based mobility simulation reflective of daily urban rhythms, and dietary-constraint filtering to preserve ecological validity in consumption patterns. In total, the system generates approximately 116 million mobility pings, 5.4 million transactions, and 1.9 million itemized purchases, yielding a reproducible benchmark for evaluating multimodal analytics, privacy-preserving computation, and secure IoT data-sharing protocols. To show the validity of this dataset, the underlying distributions of these residents were successfully validated against reported distributions in published research. We present preliminary uniqueness and cross-modal linkage indicators; comprehensive re-identification benchmarking against specific attack algorithms is planned as future work. This framework can be easily adapted to various scenarios of interest in Smart Cities and other IoT applications. By aligning methodological rigor with the operational needs of Smart City ecosystems, this work fills critical gaps in synthetic data generation for privacy-sensitive domains, including intelligent transportation systems, urban health informatics, and next-generation digital commerce infrastructures.

IoT↗

Solar Radiation Research Laboratory (SRRL) Core Project Final Report: Fiscal Years 2022-2024

The Solar Radiation Research Laboratory (SRRL) at the National Laboratory of the Rockies (NLR) is a world-leading solar calibration and measurement facility and maintains and disseminates the World Radiation Reference (essentially the W/m2) for the United States that is essential for traceable and accurate measurements of solar radiation at all solar generation facilities. SRRL operates two calibration facilities that meet International Standards Organization-17025 (ISO-17025) standards and provide unique high-quality calibrations to NREL and other U.S. Department of Energy laboratories. The Baseline Measurement System (BMS) at SRRL provides a high-quality record of solar irradiance and surface meteorological conditions. SRRL capabilities are used to develop (1) improved methods for the calibration of solar radiometers; (2) new standards through the ISO, the International Electrotechnical Commission (IEC), and the American Standards for Testing of Materials (ASTM) International; (c) solar radiation and meteorological models; and (d) advanced instrumentation and methods for operating solar measurement stations. The SRRL datasets are also critical for the validation of new models and datasets, such as the National Solar Radiation Database (NSRDB). The research and development of solar radiation measurement systems and resource modeling techniques are essential for advancing the scientific basis for producing reliable resource data. Specifically, the spatial, temporal, and spectral (wavelength dependency) characteristics of the solar resource are required in several different time frames for various project phases.

14 SOLAR ENERGY↗

Validation of the DESI DR2 Ly⁢ 𝛼 BAO analysis using synthetic datasets

The second data release (DR2) of the Dark Energy Spectroscopic Instrument (DESI), containing data from the first three years of observations, doubles the number of Lyman-α (Ly α) forest spectra in DR1 and it provides the largest dataset of its kind. To ensure a robust validation of the baryonic acoustic oscillation (BAO) analysis using Ly α forests, we have made significant updates compared to DR1 to both the mocks and the analysis framework used in the validation. In particular, we present CoLoRe-QL, a new set of Lyα mocks that use a quasilinear input power spectrum to incorporate the nonlinear broadening of the BAO peak. Here, we have also increased the number of realizations used in the validation to 400, compared to the 150 realizations used in DR1. Finally, we present a detailed study of the impact of quasar redshift errors on the BAO measurement, and we compare different strategies to mask damped Lyman-α absorbers in our spectra. The BAO measurement from the Ly α dataset of DESI DR2 is presented in a companion publication.

Casas, L. [Institut de Física d’Altes Energies (IF↗

Validation of the DESI DR2 Ly$\alpha$ BAO analysis using synthetic datasets

The second data release (DR2) of the Dark Energy Spectroscopic Instrument (DESI), containing data from the first three years of observations, doubles the number of Lyman-$\alpha$ (Ly$\alpha$) forest spectra in DR1 and it provides the largest dataset of its kind. To ensure a robust validation of the Baryonic Acoustic Oscillation (BAO) analysis using Ly$\alpha$ forests, we have made significant updates compared to DR1 to both the mocks and the analysis framework used in the validation. In particular, we present CoLoRe-QL, a new set of Ly$\alpha$ mocks that use a quasi-linear input power spectrum to incorporate the non-linear broadening of the BAO peak. We have also increased the number of realisations used in the validation to 400, compared to the 150 realisations used in DR1. Finally, we present a detailed study of the impact of quasar redshift errors on the BAO measurement, and we compare different strategies to mask Damped Lyman-$\alpha$ Absorbers (DLAs) in our spectra. The BAO measurement from the Ly$\alpha$ dataset of DESI DR2 is presented in a companion publication.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

The DECADE cosmic shear project IV: cosmological constraints from 107 million galaxies across 5,400 deg$^2$ of the sky

We present cosmological constraints from the Dark Energy Camera All Data Everywhere (DECADE) cosmic shear analysis. This work uses shape measurements for 107 million galaxies measured through Dark Energy Camera (DECam) imaging of $5,\!412$ deg$^2$ of sky that is outside the Dark Energy Survey (DES) footprint. We derive constraints on the cosmological parameters $S_8 = 0.791^{+0.027}_{-0.032}$ and $Ω_{\rm m} =0.269^{+0.034}_{-0.050}$ for the $Λ$CDM model, which are consistent with those from other weak lensing surveys and from the cosmic microwave background. We combine our results with cosmic shear results from DES Y3 at the likelihood level, since the two datasets span independent areas on the sky. The combined measurements, which cover $\approx\! 10,\!000$ deg$^2$, prefer $S_8 = 0.791 \pm 0.023$ and $Ω_{\rm m} = 0.277^{+0.034}_{-0.046}$ under the $Λ$CDM model. These results are the culmination of a series of rigorous studies that characterize and validate the DECADE dataset and the associated analysis methodologies (Anbajagane et. al 2025a,b,c). Overall, the DECADE project demonstrates that the cosmic shear analysis methods employed in Stage-III weak lensing surveys can provide robust cosmological constraints for fairly inhomogeneous datasets. This opens the possibility of using data that have been previously categorized as ``unusable'' for cosmic shear analyses, thereby increasing the statistical power of upcoming weak lensing surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

Uncertainty quantification in multivariable regression for material property prediction with Bayesian neural networks

With the increased use of data-driven approaches and machine learning-based methods in material science, the importance of reliable uncertainty quantification (UQ) of the predicted variables for informed decision-making cannot be overstated. UQ in material property prediction poses unique challenges, including multi-scale and multi-physics nature of materials, intricate interactions between numerous factors, limited availability of large curated datasets, etc. In this work, we introduce a physics-informed Bayesian Neural Networks (BNNs) approach for UQ, which integrates knowledge from governing laws in materials to guide the models toward physically consistent predictions. To evaluate the approach, we present case studies for predicting the creep rupture life of steel alloys. Experimental validation with three datasets of creep tests demonstrates that this method produces point predictions and uncertainty estimations that are competitive or exceed the performance of conventional UQ methods such as Gaussian Process Regression. Additionally, we evaluate the suitability of employing UQ in an active learning scenario and report competitive performance. The most promising framework for creep life prediction is BNNs based on Markov Chain Monte Carlo approximation of the posterior distribution of network parameters, as it provided more reliable results in comparison to BNNs based on variational inference approximation or related NNs with probabilistic outputs.

36 MATERIALS SCIENCE↗

Ensemble cure kinetics network (ECK-Net): A method to derive cure kinetics of thermosetting resin

This paper introduces an Ensemble Cure Kinetics Network (ECK-Net), a neural network (NN)–based framework for modeling the cure kinetics of thermosetting resins within a phenomenological context. ECK-Net replaces traditional analytic models, which require extensive chemical insight and multiple isothermal/non-isothermal experiments, with a data-driven surrogate that maps nonlinear relationships between temperature, degree of cure, and reaction rate from differential scanning calorimetry data. The proposed approach predicts input-dependent kinetic coefficients of a generalized nth-order reaction equation rather than reaction rates directly, enabling a single unified model to represent various epoxy systems without relying on iso-conversional analysis or predefined functional forms. To ensure robustness, multiple independently trained networks under different random initializations are blended through an ensemble strategy, effectively mitigating the stochastic variability inherent to neural networks. The framework is validated using experimental datasets from multiple resin systems, including aerospace-grade materials (Toray 3900-2, Cycom 5320-1, and Hexcel 8552) and a windmill-grade resin (RIMR 035c). The model accurately reproduces the temporal evolution of the degree of cure under manufacturers’ recommended cure cycles across all tested resins systems, yielding Pearson’s correlation coefficients of 0.992, 0.994, 0.993, 0.997, respectively. To demonstrate process-level applicability, the trained network was implemented within the Abaqus environment to simulate out-of-autoclave (OOA) curing process of the CFRP panel composed of Toray T830H-6K/3900-2D prepreg. The simulation results showed excellent agreement with experimental temperature response (maximum peak temperature, simulation: 189.6 °C, experiment: 188.5 °C) and the final degree of cure (simulation: 0.948, experiment: 0.960 ± 0.013), confirming ECK-Net’s capability as a reliable alternative to conventional cure kinetics modeling methods.

Composite curing↗

Residual stress distribution in an additively manufactured complex structure by neutron diffraction measurement

Residual stress in an aerodynamically shaped Ni-based superalloy airfoil fabricated by laser powder bed fusion was measured by neutron diffraction. The experiment was conducted by considering the complex shape, implementing computer aided experiment planning, and automatic alignment at each rapid measurement. The 3-dimensional (3D) residual stress distribution in the airfoil is presented in this work, which lacks symmetry due to the complex geometry of the airfoil. In conclusion, the results provide theoretical thermal processing models a complete residual stress dataset of simulation validation on 3D shape complex structure.

Residual stress↗

Massively parallel reporter assays and mouse transgenic assays provide correlated and complementary information about neuronal enhancer activity

High-throughput massively parallel reporter assays (MPRAs) and phenotype-rich in vivo transgenic mouse assays are two potentially complementary ways to study the impact of noncoding variants associated with psychiatric diseases. Here, we investigate the utility of combining these assays. Specifically, we carry out an MPRA in induced human neurons on over 50,000 sequences derived from fetal neuronal ATAC-seq datasets and enhancers validated in mouse assays. We also test the impact of over 20,000 variants, including synthetic mutations and 167 common variants associated with psychiatric disorders. We find a strong and specific correlation between MPRA and mouse neuronal enhancer activity. Four out of five tested variants with significant MPRA effects affected neuronal enhancer activity in mouse embryos. Mouse assays also reveal pleiotropic variant effects that could not be observed in MPRA. Our work provides a catalog of functional neuronal enhancers and variant effects and highlights the effectiveness of combining MPRAs and mouse transgenic assays.

Kosicki, Michael↗

Text-mined dataset of solid-state syntheses with impurity phases using Large Language Model

Solid-state synthesis is widely used to obtain various inorganic materials, such as battery materials and bulk thermoelectrics. Despite its prevalence, the process remains challenging due to the lack of a general theory and well-understood underlying reaction mechanisms. While prior works have successfully extracted structured datasets from literature, they often neglect product phase purity or yield. In this work, we construct a solid-state synthesis dataset consisting of 80,806 syntheses extracted with a large language model (LLM), including 18,869 reactions with impurity phase(s). Our dataset not only validates expected thermodynamic trends for impurity phase formation but also identifies challenging cases where impurity phases emerge even when the target phase is significantly more stable.

Lee, Sanghoon↗

Statistical relationships across epigenomes using large-scale hierarchical clustering

Recent advances in genomics and sequencing platforms have revolutionized our ability to create immense data sets, particularly for studying epigenetic regulation of gene expression. However, the avalanche of epigenomic data is difficult to parse for biological interpretation given nonlinear complex patterns and relationships. This attractive challenge in epigenomic data lends itself to machine learning for discerning infectivity and susceptibility. In this study, we explore over 3000 epigenomes of uninfected individuals and provide a framework to characterize the relationships among epigenetic modifiers, their modifiers, genetic loci, and specific immune cell types across all chromosomes using hierarchical clustering. Hierarchical clustering of epigenomic data revealed consistent epigenetic patterns across chromosomes, demonstrating that variation due to epigenetic modifiers is greater than variation between cell types. Gene Ontology and KEGG pathway analyses indicated significant enrichment of genes involved in chromatin remodeling, mRNA splicing, immune responses, and the regulation of microRNAs and snoRNAs. Epigenetic modifiers frequently formed biologically relevant clusters, including the cohesin complex, RNA Polymerase II transcription factors, and PRC2 complex members. These clustering behaviors remained consistent across all chromosomes, supported by entropy analysis and high Adjusted Rand Index scores, indicating robust cross-chromosomal similarity. Co-occurrence analysis further revealed specific sets of modifiers that consistently appeared together within clusters, reflecting shared biological functions and interactions. Validation using another dataset confirmed the reproducibility of these clustering patterns and modifier co-occurrence relationships, underscoring the reliability and generalizability of the methodology.

97 MATHEMATICS AND COMPUTING↗

Leveraging BERT and Network-Based Attention Analysis for Identifying Treatment Milestones in EHRs

This study introduces a sophisticated data-driven framework for analyzing Electronic Health Records (EHRs) using transformer-based models to identify and disentangle overlapping treatment contexts. The framework leverages a preprocessing pipeline that transforms structured procedural codes into semantically enriched descriptive text, enabling the use of attention mechanisms to cluster medical events into treatment milestones—cohesive and distinct components of care processes. The methodology is rigorously validated using synthetic datasets derived from the MIMIC-III database, designed to simulate the heterogeneity and overlapping procedural contexts characteristic of real-world EHR scenarios. Quantitative evaluation highlights the framework’s robustness in disentangling concurrent care pathways, with attention metrics and unsupervised clustering approaches demonstrating the ability to preserve intra-context relationships while distinguishing inter-context dependencies. By addressing challenges inherent in data heterogeneity, this approach provides a foundation for uncovering complex treatment patterns, advancing clinical decision-making, and optimizing resource allocation in diverse healthcare environments.

Kim, Minsu [ORNL] (ORCID:0000000224185535)↗

Modeling the Nucleation and Growth of Lead Sulfate Particles on Lead Electrodes

Lead-acid batteries (LABs) play a pivotal role in the energy storage sector with applications spanning from starting-lighting-ignition batteries to grid energy storage. Passivation of lead negative electrodes by PbSO 4 particles is a fundamental mechanism limiting the performance of LAB. In this regard, an electrochemical model is developed that simulates the nucleation and growth (N&G) dynamics of PbSO 4 particles on a flat lead electrode, responsible for its passivation. The model considers the electrochemical reactions between lead electrode and sulfuric acid, N&G of PbSO 4 particles, passivation of the lead surface, and the ternary transport of PbSO 4 (aq), bisulfate, and protons in H 2 SO 4 electrolyte. The model is validated with a dataset of cyclic voltammetry (CV) responses collected at several scan rates and H 2 SO 4 concentrations. The model shows remarkable qualitative and quantitative agreement with the experimental data including CV peak features, discharge capacity, and particle size. The model was employed to explore key N&G quantities, such as supersaturation, nucleation rate, particle count, growth rate, particle size, and surface coverage, and to examine how their interactions influence electrode utilization. Parametric studies were also conducted to evaluate how scan rates and acid concentrations influence the previously mentioned N&G quantities and, subsequently, the utilization of the electrode.

25 ENERGY STORAGE↗

Data From Experiments on Bubbling Fluidization of Zeolite in a Rectangular Bubbling Fluidized Bed

Fluidization experiments were conducted in a lab-scale rectangular bubbling fluidized bed with the objective of generating a high-quality dataset for model validation and artificial intelligence/machine learning (AI/ML) training. Zeolite was chosen as the bed material, and the fluidizing medium was air as supplied by a compressor. Three different flow rates at the inlet were chosen such that the particles were fluidized but not elutriated from the system. The test matrix involved randomization and replicates to provide uncertainty estimates as well as four different batches of zeolite as the bed material. The quantities of interest obtained from this study were statistics of differential pressures, interface heights, and particle velocities. Considering all the components of the elaborate test plan, the results obtained were consistent and reproducible. Characterization tests were performed to estimate particle properties including size, density, coefficient of friction, coefficient of restitution, and minimum fluidization velocity. In addition, the angle of repose from granular discharge experiments has been reported to account for rolling friction, though its effect on the overall process is expected to be negligible.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Smart Meter Data: A Gateway for Reducing Solar Soft Costs with Model-Free Hosting Capacity Maps

Public-facing solar hosting capacity (HC) maps, which show the maximum amount of solar energy that can be installed at a location without adverse effects, have proven to be a key driver of solar soft cost reductions through a variety of pathways (e.g., streamlining interconnection, siting, and customer acquisition processes). However, current methods for generating HC maps require detailed grid models and time-consuming simulations that limit both their accuracy and scalability—today, only a handful out of almost 2,000 utilities provide these maps. This project developed and validated data-driven algorithms for calculating solar HC using data from AMI without the need of detailed grid models or simulations. The algorithms were validated on utility datasets and incorporated as an application into NRECA’s Open Modeling Framework (OMF.coop) for the over 260 coops and vendors throughout the US to use. The OMF is free and open-source for everyone.

14 SOLAR ENERGY↗

The effect of heat treatment on the defect evolution in LPBF 316H stainless steel

This work investigates the effect of processing and heat treatment on defect evolution and irradiation response of laser powder bed fusion (LPBF) 316H stainless steel, with comparisons to LPBF 316L and wrought 316L/316H. Using in-situ and ex-situ ion irradiations across a wide parameter space—temperature (300–675 °C), dose (0.2–25 dpa), dose rate (10 -3 –10 -5 dpa/s), and helium co-implantation (20–2500 appm)—we correlated void swelling, dislocation loop evolution, and segregation behavior with pre-irradiation microstructures (as-built, stress-relieved, solution-annealed, and cold-worked). Results show that swelling is strongly controlled by dislocation density: intermediate densities maximize swelling, while solution annealing or cold working reduce susceptibility to levels comparable to wrought alloys. Low dose rates and helium both promote cavity nucleation and lower the incubation barrier, with helium suppressing the role of dislocation density and driving swelling behavior toward wrought-like response. Loop evolution in SA LPBF 316H resembles wrought 316L but shows localized denuded zones near low-angle grain boundaries. STEM-EDS mapping further revealed Ni segregation at voids and sparse Al-rich oxides, without evidence of Ni–Si precipitates. Collectively, these findings identify dislocation density, helium content, and dose rate as the factors governing swelling in LPBF 316H and provide mechanistic datasets for model validation, supporting LAIN and the qualification of AM austenitic steels for nuclear service.

316 stainless steel↗

Optimizing Geospatial Assessments for Nuclear Safeguards Applications with Large Language Models

A multidisciplinary team at Argonne National Laboratory evaluated the ability of large language models (LLMs) to identify geographic locations from open-source text and assessed post-processing measures to strengthen the reliability of those extractions in support of international nuclear safeguards. The study focused on addressing challenges such as toponym ambiguity, imprecise descriptions, and misinformation, which often undermine the accuracy of LLM-derived geospatial assessments. By integrating authoritative geospatial datasets, employing rigorous validation techniques, and leveraging human-in-the-loop processes, the project aimed to enhance the precision, transparency, and reproducibility of geospatial localization workflows. The findings demonstrate that while LLMs exhibit significant potential for accelerating geospatial analysis, their outputs require systematic grounding and verification to ensure reliability in high-stakes applications. This work contributes to the broader field of geospatial intelligence and supports strategic objectives of international organizations such as the International Atomic Energy Agency (IAEA) and the U.S. Department of Energy (DOE).

97 MATHEMATICS AND COMPUTING↗

Downscaled Earth System Model Data for Resilient Energy System Planning

The second-generation Sup3rCC dataset provides high-resolution meteorological data generated through the downscaling of multiple earth system models (ESMs) from the Coupled Model Intercomparison Project Phase 6 (CMIP6). This downscaling is performed through application of a generative machine learning approach called Super-Resolution for Renewable Resource Data (sup3r). This dataset builds on the first-generation Sup3rCC data by applying improved bias correction methods and adding downscaled precipitation to the output variables. In this presentation, we explore the output characteristics of the dataset and various validation analyses. We also present and discuss plans for the integration of this data into power system planning models using a decision-making under deep uncertainty (DMDU) methodology.

97 MATHEMATICS AND COMPUTING↗