Search NASA⌕ Search

SEARCH · Search NASA

Results for “Common data models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

TPCpp-10M: Simulated proton-proton collisions in a time projection chamber for AI foundation models

Scientific foundation models hold great promise for advancing nuclear and particle physics by improving analysis precision and accelerating discovery. Yet, progress in this field is often limited by the lack of openly available large scale datasets, as well as standardized evaluation tasks and metrics. Furthermore, the specialized knowledge and software typically required to process particle physics data pose significant barriers to interdisciplinary collaboration with the broader machine learning community. This work introduces a large, openly accessible dataset of 10 million simulated proton-proton collisions, designed to support self-supervised training of foundation models. To facilitate ease of use, the dataset is provided in a common NumPy format. In addition, it includes 70,000 labeled examples spanning three well defined downstream tasks: track finding, particle identification, and noise tagging, to enable systematic evaluation of the foundation model's adaptability. The simulated data are generated using the Pythia Monte Carlo event generator at a center of mass energy of $\sqrt{s}$ = 200 GeV and processed with Geant4 to include realistic detector conditions and signal emulation in the sPHENIX Time Projection Chamber at the Relativistic Heavy Ion Collider, located at Brookhaven National Laboratory. This dataset resource establishes a common ground for interdisciplinary research, enabling machine learning scientists and physicists alike to explore scaling behaviors, assess transferability, and accelerate progress toward foundation models in nuclear and high energy physics. The complete simulation and reconstruction chain is reproducible with the sPHENIX software stack. All data and code locations are provided under Data Accessibility.

Data Analysis, Statistics and Probability (physics↗

Deep Learning Reconstruction of Daily Soil CO 2 Efflux Reveals Biogeochemical Insights and Reduces Annual Estimate Uncertainty Despite Limited Daily Predictability

Soil CO 2 efflux is commonly measured monthly or seasonally, leaving daily dynamics poorly resolved and contributing to global estimation uncertainty. We trained a single Long Short-Term Memory (LSTM) model to predict daily soil CO 2 efflux across 82 globally distributed sites in COSORE, with 0.2%–46.9% daily data coverage from 2003 to 2020. Despite using far fewer sites than are typically used to train a single deep learning model, with observations biased toward temperate mesic sites, the LSTM model performed well at approximately one-third of sites, reconstructed nearly 2 decades of daily efflux, and outperformed commonly used approaches for estimating daily efflux when applied to the same data set. Performance was weakest at pronounced peaks and troughs and at non-temperate sites with <1.5 years of observations and irregular data patterns. Nevertheless, annual efflux from reconstructed daily data had <40% error even at underperforming sites, substantially improving estimates derived from monthly and seasonal sampling (maximum errors of 95% and 136%, respectively). Temperature sensitivity (Q 10 ) estimated from reconstructed daily predictions closely matched estimates from daily observations, whereas Q 10 values derived from monthly or seasonal observations deviated substantially, suggesting that coarse temporal sampling may contribute to uncertainty in reported Q 10 values. Consistent daily reconstructions further enabled trend analyses for well-performing, predominantly temperate sites and showed increasing soil CO 2 efflux at most sites from 2003 to 2020, with more variable summer trends. Despite limitations, these results demonstrate the potential of LSTM models to reconstruct daily soil CO 2 efflux and reduce estimation uncertainties from sparse observations.

Smykalov, Valerie [Pennsylvania State University, ↗

Data Format and Descriptions for the Alabama Carbon Storage: Data Sharing and Engagement Project

The Alabama Carbon Storage: Data Sharing and Engagement (ACS-DSE) project seeks to develop publicly accessible geologic carbon storage models and data across the southern Gulf Coastal Plain of Alabama. The public online platform developed for this project will include geologic, geophysical, infrastructure, and other relevant datasets and geologic models of the study area. Datasets, model surfaces (e.g. structural contour maps, isolith maps, porosity maps), and infrastructure data (e.g. offshore pipelines, field boundaries) will be downloadable in commonly used file formats. The anticipated primary geologic datasets are well headers, formation tops, average reservoir properties, and core analyses; these will be available as commaseparated values (CSV) text files and MS Excel workbooks. Geophysical logs will be available in Log ASCII Standard (LAS) file format. Modeled surfaces, such as structure contour maps, will be available in ArcGIS formats and text files. Infrastructure data will be available as ArcGIS shapefiles. This document provides information on the data sources and attributes of the datasets.

01 COAL, LIGNITE, AND PEAT↗

Simplex‐based model for nanoparticle grain identification in four‐dimensional scanning transmission electron microscopy data

Grain identification in polycrystalline nanoparticles, for example, determining which crystal phases are present at each spatial location, is fundamental to materials characterisation. This is particularly challenging when grains overlap extensively, as commonly occurs in four-dimensional scanning transmission electron microscopy (4D-STEM) datasets. We propose a simplex-based model (SBM) in which each simplex vertex represents the diffraction pattern (DP) of a pure grain, and the simplex edges and interior represent overlapping grains. Our SBM grain identification algorithm operates on the Bragg disk (BD) data matrix distilled from the 4D-STEM data to identify the grain membership at each scan position, together with a BD feature matrix whose columns represent the DPs for each constituent grain, which is important for identifying the crystal structure of each grain. We solve the model using a two-stage algorithm. In Stage 1, we adapt a linear mixing algorithm to estimate an initial BD feature matrix whose columns represent DPs of potentially overlapping grains. Our Stage 2 algorithm incorporates sparsity considerations to transform the initial BD feature matrix so that its columns represent DPs of pure grains. Using simulated datasets with various grain configurations, we demonstrate that SBM recovers both the BD feature matrix and membership maps more accurately than existing methods, even when a grain lacks any pure region and completely overlaps with other grains.

4D-STEM segmentation↗

Improving neutrino-nuclei interaction models: Recommendations and case studies on Peelle’s Pertinent Puzzle

Improving the modeling of neutrino-nuclei interactions using data-driven methods is crucial for high-precision neutrino oscillation experiments. This paper investigates Peelle’s Pertinent Puzzle (PPP) in the context of neutrino measurements, a longstanding challenge to fitting theoretical models to experimental data. Inconsistencies in data-model comparisons hinder efforts to enhance the accuracy and reliability of model predictions. We analyze various sources contributing to these inconsistencies and propose strategies to address them, supported by practical case studies. We advocate for incorporating model fitting exercises as a standard practice in cross section publications to enhance the robustness of results. We use a common analysis framework to explore PPP-related challenges with MicroBooNE and T2K data in an unified manner. Our findings offer valuable insights for improving the accuracy and reliability of neutrino-nuclei interaction models, particularly by systematically tuning models using data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Model Residuals as Shields: A Two-Level Formulation to Defend Smart Grids From Poisoning Attacks

The advancement of smart grids presents both vast opportunities and heightened cybersecurity risks. Data-driven defense mechanisms, though designed as a shield against these threats, can fall prey to poisoning attacks. We delve into regression settings, underscoring the imperative to fortify defenses against a spectrum of poison ratios, notably those above 0.5—an issue scarcely addressed in prior studies. Recognizing the susceptibilities of smart grids and their manipulable sensors, we exploit the very intent of poisoning attacks, compromising model accuracy, as our defense mechanism. Our proposed two-level optimization framework discerns between poisoned and authentic data based on model residuals, outperforming or matching existing methods in 72% to 77% of precision and 75% to 80% of recalls across various poisoning attacks, poison ratios, and datasets. Once the authentic data are identified, the trained model is adaptable for a variety of applications. Comprehensive evaluations on different smart grid datasets, pitted against myriad poisoning schemes, validate our methodology’s edge over existing methods. Here, we also shed light on the implications of model misspecification originating from temporal auto-correlation, a common feature in Internet of Things and smart grid data.

Adversarial machine learning (ML)↗

AI in Astrophysics: Tackling Domain Shift, Model Robustness and Uncertainty

Artificial Intelligence (AI) is revolutionizing physics research from probing the large-scale structure of the Universe to modeling subatomic interactions and fundamental forces. Yet, a major challenge persists: AI models trained on simulations or old experiment / astronomical survey often perform poorly when applied to new data exposing issues of dataset (domain) shift, model robustness, and uncertainty in predictions. This talk will introduce common challenges in applying AI across domains and present solutions based on domain adaptation a set of techniques designed to improve model generalization under domain shift. We will cover foundational ideas, practical strategies, and current research frontiers in this area. Through examples in astrophysics, we'll explore how domain adaptation can help bridge the gap between synthetic and real-world data, improve trust in model outputs, and advance scientific discovery.

Ciprijanvoic, Aleksandra [Fermilab] (ORCID:0000000↗

Collective excitations and low-energy ionization signatures of relativistic particles in silicon detectors

Abstract Solid-state detectors with a low energy threshold have several applications, including searches of non-relativistic halo dark-matter particles with sub-GeV masses. When searching for relativistic, beyond-the-Standard-Model particles with enhanced cross sections for small energy transfers, a small detector with a low energy threshold may have better sensitivity than a larger detector with a higher energy threshold. In this paper, we calculate the low-energy ionization spectrum from high-velocity particles scattering in a dielectric material. We consider the full material response including the excitation of bulk plasmons. We generalize the energy-loss function to relativistic kinematics, and benchmark existing tools used for halo dark-matter scattering against electron energy-loss spectroscopy data. Compared to calculations commonly used in the literature, such as the Photo-Absorption-Ionization model or the free-electron model, including collective effects shifts the recoil ionization spectrum towards higher energies, typically peaking around 4–6 electron-hole pairs. We apply our results to the three benchmark examples: millicharged particles produced in a beam, neutrinos with a magnetic dipole moment produced in a reactor, and upscattered dark-matter particles. Our results show that the proper inclusion of collective effects typically enhances a detector’s sensitivity to these particles, since detector backgrounds, such as dark counts, peak at lower energies.

Physics↗

Resource Assessment for Distributed Wind Energy: An Evaluation of Best-Practice Methods in the Continental US

Current wind resources within the United States (US) indicate a potential to profitably install nearly 1,400 gigawatts of distributed wind (DW) capacity. This amount is equivalent to over half of the United States’ current energy demand from electricity, making it enough to power millions of homes and businesses and replace countless fossil fuel-based generating plants. Despite the potential growth of DW in the US, deployments are presently hindered by a lack of confidence in resource estimation methods. One potential challenge is that smaller-scale turbines, with hub heights of 40 meters or less, are disproportionately impacted by obstacles such as buildings and vegetation. These obstacles may produce complex wake effects, best modeled with high-fidelity complex fluid dynamics (CFD) models that are too computationally expensive to use for routine siting and resource assessment. Thus, installers today make use of heuristics and simple equations to approximate the impact of obstacles while also leveraging long-term resource data from commercial or publicly available atmospheric models. This study evaluates these historical and commonly used methods alongside new lower-order obstacle models produced from CFD simulations and measurement-based bias correction. The preliminary results from this study show the importance of taking care in the choice and application of mesoscale atmospheric models and the significant value of bias correction using measurements from nearby meteorological towers. Detailed obstacle modeling provides only modest additional gains in performance and, in some cases, can add error, especially at sites where turbines have already been located to avoid obvious impact from upwind obstacles. These findings reinforce the importance of collecting in situ measurements and suggest that obstacle models may be better applied in practice to automated or computer-aided siting, rather than in economic wind resource assessments.

17 WIND ENERGY↗

Measurement of event shapes in minimum-bias events from proton-proton collisions at $\sqrt{s}$ = 13

A measurement of event-shape variables is presented, using a data sample produced in a special run with approximately one inelastic proton-proton collision per bunch crossing. The data were collected with the CMS detector at a center-of-mass energy of 13 TeV, corresponding to an integrated luminosity of 64 μ⁢b −1 . A number of observables related to the overall distribution of charged particles in the collisions are corrected for detector effects and compared with simulations. Inclusive event-shape distributions, as well as differential distributions of event shapes as functions of charged-particle multiplicity, are studied. None of the models investigated are able to satisfactorily describe the data. Moreover, there are significant features common amongst all generator setups studied, particularly showing data being more isotropic than any of the simulations. Multidimensional unfolded distributions are provided, along with their correlations.

Chekhovsky, V. [Yerevan Physics Institute]↗

Learning Constitutive Relations From Soil Moisture Data via Physically Constrained Neural Networks

Abstract The constitutive relations of the Richardson‐Richards equation encode the macroscopic properties of soil water retention and conductivity. These soil hydraulic functions are commonly represented by models with a handful of parameters. The limited degrees of freedom of such soil hydraulic models constrain our ability to extract soil hydraulic properties from soil moisture data via inverse modeling. We present a new free‐form approach to learning the constitutive relations using physically constrained neural networks. We implemented the inverse modeling framework in a differentiable modeling framework, JAX, to ensure scalability and extensibility. For efficient gradient computations, we implemented implicit differentiation through a nonlinear solver for the Richardson‐Richards equation. We tested the framework against synthetic noisy data and demonstrated its robustness against varying magnitudes of noise and degrees of freedom of the neural networks. We applied the framework to soil moisture data from an upward infiltration experiment and demonstrated that the neural network‐based approach was better fitted to the experimental data than a parametric model and that the framework can learn the constitutive relations.

54 ENVIRONMENTAL SCIENCES↗

Quality Guidelines for Energy System Studies: Process Modeling Design Parameters

The National Energy Technology Laboratory (NETL) conducts systems analysis studies that require a large number of inputs, from ambient conditions to parameters for Aspen Plus ® (Aspen) process blocks. The sheer number of assumptions required makes it impractical to document all of them in each issued report. The purpose of the Quality Guidelines for Energy System Studies (QGESS) is to document the assumptions most commonly used in system analysis studies and the basis for those assumptions. In order to develop the systems analysis models presented in various NETL reports, significant vendor data have been obtained, and these data enhance the model outputs. Much of the vendor data obtained are considered proprietary and not suitable for public release or attribution to a specific vendor. As such, several sub-systems common in NETL reports and their process parameter data are not reported in this document to protect proprietary vendor information. The values and ranges of values presented in this report represent assumptions that have been made in previous studies.

97 MATHEMATICS AND COMPUTING↗

Bias Correcting NOAA's High-Resolution Rapid Refresh (HRRR) Wind Resource Data for Grid Integration Applications [Slides]

Many weather years of high-quality wind data are widely accepted in the grid integration community to be important for studying wind energy technical potential, energy system operations, and grid resilience. NREL makes high-quality wind and solar resource data available. NREL's Grid-Atmosphere workshop (March 2024) identified NREL National Solar Radiation Database as widely used in grid integration modeling, but there is less agreement on commonly used wind datasets. One important factor identified by ESIG's 2023 report 'Weather Dataset Needs for Planning and Analyzing Modern Power Systems' for gold standard wind data is regular updates. To address the need for regular updates, NREL's team can now process all currently available and regularly updated High-Resolution Rapid Refresh (HRRR) outputs. HRRR is an hourly-updated operational forecast product produced by the National Oceanic and Atmospheric Administration (NOAA) (Dowell et al., 2022). One barrier to NREL using HRRR is systematic bias and consistency with NREL's existing wind datasets (e.g. WIND Toolkit, 'WTK') across weather years. To address this barrier, we show that the HRRR can be interpolated and bias-corrected to be consistent with NRE's existing datasets. We call the new dataset BC-HRRR (bias-corrected HRRR). As with historical datasets like the WTK, BC-HRRR is intended for use in grid integration modeling (e.g., capacity expansion, production cost, and resource adequacy modeling). BC-HRRR's (2015-present) consistency with WTK (2007-2013) allows NREL to extend internal grid integration tooling with 15+ weather years of wind data with low-overhead extensibility to future years as they are made available by NOAA. The rest of this slide deck documents the BC-HRRR processing methods, validation, and its implications for intended use.

17 WIND ENERGY↗

An improved dataset for predicting mammal infecting viruses from genetic sequence information

There have been several attempts to develop machine learning (ML) models to identify human infecting viruses from their genomic sequences, with varying degrees of success. Direct comparison between models is problematic, because these models are typically trained and evaluated on different datasets with alternative data splitting schemes, features, and model performance metrics. In this paper we present a standardized dataset of mammal infecting and non-infecting viral pathogens, refined from the previous work of Mollentze et al. to include the latest literature evidence, roughly doubling the number of curated host-virus records available to the community, and new host target labels, primate and mammal. The new host labels were included for several reasons, including previous reports that classification performance is better at broader taxonomic ranks and the idea that there may be more data for primate infection that might serve as a suitable proxy for zoonotic potential and avoidance of false positives for human infection due to absence of evidence. On this dataset, we report the performance of eight machine learning models for predicting mammal-infecting viruses from their genomic sequences. We find that randomly assigning cases in our improved dataset to training/testing sets, when compared to the original assignments into training/testing in Mollentze et al., increases the overall average ROC AUC of prediction of human infection from 0.663 ± 0.070 to 0.784 ± 0.013, consistent with the reduction in phylogenetic distance between train and test sets (relative entropy change from 3.00 to 0.08). The broadest host category of mammal infection can be predicted most reliably at 0.850 ± 0.020. We share our improved dataset and code to enable standardized comparisons of machine learning methods to predict human host infections. Overall, we have presented preliminary evidence that classification of virus host infection is more tractable at higher taxonomic ranks, that unsurprisingly reducing the phylogenetic distance between training and test sets can improve predictive performance, that peptide kmer features appear to be harmful to out of sample model performance, and we are left with the question of whether models for virus host prediction can reasonably be expected to perform well in out of sample scenarios given the likelihood that viruses do not share a common ancestor. Consistent with this concern, when the data is resampled such that there is no overlap between viral families in training and test sets (relative entropy > 24), models perform no better than random chance at prediction of human infection regardless of whether kmers are included (ROC AUC 0.50 ± 0.08) or not (ROC AUC 0.50 ± 0.04).

59 BASIC BIOLOGICAL SCIENCES↗

A Reduced-form Cost Model for Prefeasibility Analysis of Hydropower at Non-Powered Dams

This study presents a reduced-form model to support a better understanding of the capacity potential and drivers of costs for hydropower development at U.S. non-powered dams (NPD), which are existing dams that are not currently used for hydropower. With information on nineteen reference sites, a set of reduced-form design and cost equations were estimated to enable rapid assessments of aggregate costs and their components for a large number of NPD sites. The model was then applied to 36,000+ potential U.S. NPD sites using the limited data commonly available. Although the cost estimates span a wide range, there exists a significant amount of U.S. NPD hydropower capacity potential, which are considered cost-competitive in the current market using baseline technologies.

13 HYDRO ENERGY↗

W-Band ARM Scanning Cloud Radar (WSACR) 2nd Generation CF-Radial Spectral Data, Vertically-Pointing Scan, Cross-Polarization Model (a1)

ARM's scanning cloud radars are fully coherent dual-frequency, dual-polarization Doppler radars mounted on a common scanning pedestal. Each pedestal includes a Ka-band radar (2kW peak power) and the deployment location determines whether the second radar is a W-band (WSACR; 1.7 kW peak power) or X-band (XSACR; 20 kW peak power). Beamwidths at Ka and W bands are roughly matched at 0.3 degrees. Due to the narrow antenna beamwidth, ARM’s scanning cloud radars use scanning strategies that are unlike typical weather radars. Rather than focusing on plan position indicator, or PPI, scans, the Ka-SACR uses range height indicator, or RHI, scans at numerous azimuths to obtain cloud volume data. Measurements collected with the W-SACR are copolar and cross-polar radar reflectivity, Doppler velocity, spectra width and spectra when not scanning, and linear depolarization ration.

54 ENVIRONMENTAL SCIENCES↗

The Impact of Time-Aware Design Choices in ICS Anomaly Detection

Industrial control systems (ICS) remain vulnerable to increasingly sophisticated cyberattacks, yet evaluating anomaly detection models in these environments is challenging due to temporal dependencies, missing-not-at-random patterns, and extremely imbalanced datasets. These factors make common practices—especially random data splits and na¨ıve imputation— prone to severe temporal leakage, which can inflate reported performance and obscure real-world limitations. In this work, we systematically examine classical machine learning models, temporal deep learning architecture, and tensordecomposition– based methods on a gas-pipeline dataset using a fully temporally separated evaluation pipeline designed to mimic realistic deployment conditions. Our findings show that proper temporal handling and MNAR-aware preprocessing significantly alter the relative performance of popular anomaly-detection methods, providing practical guidance for designing reliable, leakage-resistant ICS intrusion-detection systems.

97 MATHEMATICS AND COMPUTING↗

A filter-dependent granular temperature model from large-scale CFD-DEM data

The computational study of strongly-coupled, gas–solid flows at scales relevant to most environmental and engineering applications requires the use of ‘coarse-grained’ methodologies such as the two-fluid model, particle-in-cell approach or the multiphase Reynolds Averaged Navier–Stokes equations. While these strategies enable computations at desirable length- and time-scales, they rely heavily on models to capture important flow physics that occur at scales smaller than the mesh. To date, the models that do exist are based on a limited set of flow conditions, such as very dilute particle phase. To this end, we leverage a large-scale repository of CFD-DEM data to develop filter-size dependent models for the mean variance in particle volume fraction, a quantity commonly used to assess the degree of clustering, and the granular temperature, a key quantity for accurately predicting gas–solid flows. In conclusion, because of its filter-size dependence, the granular temperature model can be directly translated to coarse-grained approaches and tied directly to grid size.

AMReX↗