Search NASASearch

SEARCH · Search NASA

Results for “Data-driven”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Techno-Economic Evaluation of Electrified Vehicle Options in Drayage Fleets

The electrification of drayage fleets offers potential economic and operational benefits, but the financial viability of electrified vehicles remains sensitive to battery cost, energy price, and fleet usage patterns. While total cost of ownership (TCO) is a useful benchmark, fleet operators and investors are equally concerned with investment performance metrics such as payback period (PB) and Internal Rate of Return (IRR), which better reflect financial risks and investment return timelines. This study develops a unified techno-economic framework that jointly evaluates TCO, PB, and IRR to determine when electrified trucks become cost-effective alternatives to diesel trucks. Building on a previously developed cost modeling tool and using real-world telematics data from a Class 8 drayage fleet at the Port of Savannah, the analysis incorporates projected battery cost trajectories, electricity and diesel price trends, vehicle efficiency improvements, and multiple battery capacities. Parameter ranges reflect widely cited projections and observed drayage-duty-cycle variability. A surrogate-modeling method approximates economic performance across thousands of battery cost–electricity price combinations, enabling high-resolution identification of conditions that achieve TCO parity, acceptable PB thresholds, and target IRR levels. Additionally, the study estimates the evolving share of the fleet that can feasibly electrify over time under multiple economic metrics. This integrated framework offers a novel, data-driven approach to inform risk-aware decision-making for fleet electrification and supports investment planning under evolving cost and operational conditions.

Sun, Ruixiao [ORNL] (ORCID:0000000341768676)

Leveraging machine learning to enhance aerosol classification using Single-Particle Mass Spectrometry

Advancing automated classification of atmospheric aerosols from Single-Particle Mass Spectrometry (SPMS) data remains challenging due to overlapping ion signatures, compositional diversity, and limited labeled data. This study evaluates supervised and semi-supervised learning frameworks to enhance aerosol identification by jointly leveraging labeled and unlabeled spectra. Four models were compared: a supervised Support Vector Machine (SVM), a self-training SVM, a stacked autoencoder classifier, and a stacked autoencoder trained using a temporal-ensembling Mean Teacher approach. All models achieved high and stable accuracies (90.0 %–91.1 %), surpassing previous results on the same dataset (87 %) and matching the performance of state-of-the-art deep learning methods. Despite small global metric differences (≤ 1 %), semi-supervised variants yielded up to 5 %–10 % improvements for compositionally rare particle types – such as soot (0.77 % of spectra, F1-score: 0.93–0.97) and hazelnut pollen (0.98 % of spectra, F1-score: 0.97–1.00) – equating to roughly ∼ 187 additional correctly classified spectra. These gains are scientifically significant, as such rare particles exert disproportionate influence on radiative absorption and ice nucleation processes; their improved detection reduces modeled uncertainties in aerosol absorption optical depth and mixed-phase cloud ice nucleation rates. The models' residual misclassifications (≈ 9 %) largely arise from true spectral overlap among chemically adjacent species (e.g., Na- vs. K-feldspar, coated vs. uncoated feldspars), reflecting physical compositional continuity rather than algorithmic error. Collectively, these findings demonstrate that leveraging unlabeled data to learn robust spectral representations and refine classification enhances both fidelity and interpretability, bridging data-driven analysis with aerosol–climate process understanding.

54 ENVIRONMENTAL SCIENCES

mosartwmpy sample input data; 1980 - 1985

Sample input data spanning the years 1980-1985 for running the `mosartwmpy` water routing and management model: https://github.com/IMMM-SFA/mosartwmpy. More information is available in the README.md. Changelog: v0.0.8: Update with new CAP_MIN reservoir parameter, corresponds with mosartwmpy v1.0.0 v0.0.7: new reservoir files with corrected locations based on an exhaustive review by Dan Broman for the 9505 project; corresponding long term mean demand and flow files updated based on VIC4 simulations using Daymet forcing. v0.0.6: updates reservoir parameters to indicate which reservoirs should follow generic operating rules and which should follow ISTARF data-driven operating rules. v0.0.5: adds supporting files for running the Farmer Agent Based Model of adaptive water demand.

Bracken, Cameron [Pacific Northwest National Labor

A Probabilistic Approach to Load Modeling for Central HVAC Systems in Large Commercial Buildings for Retrofit Decisions Under Uncertainty

Retrofitting central HVAC systems in large commercial buildings with advanced technologies like heat recovery chillers (HRCs) offers a significant opportunity to enhance energy efficiency. However, analyzing these retrofits is challenging with traditional whole-building simulation tools, which require intensive calibration and struggle to model innovative system configurations and controls. To overcome these limitations, this study proposes a load profilebased retrofit analysis framework that provides better decisions under uncertainty. The main focus of this paper is the development of a probabilistic load profile model that can be used in the framework by using exploratory data analysis (EDA) of measured building data to properly quantify its inherent variability. A non-parametric Gaussian Process (GP) model was employed to capture the time- and weather-dependent characteristics of the heating load while explicitly modeling its uncertainty. The model's effectiveness is demonstrated through strong predictive performance on unseen data and physically interpretable insights into load behavior. This data-driven, probabilistic load profile serves as a robust and flexible input for subsequent system simulations, enabling a more confident and statistically sound analysis of retrofit potential.

Ham, S W

Plant Bioengineering Atlas: A Knowledge Graph of Genes, DNA Constructs, and Plant Traits.

Plant bioengineering has generated tens of thousands of genotype-to-phenotype relationships, but this knowledge remains fragmented across narrative literature and difficult to use computationally. Inconsistent descriptions of DNA constructs, host species, and traits, including variable species names, omitted regulatory elements, and inconsistent gene symbols, impede data reuse, comparative analysis, and design-build-test-learn cycles. Here, we present the Plant Bioengineering Atlas, a literature-mined, ontology-grounded knowledge base assembled using an artificial intelligence (AI)-aided extraction pipeline. A large language model parsed open-access primary research articles to generate structured, provenance-anchored records of engineered genes, modification types, promoter-gene-terminator constructs, host species, target traits, and reported phenotypes, with every record traceable to its source. The current release contains 14,358 curated records encompassing 6,998 distinct genes across 436 plant species from 6,452 papers published between 2000 and 2026. Corpus analysis reveals that experiments are concentrated in a small group of model and crop species, disease and pathogen resistance is the most frequently engineered trait class, and constitutive regulatory parts (particularly the CaMV 35S promoter and NOS terminator) remain pervasive. Two in five records omit one or both flanking regulatory elements (i.e., promoter and terminator), while only 23.4% describe cassettes in which both elements resolve to named part classes, exposing a systematic reproducibility gap. We organize these data into a knowledge graph linking genes, constructs, species, and traits; provide access through an interactive web portal; and propose an AI-compatible documentation standard for AI-ready reporting. The Plant Bioengineering Atlas provides a foundation for data-driven hypothesis generation and AI-aided plant biodesign.

, Genes, DNA Constructs

Powered By SAM [Slides]

The System Advisor Model(TM) (SAM) is a free, open-source desktop application for techno-economic analysis of energy technologies. By combining detailed performance modeling with financial analysis, SAM allows users to assess technology trade-offs, explore future scenarios, and make informed decisions about energy investments. Users also have access to model details and the ability to embed SAM's core models in their own applications. This webinar, hosted by National Laboratory of the Rockies researchers Janine Keith and Matt Prilliman, highlights how this widely used modeling tool supports data-driven decision-making for energy systems.

29 ENERGY PLANNING, POLICY, AND ECONOMY

From Vehicles to Systems: Understanding Freight Transportation as a Connected Energy, Infrastructure, and Operations System

The U.S. freight system may need to handle 50% more cargo by 2050. Upgrading our freight system requires modernizing capital-intensive, long-lived assets including freight trains, ports, and terminal infrastructure. NLR is advancing freight system solutions spanning ALTRIOS, the first digital twin for the full freight rail system; ALTRIOS-LIFTS, which can create digital twins of freight terminals; INFORMES, the first national model of the intermodal freight system; MARINESim, used to simulate and optimize ocean-going vessel operations; and more. These modeling and simulation tools enable data-driven decision-making across freight modes and systems.

24 POWER TRANSMISSION AND DISTRIBUTION

Benchmarking Bayesian Optimization Frameworks and Acquisition Strategies for Materials Discovery and Autonomous Laboratories

Bayesian optimization (BO) can accelerate materials discovery by guiding expensive experiments toward the most promising processing conditions. We systematically compare five BO surrogate and framework combinations (Gaussian processes in Ax, Gaussian processes and Monte-Carlo neural networks in BayBE, random forests in Lolopy, and tree-structured Parzen (TPE) estimators in Hyperopt) on three benchmarks that mimic common materials design tasks (a discrete solid-electrolyte composition space, a hybrid discrete/continuous laminate-composite design problem solved with micromechanics modeling, and the continuous Ishigami analytic function which is a standard optimization benchmark). Each BO surrogate is paired with posterior mean, probability of improvement, and expected improvement acquisition functions and run for 100 trials from randomized initial samples with uniform random search providing a control. Across five random seeds per setting, BayBE’s Gaussian-process surrogate with expected improvement consistently reached ≥95 % of the known optimum in the fewest evaluations, while Lolopy’s random forest matched or exceeded GP performance on purely categorical or mixed spaces at a higher computational cost. Posterior mean alone often stagnated at local optima, underscoring the need for exploration, whereas probability and expected improvement balanced exploration and exploitation leading to better optimization in fewer trials. Execution times ranged from milliseconds for TPE to minutes for neural-network and random-forest surrogates. These results establish baseline expectations for BO in automated materials laboratories and highlight expected improvement with Gaussian processes as a reliable first choice, with random forests offering a strong alternative when categorical variables dominate. The benchmark suite and code are released to facilitate future surrogate, acquisition, and constraint-handling research in data-driven materials optimization.

Bayesian optimization

BEAST: Expanding Sustainable Data Infrastructure for High-Enthalpy Facilities

Reproducible, data-driven thermal protection system (TPS) research requires that experimental records from high-enthalpy testing be consistently structured, traceable, and accessible across campaigns and institutions. In practice, however, arcjet and plasma facilities data remain largely fragmented: raw diagnostics are stored in ad hoc formats, material sample histories are disconnected from test conditions, and metadata standards are absent, precluding systematic cross-campaign analysis and long-term reuse. BEAST (Backend for Experiment Analysis, Storage, and Traceability) is an open-source, web-based platform that addresses these limitations by providing a unified, queryable infrastructure for high-enthalpy ground-test data [1]. First presented at the 15th Ablation Workshop [2], BEAST has since undergone significant development. The platform ingests and structures multi-channel time-series diagnostics, facility configurations, and material property records within a common provenance model, ensuring end-to-end traceability from raw sensor acquisition to reduced experimental quantities. A versioned material library links specimen identity and processing history to the specific runs in which each sample was tested. An integrated modeling workbench enables training and evaluation of regression models directly on archived experimental data, supporting condition interpolation and the construction of empirical material response databases. Beyond its original deployment at NASA Ames Research Center, BEAST has been designed to be facility-agnostic, with ongoing efforts to extend its adoption to other facilities. Its modular architecture accommodates heterogeneous diagnostic setups and facility types, and its future open-source distribution allows institutions to build on a common data standard rather than maintaining isolated, bespoke solutions. BEAST is further integrated within a broader ecosystem of companion tools: arcjetCV [3] extracts recession rates and shock standoff distances from high-speed video using computer vision, and miniSTARscan [4] provides sub-minute, portable photogrammetric surface reconstruction of test articles before and after exposure. All tools share a common data schema, enabling seamless ingestion of surface geometry, imagery, and time-series data into a single, coherent experimental record.

Database

BEAST: Expanding Sustainable Data Infrastructure for High-Enthalpy Facilities

Reproducible, data-driven thermal protection system (TPS) research requires that experimental records from high-enthalpy testing be consistently structured, traceable, and accessible across campaigns and institutions. In practice, however, arcjet and plasma facilities data remain largely fragmented: raw diagnostics are stored in ad hoc formats, material sample histories are disconnected from test conditions, and metadata standards are absent, precluding systematic cross-campaign analysis and long-term reuse. BEAST (Backend for Experiment Analysis, Storage, and Traceability) is an open-source, web-based platform that addresses these limitations by providing a unified, queryable infrastructure for high-enthalpy ground-test data [1]. First presented at the 15th Ablation Workshop [2], BEAST has since undergone significant development. The platform ingests and structures multi-channel time-series diagnostics, facility configurations, and material property records within a common provenance model, ensuring end-to-end traceability from raw sensor acquisition to reduced experimental quantities. A versioned material library links specimen identity and processing history to the specific runs in which each sample was tested. An integrated modeling workbench enables training and evaluation of regression models directly on archived experimental data, supporting condition interpolation and the construction of empirical material response databases. Beyond its original deployment at NASA Ames Research Center, BEAST has been designed to be facility-agnostic, with ongoing efforts to extend its adoption to other facilities. Its modular architecture accommodates heterogeneous diagnostic setups and facility types, and its future open-source distribution allows institutions to build on a common data standard rather than maintaining isolated, bespoke solutions. BEAST is further integrated within a broader ecosystem of companion tools: arcjetCV [3] extracts recession rates and shock standoff distances from high-speed video using computer vision, and miniSTARscan [4] provides sub-minute, portable photogrammetric surface reconstruction of test articles before and after exposure. All tools share a common data schema, enabling seamless ingestion of surface geometry, imagery, and time-series data into a single, coherent experimental record.

Database

State, Local, and Tribal Program

NLR's State, Local, and Tribal Program delivers customized, data-driven support that strengthens local energy systems - expanding access to America's abundant energy resources, reducing costs, and supporting energy reliability across the country. NLR's world-class staff use a wide variety of cutting-edge energy tools and capabilities to deliver robust modeling, validation, and deployment support to hundreds of communities annually.

29 ENERGY PLANNING, POLICY, AND ECONOMY

Feature-Based PMU Event Classification under Variable PMU Participation and Overlapping Events

This paper is the basis for a presentation help at the 2026 Georgia Tech Fault & Disturbance Analysis Conference, which can be found at OSTI # 3168287 Paper Abstract—Phasor Measurement Units (PMUs) stream time synchronized, high-resolution measurements from the grid, enabling data-driven techniques for event detection and classification. Accurate event classification improves grid reliability and stability. Events can be detected by varying numbers of PMUs and exhibit different durations depending on the event type. This variability challenges standard classifiers that require uniform input sizes. Moreover, multiple events may coincide, which increases classification complexity. Standard classifiers assign each instance to the class with the highest predicted probability, whereas overlapping events may exhibit comparable probabilities across multiple classes. In this study, to handle data size variability, we extract a wide range of time–frequency domain features from all available PMUs for each event into a fixed-length vector, facilitating the application of standard machine learning classifiers, including Random Forest, XGBoost, LightGBM, Support Vector Machine, and Multilayer Perceptron. To account for overlapping events, a probabilistic post-processing step is applied. For a given data instance, if multiple predicted class probabilities exceed 30% and the differences between them are less than 10%, the event is assigned to multiple classes. Experiments using real-world PMU data demonstrate that the Random Forest and XGBoost models achieve the highest accuracy, while the proposed post-processing method yields perfect classification performance on external unseen test sets.

Nematirad, Reza [Danovo Energy Solutions]

Flavor as an Incomplete Structure: Conceptual Questions and the Role of DUNE

Flavor remains one of the most successful yet least understood structures of the Standard Model. The discovery of the Higgs boson completed the electroweak account of mass generation, but did not explain the origin of fermion families, mass hierarchies, or mixing patterns. In this sense, flavor can be regarded as an empirically successful but conceptually incomplete structure. Neutrinos occupy a particularly sensitive place within this problem: their masses are tiny, their mixing is large, and their mass-generation mechanism may differ from that of charged fermions. In this article, we discuss flavor as an open conceptual problem and argue that DUNE, as a phased program spanning precision oscillation measurements and sensitivity to BSM and dark-sector phenomena, provides a powerful framework for testing the self-consistency and possible limits of the present three-flavor description. In particular, the complementarity between the long-baseline program and the Phase I near-detector complex, together with the DUNE-PRISM strategy for controlling interaction-model systematics and enabling data-driven near-to-far predictions, makes DUNE especially well-suited to search for small, correlated departures from the minimal flavor framework.

Montanari, Claudio S. [Fermilab; INFN, Pavia] (ORC

Standard Candles for Supernova Neutrino Detection at DUNE

The Deep Underground Neutrino Experiment (DUNE) far detector is sensitive to $\mathcal{O}(10)$ MeV electron neutrinos through $ν_e$ charged-current reaction with argon. This capability is a unique window into the $ν_e$ component of a Galactic core-collapse supernova flux. Extracting the properties of the supernova spectrum is, however, limited by the poorly-known $ν_e$-Ar cross section. We propose a data-driven strategy that leverages $^8$B solar neutrinos and muon-decay-at-rest neutrinos as standard candles for this process. These calibration samples constrain both the low-energy and high-energy components of the cross section relevant for supernova detection. Our method reduces the reliance on nuclear models, which can bias the extraction of the spectral parameters by as much as 300$\%$.

Cheng, Ting [Fermilab]

Evaluation of a Reduced-Order Model for IBR Fault Response Representation via OEM Blackbox Models: Preprint

Driven by the need to capture the electromagnetic transients of transmission lines, inverter switching behavior, and detailed control systems, electromagnetic transient (EMT) studies have become increasingly important in industry, such as IBR interconnection study and fault study. However, original equipment manufacturer (OEM) inverter models typically include extensive parameters and proprietary settings that are unavailable to protection engineers. This paper introduces a data-driven, reduced-order model (ROM) developed as a PSCAD library component for use in EMT-based fault studies. The ROM replicates key OEM model behaviors without requiring detailed knowledge of control design or parameterization. The accompanying Python automation scripts streamline data generation, parameter fitting, and validation. The ROM's performance is demonstrated through comparison with both IEEE 2800-compliant and non-compliant OEM models in a real-world power system. Relay responses show nearly identical results, while simulation runtime is reduced by an average of 32.8\%, highlighting the ROM's practicality for protection engineers.

14 SOLAR ENERGY

Virtual Growth of SRF Materials: A Machine Learning Approach to Predict the Crystalline Structural Ordering in Nb Surface Oxides

Niobium's native surface oxide affects SRF cavity and superconducting qubit performance, motivating interest in controlling its crystalline structure. We combine a literature-derived machine-learning analysis with temperature-dependent XRD to study crystalline ordering in Nb2O5. Random Forest models, trained on 74 processing conditions from 17 papers and validated by leave-one-group-out cross-validation, predicted broad crystallinity outcomes well (balanced accuracy 0.809), but struggled with specific polymorph identity (0.577). Annealing temperature was the dominant predictor across all targets; oxygen partial pressure showed negligible importance, reflecting narrow literature coverage rather than physical irrelevance. Temperature-dependent XRD on anodized and H2O2-treated Niobium showed structural evolution consistent with the machine learning predictions. Our model and overall approach provide a data-driven framework for identifying and optimizing conditions that promote crystallization in initially amorphous oxides. This framework can guide the selection of growth and post-annealing conditions for Nb surfaces by narrowing the experimental parameter space, thereby reducing trial-and-error efforts in developing oxide structures relevant to SRF applications.

Tilkin, Anthony [Unlisted, US, IL; Fermilab]

Determining the Efficiency of EMPHATICs Silicon Strip Detectors (SSDs)

EMPHATIC is an experiment at Fermilab which aims to reduce current neutrino flux uncertainties. This report discusses the limitations current neutrino flux uncertainties places on large scale neutrino experiments, provides background on the EMPHATIC experiment, and details the project of determining the efficiency of the Silicon Strip Detectors (SSDs) used in EMPHATIC. As part of the data analysis process and in order to increase the accuracy of EMPHATIC’s simulations a representation of efficiency of each SSD is required. To achieve this a data-driven analysis was performed on EMPHATIC's collected data using the Root and Art frameworks. Visual and numerical representations of efficiency were determined. The average efficiency over all SSDs is 98.58\%, however this number deflated as it includes known bad channels.

Olson, V. [Illinois U., Urbana (main)]

Neutron detector response modeling in NOvA

Neutrons can present a significant challenge for neutrino experiments in which energy reconstruction is critical. With the ability to escape detection completely and with a weak correlation between their kinetic energy and any eventual energy deposition, it is difficult to fully account for neutrons produced in neutrino interactions. This in turn leads to significant model dependence when evaluating neutron-related systematic uncertainties. The NOvA experiment is a long-baseline neutrino oscillation experiment with a high-statistics sample of antineutrino data collected by its near detector. We report an excess relative to data of simulated neutron candidates with low energy depositions when using standard Geant4 physics lists. The simulation excess is traced to an overabundance of secondary photons produced from interactions of neutrons with kinetic energy greater than \SI{20}{\mega\eV}. Improved agreement with data is obtained by applying the data-driven neutron-on-carbon \menate model for neutrons between \SI{20}{\mega\eV} and ${\sim}$\SI{100}{\mega\eV}. With \menate, the residual oversimulation is more uniform across the calorimetric neutron energy spectrum, suggesting possible overproduction of primary neutrons by the GENIE neutrino interaction generator. These results motivate the adoption of \menate-supplemented Geant4 simulation as the nominal simulation in the production of future \nova simulation.

Abubakar, S.