Search NASASearch

SEARCH · Search NASA

Results for “Data driven”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Data driven methods to recognize patterns in EIC weak-strong simulation

Beam-Beam simulations are currently being studied in preparation for future EIC experiments to study beam-beam effects and, in turn, maximize luminosity. Weak-strong methods are studied for single-particle dynamics during collision. 1 million macro-particles for 1 million turns are typically tracked, corresponding to only 10 seconds in the EIC. The goal of this study is to predict beam properties over the scale of hours. A potential solution focuses on using data-driven methods such as machine learning methods to analyze and extend the insights of the beam properties such as long-term nonlinear effects. This would aid in long-term predictions where results would be more efficiently acquired than a typical tracking simulation. Some limitations such as inaccurate predictions and spatial complexity are also discussed. These methods can then be applied to strong-strong simulations in the future studies.

Accelerator Physics

A Data-Driven Approach for High-Impedance Fault Localization in Distribution Systems: Preprint

Accurate and quick identification of high-impedance faults (HIFs) is critical for the reliable operation of distribution systems. Unlike other faults in power grids, HIFs are very difficult to detect by conventional overcurrent relays due to the low fault current. Although HIFs can be affected by various factors, the voltage-current characteristics can substantially imply how the system responds to the disturbance and thus provides opportunities to effectively localize HIFs. In this work, we propose a data-driven approach for the identification of HIF events. To tackle the nonlinearity of the voltage-current trajectory, first, we formulate optimization problems to approximate the trajectory with piecewise functions. Then we collect the function features of all segments as inputs and use the support vector machine approach to efficiently identify HIFs at different locations. Numerical studies on the IEEE 123-node test feeder demonstrate the validity and accuracy of the proposed approach for real-time HIF identification.

explainable artificial intelligence

Data-driven method to estimate contamination from light ion beam transmutation at colliders

Collisions of relativistic light ions, such as oxygen, neon, and magnesium, have been proposed as a way to examine the system-size dependence of dynamics typically associated with the quark-gluon plasma produced in collisions of heavier ions such as xenon, gold, or lead. Recent efforts at both the Relativistic Heavy Ion Collider (RHIC) and Large Hadron Collider (LHC) have produced large datasets of proton-oxygen, oxygen-oxygen, and neon-neon collisions, catalyzing intense interest in experimental backgrounds associated with light-ion collisions. In particular, electromagnetic dissociation of light ions while they are circulating in a collider can result in beam contamination that is difficult to simulate precisely. Here we propose a data-driven method for evaluating the potential impact of beam contaminants on physics analyses. The method exploits the time dependence and smaller size of contaminant ion species to define control regions that can be used to quantify potential contamination effects. A simple model is used to illustrate the method and to study its robustness. Furthermore, this method can inform studies of recent LHC and RHIC data and could also be useful for future light-ion programs at the LHC and beyond.

Beam loss

Techno-Economic Analysis of Data-Driven and Transactive Approaches for Resilience Enhancement

As extreme weather events lead to more frequent power outages, understanding and enhancing grid resilience is critical to mitigating economic losses and non-energy impacts from service disruptions. Here, this study introduces a novel techno-economic analysis framework for evaluating resilience enhancement mechanisms. The framework combines grid response modeling with a co-simulation approach and valuation methodology to provide a comprehensive assessment. We apply this framework to a realistic case study of the Texas grid during Winter Storm Uri in February 2021. Two advanced resilience strategies are analyzed: a data-driven rolling outage mechanism and a transactive energy (TE) based allocation scheme. The rolling outage scheme selectively serves customers based on real-time curtailment needs, while the TE scheme allows customers to trade energy allocations according to their preferences. Our findings show that both the rolling outage and TE schemes significantly outperform conventional methods (i.e. controlled outages) by reducing the amount of energy not supplied to customers by 41% and 64%, respectively. These approaches also enhance flexibility and customer satisfaction, while improving energy utilization for greater resilience. Additionally, they maintain thermal comfort about 3.5 times better and substantially lower customer risk exposure. A key contribution of this study is addressing both utility and customer perspectives while considering both energy and non-energy impacts. The techno-economic analysis indicates that implementing these resilience enhancement strategies would incur an additional 1.1Bto1.6B in utility costs but has the potential to avoid 17.3Bto18B of customer losses as compared to existing solutions, thereby underscoring the value of investing in advanced resilience, as it provides significant societal benefits to customers.

42 ENGINEERING

Correlating processing variables to material properties in recycled polypropylene: A data‐driven approach

Abstract Polypropylene (PP) is one of the most widely used plastics, yet its recycling remains limited, with less than 1% of solid waste PP being reprocessed. Mechanical recycling through extrusion is the most practical method, but inconsistent reprocessing conditions introduce variability in material properties. While temperature, screw speed, and residence time influence the thermomechanical stress applied during reprocessing, there are no standardized guidelines for optimizing these parameters. This study examines how these factors shape the properties of recycled PP, using conditions designed to mimic post‐industrial recycled (PIR) scrap. Residence time was measured using colorimetric tracking and correlated with molecular weight, viscosity, and mechanical properties over multiple extrusion cycles. Data‐driven modeling, including response surface methodology, support vector machines, and artificial neural networks, identified processing temperature as the dominant factor in material degradation, followed by residence time. Mechanical properties remained stable, while viscosity decreased predictably with increasing residence time. By linking reprocessing conditions to property evolution, this study provides a method to optimize processing parameters and reduce variability in recycled PP. These findings help manufacturers improve process control, making recycled PP more predictable for reuse in manufacturing. Highlights Study of PIR‐quality PP without additives or compatibilizers. Residence time analysis shows processing temperature drives PP property changes. Mark‐Houwink enables quick molecular weight checks for quality control. Models predict mechanical and rheological shifts in reprocessing. Optimized processing parameters minimize property degradation in recycling.

Estela‐García, John E. [Polymer Engineering Center

Accurate data-driven surrogates of dynamical systems for forward propagation of uncertainty

Stochastic collocation (SC) is a well-known non-intrusive method of constructing surrogate models for uncertainty quantification. In dynamical systems, SC is especially suited for full-field uncertainty propagation that characterizes the distributions of the high-dimensional solution fields of a model with stochastic input parameters. However, due to the highly nonlinear nature of the parameter-to-solution map in even the simplest dynamical systems, the constructed SC surrogates are often inaccurate. Here, this work presents an alternative approach, where we apply the SC approximation over the dynamics of the model, rather than the solution. By combining the data-driven sparse identification of nonlinear dynamics framework with SC, we construct dynamics surrogates and integrate them through time to construct the surrogate solutions. We demonstrate that the SC-over-dynamics framework leads to smaller errors, both in terms of the approximated system trajectories as well as the model state distributions, when compared against full-field SC applied to the solutions directly. We present numerical evidence of this improvement using three test problems: a chaotic ordinary differential equation, and two partial differential equations from solid mechanics.

42 ENGINEERING

Meeting Global Health Needs via Infectious Disease Forecasting: Development of a Reliable Data-Driven Framework

Infectious diseases (IDs) have a significant detrimental impact on global health. Timely and accurate ID forecasting can result in more informed implementation of control measures and prevention policies. To meet the operational decision-making needs of real-world circumstances, we aimed to build a standardized, reliable, and trustworthy ID forecasting pipeline and visualization dashboard that is generalizable across a wide range of modeling techniques, IDs, and global locations. We forecasted 6 diverse, zoonotic diseases (brucellosis, campylobacteriosis, Middle East respiratory syndrome, Q fever, tick-borne encephalitis, and tularemia) across 4 continents and 8 countries. We included a wide range of statistical, machine learning, and deep learning models (n=9) and trained them on a multitude of features (average n=2326) within the One Health landscape, including demography, landscape, climate, and socioeconomic factors. The pipeline and dashboard were created in consideration of crucial operational metrics—prediction accuracy, computational efficiency, spatiotemporal generalizability, uncertainty quantification, and interpretability—which are essential to strategic data-driven decisions. While no single best model was suitable for all disease, region, and country combinations, our ensemble technique selects the best-performing model for each given scenario to achieve the closest prediction. For new or emerging diseases in a region, the ensemble model can predict how the disease may behave in the new region using a pretrained model from a similar region with a history of that disease. The data visualization dashboard provides a clean interface of important analytical metrics, such as ID temporal patterns, forecasts, prediction uncertainties, and model feature importance across all geographic locations and disease combinations. As the need for real-time, operational ID forecasting capabilities increases, this standardized and automated platform for data collection, analysis, and reporting is a major step forward in enabling evidence-based public health decisions and policies for the prevention and mitigation of future ID outbreaks.

60 APPLIED LIFE SCIENCES

High-throughput and data-driven search for stable optoelectronic AMSe 3 materials

The rapid advancement in emerging optoelectronic technologies demands highly efficient, affordable, and ecofriendly materials. In this context, ternary chalcogenides, especially ternary selenides, show early promise as a material class due to their stability and remarkable electronic, optical, and transport properties. In this work, we integrate first-principles-based high-throughput computations with machine learning (ML) techniques to predict the thermodynamic stability and optoelectronic properties of 920 valency-satisfied selenide compounds. Through investigating polymorphism, our study reveals the edge-sharing orthorhombic Pnma phase (NH 4 CdCl 3 -type) as the most stable structure for most ternary selenides. High-fidelity supervised ML models are trained and tested to accelerate stability and band gap predictions. These data-driven models pin down the most influential features that dominantly control key material characteristics. The multistep high-throughput computations identify the ternary selenides with optimal direct band gaps, light carrier masses, and strong optical absorption edges. The extensive materials screening considering phase stability, toxicity, and defect tolerance, finally identifies the seven most suitable candidates for photovoltaic applications. Two of these final compounds, SrZrSe 3 and SrHfSe 3 , have already been synthesized in a single-phase form, with the latter showing an optically suitable band gap, aligning well with our findings. The non-adiabatic molecular dynamics reveal sufficiently long photoexcited charge carrier lifetimes (on the order of nanoseconds) in some of these selected selenide materials, indicating their exciting characteristics. Overall, our study suggests a robust in silico framework that can be extended to screen large datasets of various material classes for identifying promising photoactive candidates.

36 MATERIALS SCIENCE

Data-Driven Modeling of High-Resolution Residential Load Profiles Using Low-Resolution Smart Meter Measurements

Accurate and high-resolution residential load profiles are essential for power system modeling, demand response planning, and effective grid operation. As the energy sector moves towards a more actively managed distribution system, the ability to understand residential energy consumption at a minute-by-minute scale becomes increasingly critical. High-resolution load profiles provide key insights into demand patterns and user behavior, enabling grid operators to design more effective energy solutions; however, residential load measurements in the field are typically recorded at low resolutions, such as 15-60 minutes, which makes it hard to study the characteristics of different residential customers. This paper addresses these challenges by introducing a data-driven approach to generate realistic, high-resolution residential load profiles based on lowre-solution measurements and weather information. The proposed method retains the key features of the actual residential load measurements while offering appliance-level energy consumption details for each residential building. The results demonstrate the effectiveness of the proposed load profile generator, proving its capability to support utilities in optimizing residential energy management and ensuring a more reliable and resilient grid.

24 POWER TRANSMISSION AND DISTRIBUTION

A data-driven method to estimate contamination from light ion beam transmutation at colliders

Collisions of relativistic light ions such as oxygen, neon, and magnesium, have been proposed as a way to examine the system-size dependence of dynamics typically associated with the quark-gluon plasma produced in collisions of heavier ions such as xenon, gold, or lead. Recent efforts at both the Relativistic Heavy Ion Collider (RHIC) and Large Hadron Collider (LHC) have produced large datasets of proton-oxygen, oxygen-oxygen, and neon-neon collisions, catalyzing intense interest in experimental backgrounds associated with light ion collisions. In particular, electromagnetic dissociation of light ions while they are circulating in a collider can result in beam contamination that is difficult to simulate precisely. Here we propose a data-driven method for evaluating the potential impact of beam contaminants on physics analyses. The method exploits the time-dependence and smaller size of contaminant ion species to define control regions that can be used to quantify potential contamination effects. A simple mode

Accelerator Physics (physics.acc-ph)

Data-driven Techniques for $v^{e}$ Signal and Background Predictions in NOνA

NOν \nu A is a long baseline neutrino oscillation experiment, using two functionally identical detectors to measure νe \nu_e appearance and νμ \nu_\mu disappearance at the Far Detector (FD) with the NuMI Beam at Fermilab. The Near Detector (ND) measures the beam before oscillations, which will allow us to make a measurement of the signal flux before it has oscillated and the background components which can mimic νe \nu_e at our FD. We use the ND in three different ways to predict the amount of signal and background expected at the FD to partially cancel systematic uncertainties. The background prediction has three components: Charged Current (CC) νμ \nu_\mu , νe \nu_e and Neutral Current (NC) events. We need to determine the fraction of the selected ND sample in each of these components because some oscillate significantly (CC νμ \nu_\mu ) and some do not oscillate at all (NC). This poster will present details of these three data-driven techniques for predicting the FD spectrum.

Yu, Shiqi [Argonne; IIT, Chicago]

Data-driven design of electrolyte additives supporting high-performance 5 V LiNi 0.5 Mn 1.5 O 4 positive electrodes

LiNi 0.5 Mn 1.5 O 4 (LNMO) is a high-capacity spinel-structured material with an average lithiation/de-lithiation potential at ca. 4.6–4.7 V vs Li + /Li, far exceeding the stability limits of electrolytes. An efficient way to enable LNMO in lithium-ion batteries is to reformulate an electrolyte composition that stabilizes both graphitic (Gr) negative electrode with solid-electrolyte-interphase and LNMO with cathode-electrolyte-interphase. In this study, we select and test a diverse collection of 28 single and dual additives for the Gr||LNMO battery system. Subsequently, we train machine learning models on this dataset and employ the trained models to suggest 6 binary compositions out of 125, based on predicted final area-specific-impedance, impedance rise, and final specific-capacity. Such machine learning-generated new additives outperform the initial dataset. This finding not only underscores the efficacy of machine learning in identifying materials in a highly complicated application space but also showcases an accelerated material discovery workflow that directly integrates data-driven methods with battery testing experiments.

batteries

Data-Driven Recommendation of Optimal Tuning Scheme for Range-Separated Hybrid Functionals in Solution-Phase UV/Vis Absorption Energy Prediction

Time-dependent density functional theory (TDDFT) combined with range-separated hybrid (RSH) functionals and a tuned range-separation parameter γ offers a computationally economical approach for high-throughput excited- state property predictions. The γ-tuning procedure in the gas phase is well established. However, no agreement on the best γ- tuning procedure has been made when considering the solvent effect with implicit solvent models like the polarizable continuum model (PCM). To answer that question, this study created a diverse dataset with 937 molecules with experimental solutionphase UV/vis absorption spectra. Three γ-tuning methods, the gasphase γ-tuning (GPγT), the partial vertical γ-tuning (PVγT), and the strict vertical γ-tuning (SVγT), were evaluated for the ωPBEh functional over the entire dataset. Additional benchmarks are done for the optimally tuned screened range-separated hybrid combined with the PCM approach (SRSH-PCM) and the solvation-mediated tuning procedure (sol-med-OT). Our findings revealed that the optimal γ-values obtained by the PVγT and the SVγT are significantly smaller than the GPγT. This trend holds consistently across all molecules in our dataset, and we explained the origin of this phenomenon. TDDFT calculations with PVγTand SVγT-tuned γ-values and default global Fock exchange fraction achieve superior performance compared to those using GPγTtuned or default γ and slightly outperform SRSH-PCM and sol-med-OT with similar or lesser computational cost. Furthermore, we found that the smaller γ-values from SVγT captured the expected 1/(εR) asymptotic behavior in the solution phase, resulting in accurate prediction of solution-phase CT excitations, consistent with the screened asymptote behavior encoded in SRSH-PCM. These results show that SVγT is the best scheme for high-throughput UV/vis absorption spectrum calculations using the ωPBEh functional from a data-driven perspective.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Data-Driven State of Health Estimation for Second-Life Batteries Using Interpolated Synthetic Data and Feature Selection

Accurate estimation of the State of Health (SOH) for second-life batteries (SLBs) is crucial given their increasing use in energy storage applications. Precise SOH prediction is essential for safe operation and robust battery management systems. A major challenge is the limited availability of datasets for building reliable degradation models. To address this, synthetic data generation through linear interpolation is performed to extend the available data, making it more representative of real-world battery operating conditions. By analyzing feature correlation with SOH, the most relevant features are selected for the model. The proposed approach employs a convolutional neural network (CNN) model trained on this interpolated, feature-selected dataset, using time series data of voltage, temperature, and current over a cycle. By focusing on highly correlated features, the model achieves over 95% accuracy, with mean absolute error and root mean squared error up to 2.27% and 2.64%, respectively, in SOH estimation for two battery datasets tested. These results highlight the potential of combining synthetic data generation and feature selection to enhance SOH predictions, showcasing the superior performance of the proposed CNN model for both new batteries and SLBs.

feature selection

Data-Driven Approach for Controlled Icosahedral Boron- Rich Compound Growth

This final technical report summarizes the research accomplishments and research highlights at the end of the funding period. This project aimed to leverage existing and new computational data produced from first-principles and molecular dynamics simulations to understand the thermodynamic, mechanical, and electronic properties of icosahedral boron compounds. The goal is to achieve targeted material properties by controlling the synthesis routes of these boron-rich compounds.

36 MATERIALS SCIENCE

Data-driven projection pursuit adaptation of polynomial chaos expansions for dependent high-dimensional parameters

Uncertainty quantification (UQ) and inference involving a large number of parameters are valuable tools for problems associated with heterogeneous and non-stationary behaviors. The difficulty with these problems is exacerbated when these parameters are statistically dependent requiring statistical characterization over joint measures. Probabilistic modeling methodologies stand as effective tools in the realms of UQ and inference. Among these, polynomial chaos expansions (PCE), when adapted to low-dimensional quantities of interest (QoI), provide effective yet accurate approximations for these QoI in terms of an adapted orthogonal basis. These adaptation techniques have been cast as projection pursuits in Gaussian Hilbert space in what has been referred to as a projection pursuit adaptation (PPA) by Xiaoshu Zeng and Roger Ghanem (2023). The PPA method efficiently identifies an optimal low-dimensional space for representing the QoI and simultaneously evaluates an optimal PCE within that space. The quality of this approximation clearly depends on the size of the training dataset, which is typically a function of the adapted reduced dimension. Here, the complexity of the problem is thus mediated by the complexity of the low-dimensional quantity of interest and not the complexity of the high-dimensional parameter space.

Data-driven

MDLoader: A Hybrid Model-Driven Data Loader for Distributed Graph Neural Network Training

Scalable data management is essential for processing large scientific dataset on HPC platforms for distributed deep learning. In-memory distributed storage is preferred for its speed, enabling rapid, random, and frequent data access required by stochastic optimizers. Processes use one-sided or collective communication to fetch remote data, with optimal performance depending on (i) dataset characteristics, (ii) training scale, and (iii) interconnection network. Empirical analysis shows collective communication excels with larger mini-batch sizes and/or fewer processes, whereas one-sided communication outperforms at larger scales. We propose MDLoader, a hybrid in-memory data loader for distributed graph neural network training. MDLoader features a model-driven performance estimator that dynamically selects between one-sided and collective communication at the beginning of training using Tree of Parzen Estimators (TPE). Evaluations on NERSC Perlmutter and OLCF Summit show MDLoader outperforms single-backend loaders by up to 2.83 × and predicts the suitable communication method with 96.3% (Perlmutter) and 94.3% (Summit) success rate.

Bae, Jonghyun