Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data driven”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

A Data-Driven Method for Synthetic Extreme Weather Generation and Solar Impact Assessment: Preprint

High-resolution, high-fidelity weather datasets are essential for testing and evaluating the resilience of power systems, particularly under extreme weather conditions. However, existing extreme weather datasets are typically derived from historical events that are localized and may lack the spatial and temporal resolution or scenario diversity needed to test largescale power systems. In this work, we propose a synthetic extreme weather simulation approach capable of generating targeted extreme events, such as hurricanes, using publicly available data sources. Preliminary results demonstrate the impact of a simulated Category 1 hurricane on renewable generation and critical infrastructure in California. The work aims to provide a flexible approach for creating multiple types of extreme weather scenarios across different regions, enabling comprehensive system stress testing, training, and resilience assessment.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Robust Data-Driven Predictive Run-to-Run Control for Automated Serial Sectioning

This letter presents a one-step predictive run-to-run controller (R2R-MPC) for the automation of mechanical serial sectioning (MSS), a destructive material analysis process. To address the inherent uncertainty and disturbances in the MSS process, a robust closed-loop approach is presented. Here, the robust R2R-MPC models the uncertainty of the MSS process using a linear differential inclusion. As an analytical model of the MSS process is unavailable, the differential inclusion is identified from historical data. The R2R-MPC is posed as an optimization problem that computes incremental changes to the control input which minimize the worst-case material removal errors. This optimization-based controller is combined with a run-to-run controller to provide integral action that rejects constant disturbances and tracks constant reference removal rates. To demonstrate the efficacy of our robust R2R-MPC, we present simulation results which compare the presented controller with a conventional non-robust R2R.

42 ENGINEERING↗

Improving Trustworthiness of Data-Driven Power Grid Contingency Analysis With Bayesian Residual Graph Neural Networks

The evolving energy landscape requires novel tools to efficiently perform contingency analysis and reliability assessment of power grids, potentially in real-time. The high computational cost of traditional power flow solvers limits their applicability in practice. Machine learning (ML) surrogates such as deep neural networks (NNs) accelerate power flow solvers computations, enabling high-order contingency analysis and real-time decision-making by learning highly nonlinear functions and integrating grid topology via graph architectures. However, (graph) NNs lack predictive power away from training data and do not provide predictive confidence estimates. Here, we present a Bayesian residual graph NN that integrates knowledge from low-fidelity data via residual training and embeds granular quantification of uncertainties, improving trustworthiness critical for high-consequence decision-making. Applying Bayesian concepts to NNs is challenging due to the high-dimensionality of both the parameter space, complicating derivation of a meaningful prior, and the output space in large grid systems, requiring enhanced techniques to assess the predicted high-dimensional uncertainties. Our contributions include: (1) Deriving a prior for fully connected and graph NNs that leverages low-fidelity data to guide mean predictions and appropriately control prior predictive uncertainty. (2) Integrating this prior within an ensembling with anchoring scheme for efficient approximate posterior inference. (3) Deriving enhanced metrics to assess accuracy of both the mean and uncertainty predictions in high dimensions, appropriately accounting for correlations propagated through graph layers. The resulting Bayesian residual graph NN is tested on a contingency analysis task for 14-bus and 118-bus grids.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

Mechanical separations of corn stover anatomical fractions in an integrated feedstock preprocessing system: An experimental and data-driven modeling study

High variabilities of material attributes in lignocellulosic biomass present risks for biofuel and biochemical productions and must be mitigated via preprocessing. Since almost no mechanical device is originally designed for processing biomass, how to operate existing apparatuses with efficient performance has not been investigated extensively. This work presents a study on an integrated screening and air classification to separate cobs and stalks from husks and leaves in corn stover. Prototype machine learning models were developed to assess the feasibility of predicting the process outcome based on the measurable parameters. The models trained upon limited experimental data rendered decent predictive accuracy of yield and purity. The experimental data and modeling results collectively suggest decreasing throughput leads to a higher purity. To the contrary, if throughput increases, a lower purity is likely. A possible trade-off between yield and purity of the separated streams indicates the need for optimal combinations of feedstock size, moisture, and throughput to achieve optimized separations. The results of this study also suggest the need to further improve model predictability by developing more accurate formulations for physics governing the integrated unit operations. To accomplish this, additional experimental data needs to be generated for model training.

09 - BIOMASS FUELS↗

A high-throughput experimentation platform for data-driven discovery in electrochemistry

Automating electrochemical analyses combined with artificial intelligence is poised to accelerate discoveries in renewable energy sciences and technologies. This study presents an automated high-throughput electrochemical characterization (AHTech) platform as a cost-effective and versatile tool for rapidly assessing liquid analytes. The Python-controlled platform combines a liquid handling robot, potentiostat, and customizable microelectrode bundles for diverse, reproducible electrochemical measurements in microtiter plates, minimizing chemical consumption and manual effort. To showcase the capability of AHTech, we screened a library of 180 small molecules as electrolyte additives for aqueous zinc metal batteries, generating data for training machine learning models to predict Coulombic efficiencies. Key molecular features governing additive performance were elucidated using Shapley Additive exPlanations and Spearman’s correlation, pinpointing high-performance candidates like cis-4-hydroxy-d-proline, which achieved an average Coulombic efficiency of 99.52% over 200 cycles. The workflow established herein is highly adaptable, offering a powerful framework for accelerating the exploration and optimization of extensive chemical spaces across diverse energy storage and conversion fields.

Lin, Dian-Zhao [Johns Hopkins University, Baltimor↗

Data Driven Commercial Building Energy Code Compliance and Technology Inventory for New York City

Building Performance Standards (BPS) are gaining national traction. A BPS will require new processes in the design, construction, and operation of buildings that take the occupants into account and enable predictive analysis to ensure compliance with current and future GHG emissions caps. In New York City, most buildings over 25,000 square feet will be regulated by a BPS starting in 2024, regardless of whether it is new construction permitted under current energy codes or an existing building. This research is one of the first to begin the evaluation of a long-term series of building policies in the context of an open data ecosystem, in cooperation with city agencies. Existing building policies enacted in NYC have ranged from building energy benchmarking and labeling to energy audits to the regulation of GHG emission in buildings. Through the development of a dataset related to building technologies and energy consumption, this project can help to evaluate if meaningful conclusions can be drawn for the data that has been largely self-reported in compliance with city regulations. This project will also provide lessons learned from a deep dive into these types of datasets to provide best practices for municipalities or states seeking to embark on policies like those enacted in NYC. In addition, a Building Automation System (BAS) Stretch Standard of Care (SSOC) for owners, designers, and building operators will enable the measurement and predictive analysis of energy consumption and GHG emissions at the plant, system, or component level, in anticipation of regulated GHG limits on buildings based on energy use. The SSOC is expected to be suitable for use on a national level. The primary feature of an SSOC is a standardized format for a set of BAS points that can be used to control and to gather data from individual plants, systems, or components that are related to building energy consumption. This project examined how measurements compare to prescriptive or simulation-based energy code targets, finding little correlation between predictive 8760-hour energy modeling and actual energy consumption for a small sample (n=27) of buildings constructed after 2015. Other analysis found that, while large multifamily housing (MFH) buildings showed a general trend similar to predicted reductions in energy use from the implementation of model commercial energy codes, this trend was not evident in the office, K-12 school, and hotel use groups in NYC. No upward or downward trends in energy consumption were found when buildings were grouped by size. Energy audit data were analyzed and it appears that there is bias by audit company on measures recommended to clients. Further research should be performed to cross-analyze this with other attributes, such as building size, vintage, and number of stories. Analysis found that for 281 buildings that were permitted and completed after 2015 and had submitted benchmarking data in 2022, between 81% and 96% (by use group) were found to be in compliance with the 2024 to 2029 NYC BPS emission caps, and between 55% and 89% were in compliance with the 2030-2034 caps. This work is beneficial to the public in helping policymakers and building stakeholders better understand the wide-ranging implications of a BPS.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Data-driven Modeling for Grid Edge IBRs: A Digital Twin Perspective of User-Defined Models

Recent events in Odessa have brought attention to the challenges associated with the interaction between Inverter- Based Resources (IBRs) and the transmission and distribution system. The NERC event diagnosis report has highlighted sev- eral issues, emphasizing the need for continuous performance monitoring of these IBRs by system operators. Key areas of concern include the mismatch of control and protection perfor- mance of IBRs between the original equipment manufacturer (OEM)-provided models and field measurements. The inability to replicate the realistic response can result in incorrect reliability and resilience studies. In this paper, we developed an approach on how to emulate the behavior of an IBR using measurement data obtained for system operators to utilize in real-time and long- term planning. Two experiments are conducted in the phasor domain and electromagnetic transients (EMT) domain to emulate the behavior for grid forming and grid following inverters under various operating conditions and the effectiveness of the proposed model is demonstrated in terms of accuracy and ease of utilizing user-defined models (UDMs)

Mahapatra, Kaveri [BATTELLE (PACIFIC NW LAB)]↗

Data-driven emulation of modal aerosol microphysics via neural operator-based modeling

The complexity and the small characteristic scales of aerosol microphysical processes pose a big challenge for accurate and efficient Earth system simulations at regional and global scales. In this work, we construct and evaluate a surrogate model: the aerosol deep operator network (ADON), a physics-inspired dual-net architecture for emulating the aerosol microphysics parameterization suite in the version 2 of the Energy Earth System Model (E3SMv2). The current version of the surrogate model is trained on a dataset comprising 9.8 million samples obtained from a global E3SMv2 simulation with the horizontal resolution of about one degree under cloud-free conditions. Incorporating domain spatial and temporal coordinates, as well as principle components extracted from training data, the dual-net surrogate model effectively captures the intricate representations of aerosol and the relationship with atmospheric state variables, achieving an R-squared score over $$95.7\%$$ for all the lognormal aerosol modes in the extrapolated regime. The validated model provides feature importance of input variables and their impact on the predictive capacity of the surrogate model in relation to the E3SM. The computational cost of online inference time deployed on CPUs and GPUs with lower precisions highlights ADON’s efficiency and potential in robust predictive modeling for large-scale Earth system computations.

Bai, Zhe↗

Data-driven organic solubility prediction at the limit of aleatoric uncertainty

Abstract Small molecule solubility is a critically important property which affects the efficiency, environmental impact, and phase behavior of synthetic processes. Experimental determination of solubility is a time- and resource-intensive process and existing methods for in silico estimation of solubility are limited by their generality, speed, and accuracy. This work presents two models derived from the FASTPROP and CHEMPROP architectures and trained on BigSolDB which are capable of predicting solubility at arbitrary temperatures for a wide range of small molecules in organic solvent. Both extrapolate to unseen solutes 2–3 times more accurately than the current state-of-the-art model and we demonstrate that they are approaching the aleatoric limit (0.5–1$$\log S$$ log S ) of available test data, suggesting that further improvements in prediction accuracy require more accurate datasets. The FASTPROP-derived model (called FASTSOLV) and the CHEMPROP-based model are open source, freely accessible via a Python package and web interface, highly reproducible, and up to 2 orders of magnitude faster than current alternatives.

Science & Technology - Other Topics↗

Continuous integration data-driven platform of industrial-scale subsurface storage for real-time analytics

This project helped address the growing need for efficient and scalable models to support geological carbon and energy storage, which are crucial for achieving net-zero emissions. Traditionally accurate high-fidelity numerical models have been used to simulate relevant storage processes under a handful of processes, however such models are computationally demanding, making uncertainty quantification impractical. Consequently, we first developed a machine learning framework, based on Graph Neural Operators (GNOs), to improving the accuracy of model predictions for a fixed computational budget. We then developed an Ensemble of Improved Neural Operators (ENO), which uses bagging and Monte Carlo dropout techniques, to further improve prediction accuracy. Lastly, we developed the way to explain progressive transfer learning methods to reduce the amount of training data and computational cost of training (i.e., reduce trainable parameters) when using our models for multiple storage sites. Our numerical investigation, which used real-world case studies, demonstrated that our framework can significantly improve the safety and efficiency of geological storage operations, with potential applications in other domains such as geothermal reservoirs and climate modeling.

54 ENVIRONMENTAL SCIENCES↗

EMPDF : inferring the Milky Way mass with data-driven distribution function in phase space

We introduce the emPDF (empirical distribution function), a novel dynamical modelling method that infers the gravitational potential from kinematic tracers with optimal statistical efficiency under the minimal assumption of steady state. emPDF determines the best-fitting potential by maximizing the similarity between instantaneous kinematics and the time-averaged phase-space distribution function (DF), which is empirically constructed from observation upon the theoretical foundation of oPDF (Han et al. 2016). This approach eliminates the need for presumed functional forms of DFs or orbit libraries required by conventional DF- or orbit-based methods. emPDF stands out for its flexibility, efficiency, and capability in handling observational effects, making it preferable to the popular Jeans equation or other minimal assumption methods, especially for the Milky Way (MW) outer halo where tracers often have limited sample size and poor data quality. We apply emPDF to infer the MW mass profile using Gaia DR3 data of satellite galaxies and globular clusters, obtaining enclosed masses of M (,r) = 26±8, 46±8, 90±13⁠, and 149±40 x 10 10 M ⊙ at r = 30, 50, 100⁠, and 200 kpc, respectively. These are consistent with the updated constraints from simulation-informed DF fitting (Li et al. 2020). While the simulation-informed DF offers superior precision owing to the additional information extracted from simulations, emPDF is independent of such supplementary knowledge and applicable to general tracer populations. emPDF is currently implemented for tracers with complete 6D kinematics within spherical potentials, but it can potentially be extended to address more general problems.

Astrophysics of Galaxies (astro-ph.GA)↗

A data-driven approach to real-time vertical position estimation for NSTX-U vertical stability control

In this paper, a database of 77 996 plasma equilibrium reconstructions from 727 discharges during the initial operation of the NSTX-U spherical tokamak is analyzed to develop a statistically robust model of the plasma vertical position for real-time control. A variety of regression models are developed and tested, ranging in complexity from linear models to deep neural networks, and including input signals ranging from the four pairs of flux loops used historically on NSTX-U up to the full set of 389 real-time signals available to the plasma control system. A linear model based on 140 real-time magnetics signals is found to offer excellent accuracy, with a coefficient of determination R 2 = 0.906. The robustness of this model to limited training data, new operating scenarios, and signal errors is tested, and a procedure is demonstrated to tune the model parameters to optimize its robustness. A time-dependent plasma equilibrium solver, TokaMaker, is used to simulate vertical stability control in NSTX-U, demonstrating that it should be possible to iteratively tune the parameters of a linear vertical position model to stabilize both positive and negative triangularity plasmas in future experiments.

magnetic diagnostics↗

A data-driven method to estimate the antiproton background in the Mu2e experiment

The Mu2e experiment at Fermilab will search for the Charged Lepton Flavour Violating (CLFV) process of coherent, neutrinoless µ− → e − conversion in the field of an aluminum nucleus. The expected signal is a monochromatic electron with the energy of 104.97 MeV, slightly below the muon rest mass. Observation of a CLFV process would provide unambiguous evidence for Beyond the Standard Model (BSM) physics. Mu2e is sensitive to a wide range of BSM models and has the capability to distinguish between them, guiding us towards the most accurate models. The key features of the Mu2e experiment are: (1) a high intensity pulsed negative muon beam with about 1010 stopped µ −/s, and (2) a sophisticated superconducting solenoid system with a gradient magnetic field to form and guide the intense muon beam to the target. The Mu2e physics data taking is expected to begin in 2027. For Run I, the expected 5σ discovery sensitivity is Rµe = 1.2 × 10−15, with a total expected background of 0.11 ± 0.03 events. In the absence of a signal, the expected upper limit is Rµe < 6.2 × 10−16 at 90% CL. The success of this experiment hinges on the accurate estimation of the background from various SM processes that could provide signal-like electrons. One of the background processes is antiprotons annihilating in the stopping target to produce signal like electrons through π0 → γγ decays followed by γ conversions, and π− → µ−ν¯ decays followed by µ− decay. It is a relatively small background with large uncertainty (100%) due to the lack of antiproton production cross section information for the Mu2e proton beam energy of 8 GeV. We have developed a novel methodology to estimate the antiproton background in-situ. This forms the main theme of the thesis. We observed that at Mu2e energies, antiproton annihilation in the stopping target is the only source of events with multiple, simultaneous particle trajectories. From Geant4 simulations, only about 0.2% of the simulated antiproton annihilation events have a signal-like electron. Meanwhile, ∼ 5% of events have multiple reconstructible particle tracks per event. Therefore, we have devised a methodology to reconstruct the multi-track events and estimate the antiproton background by exploiting the large ratio of the production rates of the two final states.

Chithirasreemadam, Namitha [Pisa U.] (ORCID:000000↗

Towards Improving luminosity using optics tuning and data-driven methods

The results of Run 24 experiments at Relativistic Heavy Ion Collider (RHIC) for improving luminosity using optics tuning are presented in this study. In the first experiment, MADx matching was used to output magnet strengths corresponding to specific s star movements around Interaction Region 8 (IR8). The corresponding Zero Degree Calorimeter (ZDC) signal was measured in place of luminosity, and Bayesian Optimization aids search of optimal movements. It was found that values retrieved from matching were inaccurate, resulting in negative feedback loops. The second experiment focused on calculating accurate s star movements. The matching method was replaced with a linear sensitivity matrix, directly relating optics to power supply, and its null space was used to fit constraints such as hysteresis effects. At the experiment, beam losses were observed at collimators around boundary of IR8, which were fixed for the third experiment. Dynamic mode decomposition was also introduced to improve quality of turn-by-turn (TBT) data as well as accuracy and consistency of optics measurements at IR8. These improvements will be tested in the experiment of next RHIC run for luminosity optimization.

Accelerator Physics↗

Physics-coupled data-driven design of high-temperature alloys

We present a materials design loop, which streamlines physics-coupled machine learning (ML) surrogate models to discover new alloy chemistries with improved properties. The efficacy is demonstrated by discovering a high-temperature alumina-forming austenitic (AFA) stainless steel with enhanced creep, followed by experimental validation. The ML models have been trained using a well-curated, highly consistent experimental dataset augmented with synthetic microstructural features from a computational thermodynamic approach. We have populated a large number of hypothetical AFA alloys to explore the high-dimensional composition space and have predicted their creep properties by providing the same synthetic input features obtained from the trained ML models. Uncertainties from the ML training were taken as thresholds for truncating predicted results to identify alloys with improved or deteriorated creep. Individual elemental compositions have been determined via probability density distribution analysis from the group of alloys at the top and bottom of the predicted creep values for further virtual and experimental validations. In conclusion, we anticipate that this workflow can be applied to screen desired conditions, such as chemistry and processing parameters, in high-dimensional space through physics-guided data analytics.

Alloy design↗

Data-Driven Clustering and Classification of Outage Patterns with Insights into their Links to Extreme Events

At a global level extreme events have increased in both scale and impact. These events have the potential to affect the electrical grid infrastructure and cause a wide range of outages, which can lead to a disruption in daily patterns, cost millions of dollars and also the loss of life. Currently, to track these outage events there have been various approaches developed ranging from regional to national level quantifications for what defines an outage. However, this variation in methods can potentially lead to subjective decision-making and a lack of proper management in relation to the event. While previous work has made strides in determining spatio-temporal patterns, minimal attention has been given to the type and number of outages an area may be exposed to. The differences in incurred cost and the overall severity of an event between a transformer box malfunction and a hurricane are drastic, and by finding historical signals, we can allow for more efficient management, potentially saving lives and millions of dollars. Here, we leverage unsupervised machine learning techniques to delineate outage patterns among 22 counties within the United States and find that there are clear, segregated clusters (0.93 silhouette) of data which are related by event behavior and underlying cause. This finding will allow for energy stakeholders, policy makers, and researchers to gain a deeper understanding of the extent and severity of historic events and to better prepare for electrical grid infrastructure planning and management.

Koob, Benjamin [ORNL]↗

Experimental and data-driven characterization of window-induced air leakage in residential buildings

Windows contributes up to 40% of envelope heat losses and around 9% of total building energy consumption due to air leakage. In the U.S., 48 million homes still use single-pane windows. Although the U.S. has an estimated 1.4 billion windows in its building stock and about 24 million windows are installed annually, only around 29 million individual window replacements (∼2%) occur each year. To address this gap, this study generates empirical evidence by (1) evaluating the contribution of windows to whole-building air leakage in 20 residential buildings using blower door tests before and after window replacement and (2) assessing whether building and window characteristics influence the measured change. Most simulation studies assume that replacing windows not only lowers the U-factor but also reduces air leakage by 10–20%. However, this assumption lacks empirical validation, highlighting the need for experimental analysis of air leakage specifically associated with windows. Using blower door tests in accordance with ASTM E779–19, the results indicated an average reduction in air infiltration of 6.1% within the range of 0.5–19.30% across all buildings and no significant correlations were found between air leakage improvements and any building/window characteristics. This research aims to help homeowners, and energy modelers to provide empirical data on importance of upgrading to more energy-efficient windows, supporting energy-efficient building standards.

Air leakage↗

A Data-Driven Exploration of the Impact of Renewable Energy on Inter-Area Oscillations in the U.S. Eastern Interconnection

As increasing amounts of renewable energy (RE) resources are incorporated into the bulk-power grid, power system oscillations are expected to change. This work investigates how RE generation impacts the frequency and damping ratio (DR) of two dominant inter-area modes in the U.S. Eastern Interconnection (EI) using regularly updated estimates collected over a 12-month period. Quantile regression is used to derive the correlation between operating conditions and mode properties, and a bootstrap method is used to quantify the uncertainty associated with the correlation estimates. Results show that with an increase in system load, the frequency of a mode decreases and DR increases. Evidence that increasing RE generation results in an increase in frequency and decline in DR was found for one of the two modes studied. This work shows that increasing RE levels will impact the properties of inter-area oscillations in the EI, but it does not indicate the presence of immediate threats to grid stability. The outlined approach can be used to periodically assess changing mode properties as RE levels continue to grow and flag stability concerns before they become serious reliability threats.

Inter-area oscillation, mode meters, quantile regr↗