Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data driven”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25

Simulation of the Solar Energetic Particle Event on 2020 May 29 Observed by Parker Solar Probe

This paper presents a stochastic three-dimensional focused transport simulation of solar energetic particles (SEPs) produced by a data-driven coronal mass ejection (CME) shock propagating through a data-driven model of coronal and heliospheric magnetic fields. The injection of SEPs at the CME shock is treated using diffusive shock acceleration of post-shock suprathermal solar wind ions. A time-backward stochastic simulation is employed to solve the transport equation to obtain the SEP time–intensity profile at any location, energy, and pitch angle. The model is applied to a SEP event on 2020 May 29, observed by STEREO-A close to ∼1 au and by Parker Solar Probe (PSP) when it was about 0.33 au away from the Sun. The SEP event was associated with a very slow CME with a plane-of-sky speed of 337 km s −1 at a height below 6 RS as reported in the SOHO/LASCO CME catalog. We compute the time profiles of particle flux at PSP and STEREO-A locations, and estimate both the spectral index of the proton energy spectrum for energies between ∼2 and 16 MeV and the equivalent path length of the magnetic field lines experienced by the first arriving SEPs. We find that the simulation results are well correlated with observations. The SEP event could be explained by the acceleration of particles by a weak CME shock in the low solar corona that is not magnetically connected to the observers.

Solar energetic particles↗

Prognostics for Systems Health Management - Model and Hybrid Based Approaches. Where are We Heading?

To facilitate and solve the prediction problem, awareness of the current state and health of the system is key, since it is necessary to perform condition-based system health predictions. To accurately predict the future state of any system, it is required to possess knowledge of its current health state and future operational conditional. In case of next generation electric aircrafts, computing remaining flying time is safety-critical, since an aircraft that runs out of power (battery charge) while in the air will eventually lose control leading to catastrophe. In order to tackle and solve the prediction problem, it is essential to have awareness of the current health state of the system, especially since it is necessary to perform condition-based predictions. To be able to predict the future state of the system, it is also required to possess knowledge of the current and future operational conditions and flight profiles for accurate estimation of end-of-discharge (EOD) for the batteries. Similar framework can be implemented to other complex systems and subsystems. Our research approach is to develop a system level health monitoring safety indicator which runs estimation and prediction algorithms to estimate remaining useful life predictions at system, subsystem swell as component levels. Given models of the current and future system behavior, a general approach of model-based prognostics is discussed as a solution to the prediction problem and further for decision making. Data driven prognostics approaches have been equally used with good results in the past, where respective approaches have their own challenges to tackle. This limits their applicability to complex real-world domains: (a) high complexity or incompleteness of physics-based models and (b) limited representativeness of the training dataset for data-driven models. With the advent of internet of things for data collection and increased use of ML algorithms, hybrid approaches are the next avenue to reduce the challenges and achieve better results. An hybrid framework for fusing information from physics-based performance models along with deep learning algorithms for prognostics of complex safety critical systems is presented. In this framework, we use physics-based performance models to infer unobservable model parameters related to the system's components health solving a calibration problem.

Prognostics↗

Roadmap and Benchmarking: Privacy in Federated Load Forecasting

Data-driven techniques for energy demand forecasting continue to emerge with promising impacts on distribution grid planning. However, the development of robust and generalizable machine learning models requires that representative high quality training data are available. Distributed energy resources have begun to embed intelligence, gathering large amounts of data on customer demand, behavior, and household devices that are connected to the grid. Though utilities aggregate meter-level demand data for load shaping, demand response, outage management, reliability planning, and billing applications, there lies an inherent privacy concern in sharing consumption data that may identify individual consumer behavioral patterns. Hence, while sharing the data is crucial, the private sensitive customer data must be safeguarded from being exposed or manipulated. In this study, we propose a roadmap for implementing a based privacy preserving framework to support the advancement of data-driven analytics in data-sensitive distributed energy resources environments. The roadmap incorporates federated learning–a distributed training framework, differential privacy–a statistical framework that provides guarantees to safeguard the leakage of sensitive data, secure multiparty computation and homomorphic encryption– techniques for encrypting model gradients and applying secure aggregation on the server. Moreover, we perform baseline experiments on the federated short-term load forecasting (STLF) task using open-source residential load profile datasets, offering insights into the challenges of integrating differential privacy into federated learning.

Abebe, Waqwoya [Oak Ridge National Laboratory (ORN↗

ARM-IRL: Adaptive Resilience Metric Quantification Using Inverse Reinforcement Learning

The resilience of safety-critical systems is gaining importance due to the rise in cyber and physical threats, especially within critical infrastructure. Traditional static resilience metrics may not capture dynamic system states, leading to inaccurate assessments and ineffective responses to cyber threats. This work aims to develop a data-driven, adaptive method for resilience metric learning. We propose a data-driven approach using inverse reinforcement learning (IRL) to learn a single, adaptive resilience metric. The method infers a reward function from expert control actions. Unlike previous approaches using static weights or fuzzy logic, this work applies adversarial inverse reinforcement learning (AIRL), training a generator and discriminator in parallel to learn the reward structure and derive an optimal policy. The proposed approach is evaluated on multiple scenarios: optimal communication network rerouting, power distribution network reconfiguration, and cyber–physical restoration of critical loads using the IEEE 123-bus system. The adaptive, learned resilience metric enables faster critical load restoration in comparison to conventional RL approaches.

97 MATHEMATICS AND COMPUTING↗

Weak-form latent space dynamics identification

Recent work in data-driven modeling has demonstrated that a weak formulation of model equations enhances the noise robustness of a wide range of computational methods. In this paper, we demonstrate the power of the weak form to enhance the LaSDI (Latent Space Dynamics Identification) algorithm, a recently developed data-driven reduced order modeling technique. We introduce a weak form-based version WLaSDI (Weak-form Latent Space Dynamics Identification). WLaSDI first compresses data, then projects onto the test functions and learns the local latent space models. Notably, WLaSDI demonstrates significantly enhanced robustness to noise. With WLaSDI, the local latent space is obtained using weak-form equation learning techniques. Compared to the standard sparse identification of nonlinear dynamics (SINDy) used in LaSDI, the variance reduction of the weak form guarantees a robust and precise latent space recovery, hence allowing for a fast, robust, and accurate simulation. We demonstrate the efficacy of WLaSDI vs. LaSDI on several common benchmark examples including viscid and inviscid Burgers', radial advection, and heat conduction. For instance, in the case of 1D inviscid Burgers' simulations with the addition of up to 100% Gaussian white noise, the relative error remains consistently below 6% for WLaSDI, while it can exceed 10,000% for LaSDI. Similarly, for radial advection simulations, the relative errors stay below 15% for WLaSDI, in stark contrast to the potential errors of up to 10,000% with LaSDI. Moreover, speedups of several orders of magnitude can be obtained with WLaSDI. For example applying WLaSDI to 1D Burgers' yields a 140X speedup compared to the corresponding full order model.

97 MATHEMATICS AND COMPUTING↗

Toward Accelerating Discovery via Physics-Driven and Interactive Multifidelity Bayesian Optimization

Both computational and experimental material discovery bring forth the challenge of exploring multidimensional and often nondifferentiable parameter spaces, such as phase diagrams of Hamiltonians with multiple interactions, composition spaces of combinatorial libraries, processing spaces, and molecular embedding spaces. Often these systems are expensive or time consuming to evaluate a single instance, and hence classical approaches based on exhaustive grid or random search are too data intensive. This resulted in strong interest toward active learning methods such as Bayesian optimization (BO) where the adaptive exploration occurs based on human learning (discovery) objective. However, classical BO is based on a predefined optimization target, and policies balancing exploration and exploitation are purely data driven. In practical settings, the domain expert can pose prior knowledge of the system in the form of partially known physics laws and exploration policies often vary during the experiment. Here, we propose an interactive workflow building on multifidelity BO (MFBO), starting with classical (data-driven) MFBO, then expand to a proposed structured (physics-driven) structured MFBO (sMFBO), and finally extend it to allow human-in-the-loop interactive interactive MFBO (iMFBO) workflows for adaptive and domain expert aligned exploration. These approaches are demonstrated over highly nonsmooth multifidelity simulation data generated from an Ising model, considering spin–spin interaction as parameter space, lattice sizes as fidelity spaces, and the objective as maximizing heat capacity. Detailed analysis and comparison show the impact of physics knowledge injection and real-time human decisions for improved exploration with increased alignment to ground truth. Here, the associated notebooks allow to reproduce the reported analyses and apply them to other systems.

97 MATHEMATICS AND COMPUTING↗

Electronic Health Management

Accelerated aging methodologies for electrolytic components have been designed and accelerated aging experiments have been carried out. The methodology is based on imposing electrical and/or thermal overstresses via electrical power cycling in order to mimic the real world operation behavior. Data are collected in-situ and offline in order to periodically characterize the devices' electrical performance as it ages. The data generated through these experiments are meant to provide capability for the validation of prognostic algorithms (both model-based and data-driven). Furthermore, the data allow validation of physics-based and empirical based degradation models for this type of capacitor. A first set of models and algorithms has been designed and tested on the data.

Celaya, Jose R.↗

Global CO2 Distributions over Land from the Greenhouse Gases Observing Satellite (GOSAT)

January 2009 saw the successful launch of the first space-based mission specifically designed for measuring greenhouse gases, the Japanese Greenhouse gases Observing SATellite (GOSAT). We present global land maps (Level 3 data) of column-averaged CO2 concentrations (X(sub CO2)) derived using observations from the GOSAT ACOS retrieval algorithm, for July through December 2009. The applied geostatistical mapping approach makes it possible to generate maps at high spatial and temporal resolutions that include uncertainty measures and that are derived directly from the Level 2 observations, without invoking an atmospheric transport model or estimates of CO2 uptake and emissions. As such, they are particularly well suited for comparison studies. Results show that the Level 3 maps for July to December 2009 on a lO x 1.250 grid, at six-day resolution capture much of the synoptic scale and regional variability of X(sub CO2), in addition to its overall seasonality. The uncertainty estimates, which reflect local data coverage, X(sub CO2) variability, and retrieval errors, indicate that the Southern latitudes are relatively well-constrained, while the Sahara Desert and the high Northern latitudes are weakly-constrained. A probabilistic comparison to the PCTM/GEOS-5/CASA-GFED model reveals that the most statistically significant discrepancies occur in South America in July and August, and central Asia in September to December. While still preliminary, these results illustrate the usefulness of a high spatiotemporal resolution, data-driven Level 3 data product for direct interpretation and comparison of satellite observations of highly dynamic parameters such as atmospheric CO2.

Hammerling, Dorit M.↗

Datalist: A Value Added Service to Enable Easy Data Selection

Imagine a user wanting to study hurricane events. This could involve searching and downloading multiple data variables from multiple data sets. The currently available services from the Goddard Earth Sciences Data and Information Services Center (GES DISC) only allow the user to select one data set at a time. The GES DISC started a Data List initiative, in order to enable users to easily select multiple data variables. A Data List is a collection of predefined or user-defined data variables from one or more archived data sets. Target users of Data Lists include science teams, individual science researchers, application users, and educational users. Data Lists are more than just data. Data Lists effectively provide users with a sophisticated integrated data and services package, including metadata, citation, documentation, visualization, and data-specific services, all available from one-stop shopping. Data Lists are created based on the software architecture of the GES DISC Unified User Interface (UUI). The Data List service is completely data-driven, and a Data List is treated just as any other data set. The predefined Data Lists, created by the experienced GES DISC science support team, should save a significant amount of time that users would otherwise have to spend.

Datalist↗

On the Impact of Granularity of Space-Based Urban CO2 Emissions in Urban Atmospheric Inversions: A Case Study for Indianapolis, IN

Quantifying greenhouse gas (GHG) emissions from cities is a key challenge towards effective emissions management. An inversion analysis from the INdianapolis FLUX experiment (INFLUX) project, as the first of its kind, has achieved a top-down emission estimate for a single city using CO2 data collected by the dense tower network deployed across the city. However, city-level emission data, used as a priori emissions, are also a key component in the atmospheric inversion framework. Currently, fine-grained emission inventories (EIs) able to resolve GHG city emissions at high spatial resolution, are only available for few major cities across the globe. Following the INFLUX inversion case with a global 1x1 km ODIAC fossil fuel CO2 emission dataset, we further improved the ODIAC emission field and examined its utility as a prior for the city scale inversion. We disaggregated the 1x1 km ODIAC non-point source emissions using geospatial datasets such as the global road network data and satellite-data driven surface imperviousness data to a 3030 m resolution. We assessed the impact of the improved emission field on the inversion result, relative to priors in previous studies (Hestia and ODIAC). The posterior total emission estimate (5.1 MtC/yr) remains statistically similar to the previous estimate with ODIAC (5.3 MtC/yr). However, the distribution of the flux corrections was very close to those of Hestia inversion and the model-observation mismatches were significantly reduced both in forward and inverse runs, even without hourly temporal changes in emissions. EIs reported by cities often do not have estimates of spatial extents. Thus, emission disaggregation is a required step when verifying those reported emissions using atmospheric models. Our approach offers gridded emission estimates for global cities that could serves as a prior for inversion, even without locally reported EIs in a systematic way to support city-level Measuring, Reporting and Verification (MRV) practice implementation.

EIs↗

Developing Deep Learning Models for System Remaining Useful Life Predictions: Application to Aircraft Engines

Prognostics and health management (PHM) is an important part of ensuring reliable operations of complex safety- critical systems. System-level remaining useful life (RUL) estimation is a much more complex problem than making estimations at the component level, and system-level RUL methodologies remain sparse in the literature. Model-based approaches have traditionally worked in the past for components such as capacitors, MOSFETs, batteries, or hard-drives (to name a few examples), but developing high fidelity dynamics models of cyber physical systems that can be used to study the effects of multiple degrading components in the system remains a challenging task. Some initial work on model-based System RUL predictions was demonstrated in Khorasgani, et al [1], but, to generalize the system-level prognostics problem, we have to resort to pure data driven and hybrid approaches. In this work, we propose an end-to-end data- driven framework for developing deep learning models to predict remaining useful life of cyber physical systems operating under unknown faulty conditions. The raw data is organized with a data schema that improves the model development process and down stream data analysis tasks. Due to the unknown faulty conditions, the raw sensor data is transformed into signals that expose the underlying degradation processes, which are then used for model development. Bayesian Optimization is used to tune the model parameters prior to training and validation. We show that this approach results in accurate predictions within 3 cycles to end of life (EOL). We demonstrate the effectiveness of our approach by applying it to the N-CMAPSS turbofan engine dataset recently released by NASA, which includes high fidelity degradation modeling, real world operating conditions, and a large set of fault operating modes.

Prognostics↗

Pavement condition and climatic data in southeast Texas: A dataset for evaluating flood impacts on pavement performance

Effective pavement maintenance is essential for economic stability, optimal network performance, and roadway safety. Achieving this requires thorough evaluation of pavement conditions, including structural integrity, surface roughness, and distress characteristics. Pavement performance indicators play a critical role in influencing vehicle safety and ride quality. Recent advances have emphasized the use of data-driven modeling to anticipate pavement behavior, with the goal of optimizing resource allocation and refining Maintenance and Rehabilitation (M&R) strategies through accurate condition assessment. A foundational requirement for these modeling efforts is the availability of standardized, high-quality datasets that can support robust and reproducible infrastructure analysis. This data article presents a comprehensive dataset assembled to facilitate pavement performance prediction, with a geographic focus on Southeast Texas, particularly the flood-vulnerable area of Beaumont. The dataset encompasses pavement and traffic attributes, meteorological records, flood simulation outputs, ground deformation measurements, and topographic indices, enabling detailed examination of both load-associated and non-load-associated degradation mechanisms. Data preprocessing was performed using ArcGIS Pro, Microsoft Excel, and Python to ensure consistency and usability in data-driven modeling applications, including machine learning workflows. Key contributions of this dataset include its utility in analyzing the climatic and environmental factors affecting pavement conditions, identifying critical predictive features, and enabling in-depth correlation analysis across diverse variables. By filling existing gaps in input variable selection resources, this dataset supports the development of predictive tools for estimating future maintenance demand and enhancing the resilience of pavement networks in flood-impacted areas. The resource highlights the importance of standardized datasets for advancing pavement management practices and provides a robust foundation for ongoing infrastructure performance modeling.

42 ENGINEERING↗

Equipping Neural Network Surrogates with Uncertainty for Propagation in Physical Systems

Coarse-grained or filtered models typically rely on closure models to account for unresolved scales. For instance, large eddy simulation for modeling turbulent fluid flows explicitly resolves the largest scales, but requires modeling closure terms to account for the sub-filter scales. With the vast amount of data available from high-fidelity simulations, there are unique opportunities to leverage data-driven modeling techniques to formulate expressive and flexible closure models. Despite their flexibility, data-driven models struggle in domain shift settings, i.e. when deployed in configurations not captured in the training dataset. In particular, the efficacy of neural network surrogates is difficult to assess a priori due to the deterministic, point-estimate nature of predictions. In high-consequence applications, such models require reliable uncertainty estimates in the data-informed and out-of-distribution regimes. To quantify uncertainties in both regimes, we employ Bayesian neural networks which are able to capture both epistemic and aleatoric uncertainties. We will discuss challenges associated with the training and evaluation of these networks. Furthermore, we will discuss uncertainty embedding strategies to enable efficient sampling and propagation of uncertainty through high-fidelity simulations.

Bayesian neural networks↗

From disorganized data to emergent dynamic models: Questionnaires to partial differential equations

Starting with sets of disorganized observations of spatially varying and temporally evolving systems, obtained at different (also disorganized) sets of parameters, we demonstrate the data-driven derivation of parameter dependent, evolutionary partial differential equation (PDE) models capable of generating the data. This tensor type of data is reminiscent of shuffled (multidimensional) puzzle tiles. The independent variables for the evolution equations (their “space” and “time”) as well as their effective parameters are all emergent , i.e. determined in a data-driven way from our disorganized observations of behavior in them. We use a diffusion map based questionnaire approach to build a smooth parametrization of our emergent space/time/parameter space for the data. This approach iteratively processes the data by successively observing them on the “space,” the “time” and the “parameter” axes of a tensor. Once the data become organized, we use machine learning (here, neural networks) to approximate the operators governing the evolution equations in this emergent space. Our illustrative examples are based (i) on a simple advection–diffusion model; (ii) on a previously developed vertex-plus-signaling model of Drosophila embryonic development; and (iii) on two complex dynamic network models (one neuronal and one coupled oscillator model) for which no obvious smooth embedding geometry is known a priori. This allows us to discuss features of the process like symmetry breaking, translational invariance, and autonomousness of the emergent PDE model, as well as its interpretability.

generative models↗

SODAs: sparse optimization for the discovery of differential and algebraic equations

Differential-algebraic equations (DAEs) integrate ordinary differential equations (ODEs) with algebraic constraints, providing a fundamental framework for developing models of dynamical systems characterized by time-scale separation, conservation laws and physical constraints. While sparse optimization has revolutionized model development by allowing data-driven discovery of parsimonious models from a library of possible equations, existing approaches for dynamical systems assume DAEs can be reduced to ODEs by eliminating variables before model discovery. This assumption limits the applicability of such methods for DAE systems with unknown constraints and time scales. We introduce sparse optimization for differential-algebraic systems (SODAs), a data-driven method for the identification of DAEs in their explicit form. By discovering the algebraic and dynamic components sequentially without prior identification of the algebraic variables, this approach leads to a sequence of convex optimization problems. It has the advantage of discovering interpretable models that preserve the structure of the underlying physical system. To this end, SODAs improves since SODAs is singular numerical stability when handling high correlations between library terms, caused by near-perfect algebraic relationships, by iteratively refining the conditioning of the candidate library. We demonstrate the performance of our method on biological, mechanical and electrical systems, showcasing its robustness to noise in both simulated time series and real-time experimental data.

DAE↗

Optimal Control of an Oscillating Surge Wave Energy Converter

During this project, we experimentally investigated the hydrodynamics and performance of a laboratory-scale oscillating surge wave energy converter (OSWEC).We looked at how flap buoyancy and driveline losses (primarily in the form of stiction) affected the dynamics and performance of the device. In addition, we assessed the influence of flap profile (rounded vs. square edges) on OSWEC hydrodynamics. Through this, we were able to develop a deeper understanding of OSWEC performance and provide guidance on strategies to counteract artifacts that may be present in laboratory models, but are absent in field-scale devices. To do this, we tested a laboratory-scale OSWEC in the Sea Wave Environmental Lab (SWEL) wave tank at the National Renewable Energy Laboratory (NREL). We ran several types of experiments to investigate the hydrodynamics and performance of the device. Overall, we achieved the overall goal of experimentally investigating the hydrodynamics and performance of this device. We discovered important and unexpected trends in performance, and collected time-resolved data to help us further investigate the underlying hydrodynamics responsible for these trends. In addition, we are currently using the time-resolved data from these experiments to build data-driven models of the dynamics, which can in turn be used to inform data-driven model predictive control of this device and address this objective in the future.

16 TIDAL AND WAVE POWER↗

Harnessing Satellite Data Alone for Mapping Global Thermal Anisotropy

Mapping thermal anisotropy across global lands is critical for advancing a wide range of Earth science studies. However, a comprehensive understanding of global thermal anisotropy intensity (TAI) and its governing factors remains missing. We introduce a novel data-driven methodology to quantify global TAI exclusively using multi-angle MODIS land surface temperature time series observations. Our analysis reveals distinct seasonal and diurnal TAI patterns, with global mean summertime TAI exceeding 2.9°C. Furthermore, we identify strong associations between TAI and key surface and atmospheric parameters, such as leaf area index and downward shortwave radiation. Our findings advocate for a paradigm shift from model-based to data-driven approaches in correcting thermal anisotropy, thereby addressing a critical bottleneck in Earth observation.

54 ENVIRONMENTAL SCIENCES↗

Hybrid Modeling of Three-Phase Grid-Supporting Inverters for Dynamic Studies

Grid technologies connected by power electronic converter (PEC) interfaces continually implement grid support functions mandated by grid codes and standards. The transition to converter-based generation demands precise PEC models to assess system dynamics, which have been previously overlooked in conventional power systems. This study proposes a hybrid method for analyzing grid-connected three-phase PEC dynamics with the IEEE standard 1547-2018 Volt-VAr mode that combines physics and data-driven techniques. The physics model reflects the PEC’s internal behavior, whereas the data-driven modeling technique evaluates the grid-supporting capabilities of the smart PEC. The system identification approach is used to generate dynamic PEC models based on changing grid voltage and measured current injected into the grid by the PEC. In the Volt-VAr support mode, a detailed topological model including switches is utilized to compare the goodness-of-fit of the extracted hybrid dynamic model. The results demonstrate that the hybrid PEC model in the Volt-VAr mode accurately matches the dynamics with the topological model.

Subedi, Sunil↗