Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data driven”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Optimal Control of an Oscillating Surge Wave Energy Converter

During this project, we experimentally investigated the hydrodynamics and performance of a laboratory-scale oscillating surge wave energy converter (OSWEC).We looked at how flap buoyancy and driveline losses (primarily in the form of stiction) affected the dynamics and performance of the device. In addition, we assessed the influence of flap profile (rounded vs. square edges) on OSWEC hydrodynamics. Through this, we were able to develop a deeper understanding of OSWEC performance and provide guidance on strategies to counteract artifacts that may be present in laboratory models, but are absent in field-scale devices. To do this, we tested a laboratory-scale OSWEC in the Sea Wave Environmental Lab (SWEL) wave tank at the National Renewable Energy Laboratory (NREL). We ran several types of experiments to investigate the hydrodynamics and performance of the device. Overall, we achieved the overall goal of experimentally investigating the hydrodynamics and performance of this device. We discovered important and unexpected trends in performance, and collected time-resolved data to help us further investigate the underlying hydrodynamics responsible for these trends. In addition, we are currently using the time-resolved data from these experiments to build data-driven models of the dynamics, which can in turn be used to inform data-driven model predictive control of this device and address this objective in the future.

16 TIDAL AND WAVE POWER↗

Harnessing Satellite Data Alone for Mapping Global Thermal Anisotropy

Mapping thermal anisotropy across global lands is critical for advancing a wide range of Earth science studies. However, a comprehensive understanding of global thermal anisotropy intensity (TAI) and its governing factors remains missing. We introduce a novel data-driven methodology to quantify global TAI exclusively using multi-angle MODIS land surface temperature time series observations. Our analysis reveals distinct seasonal and diurnal TAI patterns, with global mean summertime TAI exceeding 2.9°C. Furthermore, we identify strong associations between TAI and key surface and atmospheric parameters, such as leaf area index and downward shortwave radiation. Our findings advocate for a paradigm shift from model-based to data-driven approaches in correcting thermal anisotropy, thereby addressing a critical bottleneck in Earth observation.

54 ENVIRONMENTAL SCIENCES↗

Hybrid Modeling of Three-Phase Grid-Supporting Inverters for Dynamic Studies

Grid technologies connected by power electronic converter (PEC) interfaces continually implement grid support functions mandated by grid codes and standards. The transition to converter-based generation demands precise PEC models to assess system dynamics, which have been previously overlooked in conventional power systems. This study proposes a hybrid method for analyzing grid-connected three-phase PEC dynamics with the IEEE standard 1547-2018 Volt-VAr mode that combines physics and data-driven techniques. The physics model reflects the PEC’s internal behavior, whereas the data-driven modeling technique evaluates the grid-supporting capabilities of the smart PEC. The system identification approach is used to generate dynamic PEC models based on changing grid voltage and measured current injected into the grid by the PEC. In the Volt-VAr support mode, a detailed topological model including switches is utilized to compare the goodness-of-fit of the extracted hybrid dynamic model. The results demonstrate that the hybrid PEC model in the Volt-VAr mode accurately matches the dynamics with the topological model.

Subedi, Sunil↗

Stoichiometrically-informed symbolic regression for extracting chemical reaction mechanisms from data

A data-driven computational method is introduced to extract chemical reaction mechanisms from time series chemical concentration data. It is realized through the use of dynamic symbolic regression in which a sparse analytical form for a dynamical system is discoverable from the underlying data. We specifically develop the stoichiometrically-informed symbolic regression (SISR) method to address a standing challenge in complex chemical reaction networks: given a time-series dataset of concentrations of several components, what is the mechanism and the associated rate constants? SISR finds the optimal mechanism, kinetic equations and rate constants by combining differential optimization with a genetic optimization approach that searches a symbolic space of possible reaction mechanisms. Use of SISR in several paradigmatic examples spanning linear and nonlinear reaction schemes results in excellent agreement between true and predicted mechanisms, including when the method is applied to noisy data. The advantages of a stoichiometrically-informed approach such as SISR to address reaction discovery is illustrated through comparison with the use of generic state-of-the-art data-driven approaches.

36 MATERIALS SCIENCE↗

Nonperturbative Guiding Center Model for Magnetized Plasmas

Perturbative guiding center theory adequately describes the slow drift motion of charged particles in the strongly magnetized regime characteristic of thermal particle populations in various magnetic fusion devices. However, it breaks down for particles with large-enough energy. Here, we report on a data-driven method for learning a nonperturbative guiding center model from full-orbit particle simulation data. We show the data-driven model significantly outperforms traditional asymptotic theory in magnetization regimes appropriate for fusion-born α particles in stellarators, thus opening the door to nonperturbative guiding center calculations.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Codiscovering graphical structure and functional relationships within data: A Gaussian Process framework for connecting the dots

Most problems within and beyond the scientific domain can be framed into one of the following three levels of complexity of function approximation. Type 1: Approximate an unknown function given input/output data. Type 2: Consider a collection of variables and functions, some of which are unknown, indexed by the nodes and hyperedges of a hypergraph (a generalized graph where edges can connect more than two vertices). Given partial observations of the variables of the hypergraph (satisfying the functional dependencies imposed by its structure), approximate all the unobserved variables and unknown functions. Type 3: Expanding on Type 2, if the hypergraph structure itself is unknown, use partial observations of the variables of the hypergraph to discover its structure and approximate its unknown functions. These hypergraphs offer a natural platform for organizing, communicating, and processing computational knowledge. While most scientific problems can be framed as the data-driven discovery of unknown functions in a computational hypergraph whose structure is known (Type 2), many require the data-driven discovery of the structure (connectivity) of the hypergraph itself (Type 3). We introduce an interpretable Gaussian Process (GP) framework for such (Type 3) problems that does not require randomization of the data, access to or control over its sampling, or sparsity of the unknown functions in a known or learned basis. Its polynomial complexity, which contrasts sharply with the super-exponential complexity of causal inference methods, is enabled by the nonlinear ANOVA capabilities of GPs used as a sensing mechanism.

Science & Technology - Other Topics↗

Hydropower potential derived from streamflow extremes for Alaska, USA

Alaska is an expansive region known for its abundant natural resources, including thousands of miles of streams and rivers. These rivers represent potential opportunities for future hydropower development that could provide reliable energy supply for local communities. There is limited long-term high temporal resolution streamflow data available for the region, making data-driven estimates of potential hydropower and its variability across the state challenging. This study provides a novel data-driven approach for hydropower capacity estimation across Alaska. We use supervised machine learning to develop a relationship between the daily and peak flow duration curves in order to augment the size of our dataset from 44 sites to 67 sites. We perform a stochastic hydropower estimation across the 67 sites and identify approximately 1000 MW of total potential hydropower capacity distributed across these sites. Our study provides the first step towards more comprehensive hydropower estimation for this critical region, highlighting the need for future work integrating high-resolution spatial data, community needs, and economic constraints in estimates of potential hydropower development in Alaska.

Hydropower↗

HTESP (High-throughput electronic structure package): A package for high-throughput ab initio calculations

High-throughput ab initio calculations are the indispensable parts of data-driven discovery of new materials with desirable properties, as reflected in the establishment of several online material databases. The accumulation of extensive theoretical data through computations enables data-driven discovery by constructing machine learning and artificial intelligence models to predict novel compounds and forecast their properties. Efficient usage and extraction of data from these existing online material databases can accelerate the next stage materials discovery that targets different and more advanced properties, such as electron–phonon coupling for phonon-mediated superconductivity. However, extracting data from these databases, generating tailored input files for different ab initio calculations, performing such calculations, and analyzing new results can be demanding tasks. Here, in this work, we introduce a software package named “HTESP” (High-Throughput Electronic Structure Package) written in Python and Bash languages, which automates the entire workflow including data extraction, input file generation, calculation submission, result collection and plotting. Our HTESP will help speed up future computational materials discovery processes.

36 MATERIALS SCIENCE↗

A Privacy First Path Analysis using Clickstream Data

In the modern digital economy, data-driven decision making is crucial for effectively meeting the ever-evolving demands of consumer engagement and satisfaction. Clickstream data has become invaluable for understanding customer behavior, yet concerns over privacy and security persist, especially with some internet service providers profiting from its sale. This article introduces an innovative methodology that blends experiential learning with advanced cryptographic techniques, including differential privacy and graph analytics. The core objective of this methodology is to estimate Customer Lifetime Value (CLV) by analyzing clickstream data, achieving an average prediction accuracy of 92.4% in user engagement levels while ensuring user anonymity through Recency, Frequency, and Monetary (RFM) analysis. Our study introduces the concept of a “data depositor” and a privacy manager, employing the composition theorem to merge non-adaptive queries effectively. Privacy budgets (? = 1.0, d = 10-5), sensitivity-specific techniques, and data partitioning were applied. Randomization and noise addition protect data integrity, with special handling for categorical values. This approach, differing from prior studies, offers a 12.6% improvement in privacy-preserving targeting accuracy while maintaining strict confidentiality, presenting a novel path forward in data-driven decision-making.

Frequency and Monetary (RFM) analysis↗

Unsupervised Learning for Improved Gamma-Ray Spectrometry in Pixelated Cadmium Zinc Telluride (CZT) Detectors

Machine learning has been found to be ubiquitously useful across many industries, presenting an opportunity to improve radiation detection performance using data-driven algorithms. Improved detector resolution can aid in the detection, identification, and quantification of radionuclides. Here, in this work, a novel, data-driven, unsupervised learning approach is developed to improve detector spectral characteristics by learning, and subsequently rejecting, poorly performing regions of the pixelated detector. Feature engineering is used to fit individual characteristic photo peaks to a Doniach lineshape with a linear background model. Then, principal component analysis is used to learn a lower-dimension latent space representation of each photo peak where the pixels are clustered, and subsequently ranked, based on the cluster mean distance to an optimal point. Pixels within the worst cluster(s) are rejected to improve the full-width at half-maximum (FWHM) by 10% to 15% (relative to the bulk detector) at 50% net efficiency when applied to training data obtained from measurements of a 100 μCi 154 Eu source using a H3D M400i pixelated cadmium zinc telluride detector. These results compare well with, but do not outperform, a greedy algorithm that accumulates pixels in order of FWHM from lowest to highest used as a benchmark. In the future, this approach can be extended to include the detector energy and angular response. Finally, the model is applied to newly seen natural and enriched uranium spectra relevant for nuclear safeguards applications.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Deployment of Traditional and Hybrid Machine Learning for Critical Heat Flux Prediction in the CTF Thermal-Hydraulics Code

Critical heat flux (CHF) marks the transition from nucleate to film boiling, where heat transfer to the working fluid can rapidly deteriorate. Accurate CHF prediction is essential for efficiency, safety, and preventing equipment damage, particularly in nuclear reactors. Although widely used, empirical correlations frequently exhibit discrepancies when compared to experimental data, limiting their reliability in diverse operational conditions. Traditional machine learning (ML) approaches have demonstrated potential for CHF prediction but often suffer from limited interpretability, data scarcity, and insufficient knowledge of physical principles. Hybrid model approaches, which combine data-driven ML with base models, mitigate these concerns by incorporating prior knowledge of the domain. This study integrates an externally trained purely data-driven ML model and two hybrid models (using the Biasi and Bowring CHF correlations) within the CTF subchannel code via a custom Fortran framework. Performance was evaluated using two validation cases: a subset of the Nuclear Regulatory Commission (NRC) CHF database and the Bennett dryout experiments. In both cases, the hybrid models demonstrated significantly lower error metrics compared to conventional empirical correlations, with the best models often reducing relative error by about 5 percentage points. The pure ML model achieved comparable accuracy, outperforming the hybrid Biasi model in the NRC test case (3.3% versus 5.5% relative error) but exhibiting slightly higher error against the hybrid Bowring model in the Bennett test case (7.7% versus 6.1%). Trend analysis of error parity indicated that ML-based models reduced the tendency for CHF overprediction, improving overall accuracy. These results demonstrate that ML-based CHF models can be effectively integrated into subchannel codes and could potentially increase performance compared to conventional methods.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

mphys-surrogate-model

This repository contains python scripts for building and studying reduced-order-modeling representations of droplet coalescence for eventual use in atmospheric models. The included data are generated from high-fidelity superdroplet methods and are utilized by machine learning pipelines to build data-driven models of droplet size distributions that evolve under coalescence. This repository further includes scripts to determine prediction (uncertainty) intervals on the data-driven model products based on conformal prediction.

Katona, JonasE [Lawrence Livermore National Labora↗

Deep Learning-Based Dynamic Modeling of Three-Phase Voltage Source Inverters

Inverter-based resource (IBR) models are necessary to analyze modern power system stability and create effective control strategies. Modeling IBRs in converter-rich power systems is crucial, yet challenging due to the lack of commercial information on converter topologies and control parameters. This paper proposes novel convolutional neural network (CNN)–based data-driven techniques for modeling IBRs, addressing adaptability and proprietary concerns without requiring internal system physics knowledge. The proposed method is tested using real grid-tied commercial IBR transient data and demonstrates effectiveness and accuracy. Furthermore, the developed modeling approach is integrated and implemented in the open-source power distribution simulation and analysis tool, GridLAB-D, to illustrate the potentiality of dynamic analysis of large-scale power systems with high IBRs.

deep learning, artificial intelligence↗

Learning Physically Interpretable Atmospheric Models From Data With WSINDy

The multiscale and turbulent nature of Earth's atmosphere has historically rendered accurate weather modeling a hard problem. Recently, there has been an explosion of interest surrounding data-driven approaches to weather modeling, which in many cases show improved forecasting accuracy and computational efficiency when compared to traditional methods. However, many of the current data-driven approaches employ highly parameterized neural networks, often resulting in uninterpretable models and limited gains in scientific understanding. In this work, we address the interpretability problem by explicitly discovering partial differential equations governing atmospheric phenomena, identifying symbolic mathematical models with direct physical interpretations. The purpose of this paper is to demonstrate that, in particular, the weak-form sparse identification of nonlinear dynamics (WSINDy) algorithm can learn effective atmospheric models from both simulated and assimilated data. Our approach adapts the standard WSINDy algorithm to work with high-dimensional fluid data of arbitrary spatial dimension.

58 GEOSCIENCES↗

High-Fidelity Dataset Generation for Sensor Anomalies in Power Grids using Hardware-in-the-Loop Testbed

Sensor anomalies in power grids can have significant impacts on the operation of the grid due to the increased reliance of the grid operation on data-driven applications. However, there is a lack of datasets that accurately capture these anomalies as many of the anomalies go undetected using the current bad data detectors. High-fidelity labeled datasets are essential for developing robust applications that can detect and mitigate the impacts of anomalies. In this paper, we propose a hardware-in-the-loop testbed model that can emulate the grid behavior with high-fidelity. This testbed is used to inject anomalies at various levels in the grid architecture and generate labeled datasets. These high-fidelity datasets can be used for development and validation of data-driven applications for detection and mitigation of anomalies in grids and other cyber-physical systems.

Hyder, Burhan↗

Data-Enabled Fusion Technology (Final Scientific/Technical Report)

Advancing Scientific Understanding in Fusion Energy and Machine Learning This research represented a significant step forward in machine learning (ML) applications for fusion energy experiments. The project integrated advanced data-driven modeling, optimization techniques, and artificial intelligence to enhance the predictive capabilities and operational efficiency of plasma-based fusion systems. Specifically, tasks focused on ML-enhanced diagnostics, operator guidance tools, and predictive modeling helped improve the ability to interpret complex fusion experiments. Key areas of advancement included: 1) data-driven plasma control, i.e., using ML algorithms to optimize experimental conditions and classify plasma behaviors based on historical data; 2) spectroscopy and diagnostics, i.e., applying AI models to extract previously inaccessible insights from experimental spectroscopy data; and 3) configuration mapping and operator guidance, i.e., developing a predictive framework to assist scientists in identifying the most effective experimental parameters, reducing reliance on manual adjustments. By refining these ML-driven techniques, the project contributed to the broader scientific community’s understanding of plasma dynamics and fusion energy viability. Technical Effectiveness and Economic Feasibility The methods investigated demonstrated high technical effectiveness, as reflected in milestones assessing the predictive accuracy, performance, and optimization of fusion configurations. The development of an Operator Guidance Tool (OGT), for example, led to more precise control of plasma conditions by learning from experimental data and offering real-time adjustments. From an economic standpoint, DeFT provided: 1) the ability to reduce trial-and-error experimentation, which lowered operational costs; 2) improved data interpretation methods, which enabled more efficient resource allocation in large-scale fusion research projects; and 3) the automation of key diagnostic tasks, which reduced manual labor and human error, increasing overall efficiency. 13 The final assessments of predictive models and optimization strategies demonstrated that these approaches were scalable and could be implemented across multiple fusion energy research programs. Public Benefit and Societal Impact This project contributed directly to the broader goal of achieving sustainable and commercially viable fusion energy, which had profound implications for clean energy production and climate change mitigation. The integration of AI-driven solutions into fusion research: 1) sped up scientific discovery, accelerating progress towards achieving energy breakthroughs; 2) reduced the cost of experimentation, making fusion research more accessible; and 3) provided a framework for future AI applications in high-energy physics, benefiting adjacent fields like space exploration, material science, and renewable energy. Additionally, by fostering collaborations between AI researchers and plasma physicists, this project promoted interdisciplinary innovation that could lead to broader applications beyond fusion research.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Low‐dimensional manifold learning for uncertainty quantification in complex multi‐scale stochastic systems

Broadly speaking, the goals of the project are to develop techniques to use manifold learning to develop reduced‐order and surrogate models for "hyper‐reduction" of very high‐dimensional complex multi‐scale systems. This is being achieved by employing a newly proposed form of manifold projection and learning that leverages recent advancements in computational geometry and data‐driven modeling. In particular, we are applying a manifold projection technique to project the solutions of very high‐dimensional systems onto the so‐called Grassmannmanifold, a Reimannian manifold comprised of orthonormal matrices. We then apply data‐driven machine learning techniques to classify the solutions on the manifold (e.g. clustering techniques) according to their proximity on the manifold and leverage a further nonlinear dimension reduction to organize the structured data on the manifold. Finally, we are developing novel techniques that enable us to directly interpolate the hyper‐reduced data such that we can predict the solution of the complex, high‐ dimensional system without need to call the full expensive computational model. Given their adherence to the underlying structure of the solution of the physical system, it is expected that these approximate solutions will be sufficiently constrained so as to (approximately) adhere to physical principles.

97 MATHEMATICS AND COMPUTING↗

Diesel Fuel Consumption in Prominent U.S. Open-Pit Mines: Site-Level Estimates

This report presents a comprehensive framework for estimating diesel fuel consumption and prices at open-pit mines in the United States. The framework includes transparent methods for calculating site-level diesel energy use when direct reporting is unavailable, and a structured confidence evaluation for each method. The framework is demonstrated to estimate current diesel consumption at 21 open-pit mines in the United States. Initial findings support ongoing efforts to strengthen the competitiveness and security of the U.S. industrial base by supporting data-driven supply chain analysis and decision-making, improved transparency in mining sector energy use, and targeted deployment of energy innovation and cost-reduction strategies. Future updates to the dataset—coupled with expanded data transparency and method validation—will help ensure that the findings remain relevant as the sector continues to evolve.

02 PETROLEUM↗