Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data driven”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26

Physics-informed machine learning exploration of Na storage mechanisms in disordered carbon

Sodium-ion batteries are a cost-effective, sustainable alternative to lithium-ion systems for large-scale energy storage. However, optimizing sodium storage in carbon-based anodes with microstructural complexity and atomic disorder remains a major challenge. The intrinsic inhomogeneity of these materials produces diverse local environments, making it difficult for conventional methods to predict and control ion dynamics. Hard carbon (HC) anodes, composed of ranges of ordered-to-disordered graphitic and amorphous nanodomains, offer tunable ion storage and rate capacity, yet rationale design remains a challenge due to poorly understood correlation between local atomic feature and ion transport mechanism. Here, to address this challenge, we introduce a data-driven framework that integrates validated machine-learned interatomic potentials, large-scale molecular dynamics simulations, and machine learning to elucidate sodium transport mechanisms as a function of carbon and sodium loading densities. By computing per-ion structural descriptors and applying unsupervised learning, we identify distinct diffusion modes governed by microscopic features. Supervised analysis and correlation mapping then establish quantitative links between these transport regimes and processing variables such as bulk carbon density and sodium content. This physics-informed approach establishes quantitative structure–transport relationships and offers actionable design principles for engineering high-performance HC anodes.

Data-driven framework↗

Interpretable and flexible non-intrusive reduced-order models using reproducing kernel Hilbert spaces

This paper develops an interpretable, non-intrusive reduced-order modeling technique using regularized kernel interpolation. Existing non-intrusive approaches approximate the dynamics of a reduced-order model (ROM) by solving a data-driven least-squares regression problem for low-dimensional matrix operators. Our approach instead leverages regularized kernel interpolation, which yields an optimal approximation of the ROM dynamics from a user-defined reproducing kernel Hilbert space. We show that our kernel-based approach can produce interpretable ROMs whose structure mirrors full-order model structure by embedding judiciously chosen feature maps into the kernel. The approach is flexible and allows a combination of informed structure through feature maps and closure terms via more general nonlinear terms in the kernel. We also derive a computable a posteriori error bound that combines standard error estimates for intrusive projection-based ROMs and kernel interpolants. In conclusion, the approach is demonstrated in several numerical experiments that include comparisons to operator inference using both proper orthogonal decomposition and quadratic manifold dimension reduction.

Data-driven model reduction↗

Search for single-production of vector-like quarks decaying into Wb in the fully hadronic final state in pp collisions at $\sqrt{s}$ = 13 TeV with the ATLAS detector

A search for T and Y vector-like quarks produced in proton-proton collisions at a centre-of-mass energy of 13 TeV and decaying into Wb in the fully hadronic final state is presented. The search uses 139 fb −1 of data collected by the ATLAS detector at the LHC from 2015 to 2018. The final state is characterised by a hadronically decaying W boson with large Lorentz boost and a b-tagged jet, which are used to reconstruct the invariant mass of the vector-like quark candidate. The main background is QCD multijet production, which is estimated using a data-driven method. Upon finding no significant excess in data, mass limits at 95% confidence level are obtained as a function of the global coupling parameter, κ. The observed lower limits on the masses of Y quarks with κ = 0.5 and κ = 0.7 are 2.0 TeV and 2.4 TeV, respectively. For T quarks, the observed mass limits are 1.4 TeV for κ = 0.5 and 1.9 TeV for κ = 0.7.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

ESnet Data and AI Workshop Report

In February 2025, the DOE user facility Energy Sciences Network (ESnet) held a three-day Data and AI Workshop in Berkeley, California. The objective of the workshop was to identify challenges within ESnet that could be addressed through data-driven methods, to help define ESnet’s data-analysis requirements, and to shape its AI strategy, guiding data-stewardship efforts and the direction of AI research and AIOps exploration for ESnet7, the next iteration of ESnet’s network. This report summarizes the multi-faceted discussions and findings and presents a set of recommendations for next steps.

97 MATHEMATICS AND COMPUTING↗

Addressing Rising Energy Demand Through Innovation

The U.S. is facing a significant increase in energy demand, driven by AI advancements, the rapid expansion of data centers, manufacturing and industrial growth, and the electrification of transportation and buildings. Buildings alone account for approximately 75% of U.S. electricity consumption and 40% of total energy use. To address these challenges, NLR leverages its state-of-the-art research facilities, advanced energy modeling, hardware-in-the-loop emulation, and real-world demonstrations to provide data-driven insights that de-risk emerging energy solutions, increase efficiency and demand flexibility, optimize grid controls, and identify vulnerabilities to enhance energy security. This presentation will highlight our research ecosystem and its role in supporting a more reliable, affordable, and adaptive energy infrastructure in the face of accelerating demand.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

EMPHATIC Silicon Strip Detector Efficiencies

EMPHATIC is an experiment at Fermilab which aims to reduce current neutrino flux uncertainties. This report discusses the limitations current neutrino flux uncertainties places on large scale neutrino experiments, provides background on the EMPHATIC experiment, and details the project of determining the efficiency of the Silicon Strip Detectors (SSDs) used in EMPHATIC. As part of the data analysis process and in order to increase the accuracy of EMPHATIC’s simulations a representation of efficiency of each SSD is required. To achieve this a data-driven analysis was performed on EMPHATIC's collected data using the Root and Art frameworks. Visual and numerical representations of efficiency were determined. The average efficiency over all SSDs is 98.58\%, however this number deflated as it includes known bad channels.

Olson, Virginia [Illinois U., Urbana (main)]↗

Demonstration and performance of an online data selection algorithm for liquid argon time projection chambers using MicroBooNE

The MicroBooNE detector is a liquid argon time projection chamber (LArTPC) that produces three-dimensional images of particle interactions using ionization charge collected by anode wire plane arrays and scintillation light collected by a light detection system. In addition to testing long-standing experimental neutrino anomalies and performing measurements of neutrino interactions with argon nuclei using the Fermilab Booster Neutrino Beam, MicroBooNE aims to develop methodologies for rare beyond the Standard Model and off-beam physics searches. Looking ahead to the upcoming Deep Underground Neutrino Experiment (DUNE), with MicroBooNE serving as a valuable testbed, achieving high sensitivity and livetime for off-beam physics while satisfying data processing and storage constraints will require data-driven, intelligent, and online or real-time data selection techniques. These techniques are essential for reducing data rates and preserving rare signals with high accuracy. In this paper, we describe a fast data selection algorithm suitable for online execution to identify electrons from stopping cosmic ray muons in the MicroBooNE detector utilizing ionization charge information, and present its performance. This represents the first demonstration of online data selection in a LArTPC using real data and charge information exclusively and provides an important proof-of-principle for applying such techniques to other LArTPC experiments such as the Short-Baseline Near Detector and DUNE.

Abratenko, P. [Tufts U. (main)]↗

Determining the Efficiency of EMPHATICs Silicon Strip Detectors (SSDs)

EMPHATIC is an experiment at Fermilab which aims to reduce current neutrino flux uncertainties. This report discusses the limitations current neutrino flux uncertainties places on large scale neutrino experiments, provides background on the EMPHATIC experiment, and details the project of determining the efficiency of the Silicon Strip Detectors (SSDs) used in EMPHATIC. As part of the data analysis process and in order to increase the accuracy of EMPHATIC’s simulations a representation of efficiency of each SSD is required. To achieve this a data-driven analysis was performed on EMPHATIC's collected data using the Root and Art frameworks. Visual and numerical representations of efficiency were determined. The average efficiency over all SSDs is 98.58\%, however this number deflated as it includes known bad channels.

Olson, V. [Illinois U., Urbana (main)]↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗

Securing the Modern Grid: Federal Investments, Digitization, and Supply Chain Strategy

Across the United States (U.S.) grid expansion and modernization is underway, paving the way for accelerated load growth and intelligent resource management. Digitization of the grid is supported by several state and federal programs, providing support for utilities installing advanced metering infrastructure (AMI), AI-powered analytics systems, battery energy storage systems (BESS), and distributed energy resource management systems (DERMS) to transform the grid from a one-way power delivery system into an intelligent, responsive network that will enable faster load growth and power expansion of data centers for advanced artificial intelligence (AI) applications. The digital transformation of America's grid presents opportunity for increased efficiency and resiliency but also introduces new digital risks that require careful management. Digital equipment often contains several vulnerabilities such as unencrypted communication protocols, and persistent remote access capabilities that could be exploited to manipulate device settings, coordinate service disruptions, or inject false data into grid operations. These digital risks become particularly important as the grid must rapidly scale to support AI-driven data centers, which the administration has identified as essential for maintaining U.S. technological leadership and economic competitiveness. These vulnerabilities are compounded by supply chain realities: Chinese manufacturers currently produce 70-90% of essential grid components including inverters, batteries, and control systems, with the U.S. lacking domestic manufacturing capacity for critical assets like extra-high voltage transformers. Recent federal legislation has established Foreign Entity of Concern (FEOC) restrictions to address these risks, requiring projects to achieve escalating thresholds of non-FEOC content to receive tax credits while utilities work to expand sourcing channels for their supply chains and strengthen security measures. These restrictions arrive precisely when utilities face unprecedented electricity demand growth driven by the rapid growth in data centers, creating a considerable challenge: rapidly expanding infrastructure while navigating complex compliance requirements while lacking viable alternatives for many critical components. Idaho National Laboratory (INL) and its partners have developed practical approaches to help utilities navigate these intersecting challenges as they leverage federal investment to strengthen and grow the grid. These solutions include Cyber-Informed Engineering (CIE) principles that build resilience directly into systems, the Cirrus tool for secure cloud migration, and enhanced procurement guidance that embeds security requirements throughout equipment lifecycles. Federal initiatives, such as the Technical Assistance for Digital Assurance (TADA) project, provide direct support to utilities implementing these approaches while facilitating knowledge sharing across the industry. While these tools and frameworks cannot eliminate all risks inherent in foreign supply chain dependencies, they offer pragmatic pathways for strengthening security posture without sacrificing the deployment momentum essential to meeting surging electricity demand. Ultimately, securing America's digital energy infrastructure demands dedicated coordination across multiple fronts: building domestic supply chains, implementing robust digital assurance practices, and maintaining the aggressive modernization timeline necessary for reliability, resilience, and energy independence.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Power generation forecasting for solar plants based on Dynamic Bayesian networks by fusing multi-source information

A Dynamic Bayesian network (DBN) model for solar power generation forecasting in solar plants is proposed in this paper. The key idea is to fuse sensor data, operational indicators, meteorological data, lagged output power information, and model errors for more accurate short-term (e.g., hours) and mid-term (e.g., days to weeks) power generation forecasting. The proposed DBN augments automated data-driven structure learning with expert knowledge encoding using continuous and categorical data given constraints to represent causal relationships within a solar inverter system. Additionally, an error compensation mechanism is proposed to capture temporal fluctuation. The effectiveness of the DBN on solar power generation forecasting was evaluated by rolling window analysis with one-year testing data collected from a local solar plant. The proposed DBN is compared with four state-of-art methods including support-vector regression (SVR), k-nearest neighbors (kNN), artificial neural network (ANN), and long short-term memory (LSTM) models. The result show that the proposed DBN achieves better accuracy in general, and it is not as data-hungry as some neural network-based models. The proposed DBN is also shown to have robust and consistent forecasting power with different forecasting horizons. The accuracy is 92% - 95% from one hour to one week ahead forecasting.

14 SOLAR ENERGY↗

2025 Workshop on Envisioning Frontiers in AI and Computing for Biological Research: Position Papers

This workshop aims to identify key research directions for transforming biology using artificial intelligence (AI), machine learning (ML) and computational methods to facilitate the discovery of new behaviors, mechanisms, and designs of biological processes relevant to DOE missions, underpinning a broader U.S. bioeconomy. By developing novel AI/ML technologies to analyze and interpret complex biological data, researchers can organize and simulate biological processes at various scales as well as advance predictive understanding and manipulation of biological systems. This integration of computation, experimentation, and next-generation experimental technologies can lead to discoveries in new biological behaviors and mechanisms relevant to DOE missions. The focus is on how advanced computational and mathematical methods can impact this mission by exploring digital twins, foundation models, automated laboratory experiments, modeling of complex living systems, and data-driven approaches for the biodesign of plants and microbial systems. While data management is important, it is not the primary focus of this workshop, which will assess the current state, trends, and AI/ML challenges at the interface between biology and computational science to identify opportunities for high-impact research at their intersection. The goal is to define research needs and opportunities that align with biological sciences, computational sciences, and applied mathematics research.

59 BASIC BIOLOGICAL SCIENCES↗

Equation-Free Coarse Control of Distributed Parameter Systems via Local Neural Operators

The control of high-dimensional distributed parameter systems (DPS) remains a challenge when explicit coarse-grained equations are unavailable. Classical equation-free (EF) approaches rely on fine-scale simulators treated as black-box timesteppers. However, repeated simulations for steady-state computation, linearization, and control design are often computationally prohibitive, or the microscopic timestepper may not even be available, leaving us with data as the only resource. We propose a data-driven alternative that uses local neural operators, trained on spatiotemporal microscopic/mesoscopic data, to obtain efficient short-time solution operators. These surrogates are employed within Krylov subspace methods to compute coarse steady and unsteady-states, while also providing Jacobian information in a matrix-free manner. Krylov-Arnoldi iterations then approximate the dominant eigenspectrum, yielding reduced models that capture the open-loop slow dynamics without explicit Jacobian assembly. Both discrete-time Linear Quadratic Regulator (dLQR) and pole-placement (PP) controllers are based on this reduced system and lifted back to the full nonlinear dynamics, thereby closing the feedback loop.

93B52, 93C20, 47N70, 65J15, 65M32, 68T07, 68T20, 6↗

Integrating Data Centers and Grid Technologies at Scale

This presentation focuses on the challenge of integrating AI-driven data centers with the power grid at scale. It examines the AI data center capacity challenge and the role of new Medium Voltage Direct Current (MVDC) and other grid-enhancing technologies in enabling efficient and reliable power delivery. The session will highlight the National Laboratory of the Rockies' ARIES capabilities and planning tools, along with collaborative examples involving Verrus, Compass, and Schneider through the Agora test bed for grid-friendly data center evaluations, and ON. Energy for UPS evaluation. It will showcase the NLR Stable Grid Platform for studying oscillations caused by large-scale data centers, along with planning tools to assess grid security and reliability. Additionally, the presentation covers reconductoring strategies to increase grid capacity and explores innovative data center architectures, including the Advanced DC Architectures with Power-electronic Transformers (ADAPT) platform, which enables testing of complete DC architectures for data centers.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Multi-system analysis of offshore geologic carbon storage: a review of open-source data science solutions

Geologic carbon storage projects are maturing worldwide and the footprint of deployment in the offshore is expanding. At present, there are ten projects in operation or that have been completed, more than 50 in construction and development, and dozens of characterization studies completed or underway. Offshore geologic carbon storage offers potential benefits over onshore geologic carbon storage. These offshore projects are generally remote in location, distant from population centers, and avoid complicated pore space rights while having abundant prospective storage potential. Some offshore fields targeted for carbon storage have comparatively fewer prior borehole penetrations except for areas that have been explored for petroleum production, minimizing potential issues such as pressure interference and infrastructure impacts. Yet offshore geologic carbon storage projects face distinctive technical and economic challenges, such as seafloor geohazards (e.g., seabed instability), expensive maritime transport, and meteorological-oceanographic conditions that can damage infrastructure and impact operations. Analytical capabilities and improved computational speeds have advanced engineering, earth and energy sciences in the wake of the arrival of modern data science over the last decade. These advancements have created an opportunity for integrated, multi-systems modeling approaches utilizing artificial intelligence and machine learning that are no longer limited by computational issues. Analytical tools developed alongside this advancement in data science can be leveraged to calibrate the potential advantages and challenges of carbon storage operations in the offshore. New methods and approaches that incorporate data science to analyze multiple aspects of engineered and natural systems can provide insights that complement the characterization and onsite engineering that traditional commercial and operational software addresses. These new methods and approaches can potentially improve the outcome of energy operations and carbon storage. Providing multi-system, science-driven data analytics enhances the knowledge base that offshore developers, operators, and regulatory bodies may draw from to improve offshore site selection and operational efficiency. Here, we provide a brief synopsis of geologic carbon storage efforts to date, an overview of the engineered and natural systems involved in offshore geologic carbon storage, and a review of publicly available, open-source, offshore and/or carbon storage related data- and science-driven tools developed by 2010 or later that are suitable for screening and assessing regions for offshore geologic carbon storage.

artificial intelligence↗

Latency Analysis of the Nexus Digital Twin Framework

Real-time digital catalogs are increasingly relied upon to track metadata and connect disparate data sources for cloud-based data integration efforts. One such tool, Deeplynx Nexus is supporting real-time digital twin efforts through event-driven data integration and time-series queries. Nexus’s usefulness for these applications depends critically on how quickly individual records can be uploaded and downloaded, since delays directly affect the responsiveness of any system built on top of it. However, the actual latency a user should expect from Nexus has not been systematically measured before, particularly for the small, frequent transactions typical of live sensor feeds. Here we show that single-record round-trip latency is 61.1 ms on a local Nexus instance and 391.7 ms on the hosted production infrastructure, a roughly 6.4x difference driven primarily by fixed per-request overhead rather than data volume. This overhead dominates at small scale: comparing single-record and ten-record trials suggests approximately 56 ms of each single-record request is fixed connection and authentication cost rather than data-transfer time, meaning batching even a handful of records is substantially more efficient than transmitting them individually. At large batch sizes, this pattern reverses for uploads, which converge to near parity between local and hosted environments by 25,000-50,000 records, while download latency remains persistently 5.7-6.4x slower on hosted infrastructure even at scale. These results suggest that Nexus deployments intended for real-time digital twin applications should prioritize record batching over single-record transactions, and that download-path optimization on hosted infrastructure offers the largest remaining opportunity to reduce latency at scale. We anticipate these baseline measurements will serve as a reference point for future digital twin projects evaluating whether Nexus’s latency profile meets their real-time requirements, and as a benchmark for tracking the effect of future infrastructure or API changes.

99 - GENERAL AND MISCELLANEOUS↗

Hyperplane decision trees as piecewise linear surrogate models for chemical process design

Recent trends in chemical engineering research point towards an increasing reliance on data-driven modeling approaches. Neural networks, for instance, have proven to be accurate when data is plentiful and high-dimensional, but in many cases, they require computationally-intensive training procedures. Here, in this work, we describe hyperplane decision trees (HT) as a highly expressive and low-compute machine learning model architecture. These models are locally linear and have linear decision boundaries, resulting in a piecewise linear model of the data. This property allows them to be converted into mixed-integer linear constraints which can be globally optimized. Our open-source PyTorch implementation of this method is a fast, flexible, and accessible way to build accurate piecewise linear models of data.

Decision trees↗

Digital Twins for Data Centers

Fueled by an unprecedented adoption of AI (Artificial Intelligence), data centers are becoming the largest growing consumers of energy. Digital Twins provide living digital models of physical systems that enable data-driven analysis and application of AI to better manage selective aspects of the data center and drive efficiency for sustainability. Digital twins have emerged as a way to create virtual prototypes of physical artifacts, which may be used in a variety of contexts. Physical artifacts include airplanes, factories, or even static objects, such as bridges or dams. Digital twin helps monitor changes and assist in predicting planned or unplanned behaviors of physical objects. In this paper, we discuss digital twins for data centers.

97 MATHEMATICS AND COMPUTING↗