Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Expanding the Domain of Applicability of Machine Learning Models with Limited Data for Drug Property Prediction

Accurate machine learning models for predicting small molecule interactions with biological targets are essential for therapeutic discovery, biothreat response, and computational drug design, but their performance is often limited for understudied targets with sparse experimental data. To address this challenge, we developed and evaluated methods to improve molecular property prediction under low-data conditions, using the NimA-related kinase (NEK) family as a proof-of-concept. This work focused on two complementary goals within the ATOM Modeling PipeLine (AMPL) and the Generative Molecular Design (GMD) loop: expanding model applicability through transfer learning, representation learning, feature scaling, sampling strategies, and active-learning-inspired compound selection; and enabling efficient virtual screening to prioritize compounds that balance predicted activity, design objectives, and synthetic accessibility.

organic↗

Workbooks for Cambium 2024 Data

These workbooks contain a subset of cost and emissions data from the 2024 Cambium datasets, with levelization calculations to assist users in translating Cambium’s year-over-year values to a representative value for a user-specified project timeline. These workbooks provide modeled data for 18 GEA regions covering the contiguous United States, projected forward through 2050. Mappings of these regions to ZIP codes and counties is given within this workbook in the corresponding tabs. For the full Cambium 2024 data sets, see the Cambium 2024 project on NREL's Scenario Viewer. For more details on input assumptions and methodology see the associated report: Cambium 2024 Scenario Descriptions and Documentation. Users are advised to review section 4 of the report, which discusses limitations and caveats of the data. This data is planned to be updated annually. Information on the latest versions can be found on the NREL energy analysis page on Cambium.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

A Perspective on Traditional and Data Driven Electrochemical Modeling and Analysis

To understand the behavior of electrochemical systems, we need to reduce the dimensionality of the measured current-voltage-time (I-V-t) data by fitting models, thus enabling us to analyze and compare the governing physics. Traditionally, the process for this is an 'expert first' approach: defining the model and its explicit assumptions based on inductive reasoning or empirical observation, fitting small portions of the I-V-t data where assumptions are most valid or carefully designing experiments to enforce key assumptions, and then interpreting the model parameters. However, modern data-driven methods enable a new paradigm: a 'data first' approach, where the latent behaviors governing the system's measured response are identified directly using machine-learning models that optimize both model structure and parameters from the I-V-t data, guaranteeing that the learned model explains as much of the observed system response as possible. After model identification, the model can then be interrogated by an expert to connect observed behaviors with underlying physics. This talk will review several different types of electrochemical analysis (electrochemical impedance, differential voltage-capacity, electrochemical kinetics) and compare the traditional and data-driven methods for analyzing the data.

42 ENGINEERING↗

Illinois Storage Corridor, CarbonSAFE Phase III: UIC Class VI Permitting Plan

The Illinois Storage Corridor (ISC) project evaluated two distinct sites to determine the feasibility of commercial-scale CO₂ storage at each. The project leveraged the region's exceptional geological characteristics, particularly the well-characterized Cambro-Ordovician Storage Complex, to enable permanent geological storage of more than 50 million tonnes of CO₂ over 30 years. The two storage sites are located at One Earth Energy (OEE) facility in northcentral Illinois and Prairie State Generating Company (PSGC) in southcentral Illinois. Once operational, these facilities will combine to capture and store more than 6.5 million tonnes of CO₂ per year, positioning the ISC among the largest carbon storage regions globally. Preliminary homogeneous dynamic modeling based on regional and site-specific reservoir characteristics indicates promising injection capabilities at both locations. For the OEE site, modeling predicts a maximum allowable injection rate of 3.7 MTPA, with a baseline scenario of 1.7 MTPA over 30 years producing a CO₂ plume radius of 1.4 miles at end of injection. For the PSGC site, incorporating recent well data, modeling indicates a single-well maximum injection rate of 2.1 MTPA, with a plume radius of 4.2 miles at end of injection for the 60 MT over 30 years scenario. The primary objective of this CarbonSAFE Phase III project is to develop and submit Class VI Underground Injection Control (UIC) permit applications to the U.S. Environmental Protection Agency Region 5. The permitting plan outlines comprehensive site characterization, Area of Review delineation, monitoring programs, well construction designs, financial responsibility provisions, and post-injection site care procedures necessary to demonstrate safe, permanent CO₂ storage protective of underground sources of drinking water. Three UIC Class VI permit applications for the OEE site were submitted to EPA in October 2022 and are progressing through technical review, with final permit decision projected by June 2026. Through five rounds of Requests for Additional Information and responses, the applications have been refined to address computational modeling, area of review delineation, well integrity, monitoring protocols, and financial assurance requirements. For the PSGC site, finalized characterization and permitting documentation was delivered directly to the facility in July 2023 due to business constraints precluding formal federal regulatory submission. This comprehensive permitting effort builds upon extensive prior subsurface evaluations and demonstration projects that have confirmed the feasibility of widespread commercial-scale carbon storage in the region.

01 COAL, LIGNITE, AND PEAT↗

Atmospheric Modeling and Denoising for Millimeter-Wave Line Intensity Mapping

Line-intensity mapping (LIM) offers a promising approach to mapping large-scale cosmic structure, and the greatest obstacle for ground-based observations at millimeter wavelengths is foreground contamination from atmospheric emission. In this work, we present a simulation and denoising framework designed to isolate and subtract atmospheric fluctuations from LIM data, modeled after the instrument parameters of the South Pole Telescope Summertime Line Intensity Mapper (SPT-SLIM). We generate mock observations spanning 125-175 GHz containing cosmic signals, precipitable water vapor screens, ice crystal fluctuations, and photon noise. We then implement a spatial-spectral atmospheric removal pipeline combining per-pixel linear template regression with a two-dimensional Fourier-domain filter. The framework is evaluated under simulated conditions in the South Pole and the Atacama Desert across three key metrics: cosmic signal preservation, foreground subtraction efficiency, and instrument noise injection. Our pipeline achieves atmospheric suppression at large spatial scales, and these results establish a physically grounded foundation for atmosphere removal in ground-based LIM data collection.

Saye, Laney [UC, Berkeley (main)] (ORCID:000900078↗

Atmospheric Modeling and Denoising for Millimeter-Wave Line Intensity Mapping

Line-intensity mapping (LIM) offers a promising approach to mapping large-scale cosmic structure, and the greatest obstacle for ground-based observations at millimeter wavelengths is foreground contamination from atmospheric emission. In this work, we present a simulation and denoising framework designed to isolate and subtract atmospheric fluctuations from LIM data, modeled after the instrument parameters of the South Pole Telescope Summertime Line Intensity Mapper (SPT-SLIM). We generate mock observations spanning 125-175 GHz containing cosmic signals, precipitable water vapor screens, ice crystal fluctuations, and photon noise. We then implement a spatial-spectral atmospheric removal pipeline combining per-pixel linear template regression with a two-dimensional Fourier-domain filter. The framework is evaluated under simulated conditions in the South Pole and the Atacama Desert across three key metrics: cosmic signal preservation, foreground subtraction efficiency, and instrument noise injection. Our pipeline achieves atmospheric suppression at large spatial scales, and these results establish a physically grounded foundation for atmosphere removal in ground-based LIM data collection.

Saye, L. K. [UC, Berkeley (main)] (ORCID:000900078↗

HPC-FAIR: A Framework Managing Data and AI Models for Analyzing and Optimizing Scientific Applications

The increasing reliance on machine learning (ML) to analyze and optimize large-scale scientific applications on supercomputers faces a significant bottleneck: the lack of readily available, high-quality training datasets and the difficulty in reusing existing AI models. This project was motivated by the urgent need to address the “FAIR” principles (Findability, Accessibility, Interoperability, Reusability) for both training datasets and AI models in the high-performance computing (HPC) domain. The project developed HPC-FAIR, a high-performance computing data management framework designed to centralize HPC-related datasets and AI models within a unified hub. To ensure interoperability, the framework established a standardized representation and vocabulary (ontology) for both data and models. HPC-FAIR also implemented automated workflows to streamline data processing, model access, and benchmarking. Additionally, the project focused on optimizing data harnessing efficiency through advanced techniques like deep reuse and compression-based analytics.

97 MATHEMATICS AND COMPUTING↗

Bim-to-fea Conversion Program

The purpose of this program is to enable interoperability between BIM-based architectural design software (i.e., Revit, ArchiCAD, AVEVA E3D) to structural analysis software (i.e., SAP2000). The program takes in BIM building model data via the IFC file format, automatically transforms the architectural coordination entities (structural beams, columns, slabs, walls) to structural analysis entities (i.e., finite element space frames and shells), automatically adjusts the connectivity of the structural analysis entities, and finally exports the structural analysis entities as a structural analysis model contained within a new IFC file. For example, a 3D building in Revit can be exported to an IFC file, run through this BIM-to-FEA program, then the exported IFC can be inputted into SAP2000.

Crowder, Nicholas [Idaho National Laboratory (INL)↗

A multi-algorithm approach for modeling coastal wetland eco-geomorphology

Coastal wetlands play an important role in the global water and biogeochemical cycles. Climate change makes it more difficult for these ecosystems to adapt to the fluctuation in sea levels and other environmental changes. Given the importance of eco-geomorphological processes for coastal wetland resilience, many eco-geomorphology models differing in complexity and numerical schemes have been developed in recent decades. However, their divergent estimates of the response of coastal wetlands to climate change indicate that substantial structural uncertainties exist in these models. To investigate the structural uncertainty of coastal wetland eco-geomorphology models, we developed a multi-algorithm model framework of eco-geomorphological processes, such as mineral accretion and organic matter accretion, within a single hydrodynamics model. The framework is designed to explore possible ways to represent coastal wetland eco-geomorphology in Earth system models and reduce the related uncertainties in global applications. We tested this model framework at three representative coastal wetland sites: two saltmarsh wetlands (Venice Lagoon and Plum Island Estuary) and a mangrove wetland (Hunter Estuary). Through the model–data comparison, we showed the importance of using a multi-algorithm ensemble approach for more robust predictions of the evolution of coastal wetlands. We also found that more observations of mineral and organic matter accretion at different elevations of coastal wetlands and evaluation of the coastal wetland models at different sites in diverse environments can help reduce the model uncertainty.

58 GEOSCIENCES↗

Introduction to Engage: NASA Training Session

Welcome to Engage! Engage is a capacity expansion modeling tool supported by the National Renewable Energy Laboratory and based on the Calliope open-source capacity expansion model developed by the ETH Zurich University, maintained at the TU Delft University. Engage is an accessible (free, open-access, web-hosted) and flexible web-based energy system planning application for rapid multiple-energy-form energy system scenario exploration. Its cloud-based, collaborator-sharable data model, intuitive interface and visualization capabilities facilitate collaboration and communication among teams, with experts, and among diverse stakeholder groups exploring energy system implications from district to national-scale models. This training session was presented to the National Aeronautics and Space Administration (NASA) to help them understand how capacity expansion modeling can help them develop single site/distribution analysis of energy to regional airports.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Intersection of Hydrologic Change and Hydropower in the United States: Needs for Future Research and Practice

Hydropower is crucial for electric‐grid stability in the context of variable renewables but faces threats from changing hydrology. Here, we summarize the state of the science at the intersection of hydropower operations and planning, hydrologic science, and climate. We focus on the United States, outlining research, development, and training needs. Key knowledge gaps include the risk that intensification of compound extreme events poses to future generation, as well as uncertainties surrounding greenhouse gas emissions from hydropower reservoirs with relevance to hydropower's role in energy decarbonization. Quantifying such impacts and reducing uncertainty are critical where possible, but remaining irreducible or deep uncertainty will require new approaches. Future monitoring and modeling methods must provide a better understanding of the complexity inherent in large watersheds that is critical to managing both hydropower and watersheds in the context of hydrologic change. Yet, research and development will have little impact if they do not inform practice. Standardization and consolidation of platforms are essential for data, modeling, and tool translation to local scales and small operators. An enhanced industry‐academia dialog is pivotal for fostering a robust pipeline of hydropower professionals. Collaboration among researchers, policymakers, authorities, and industry stakeholders emerges as a recurring theme, highlighting the imperative for collective efforts.

13 HYDRO ENERGY↗

DOE FAIR Surrogate Benchmarks Supporting AI and Simulation Research (SBI Surrogate Benchmark Initiative) (Final Report)

Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia(UVA). SBI repositories include data, code, and all relevant collateral artifacts, that the science and engineering community needs to use and reuse these data sets and surrogates. SBI repositories generate active research from both participants in SBI and the broader AI and domain science communities. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and capture them as surrogate benchmarks with a rich set of metadata, covering. Data; Model; Metrics specification; Machine specification; Science, Speed, Power Results, We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non Surrogate benchmarks that have many common features and similar issues regarding FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, benchmarks have datasets, models, and metadata, and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates, including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).

97 MATHEMATICS AND COMPUTING↗

FAIR Surrogate Benchmarks Supporting AI and Simulation Research (Final Report)

Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia (UVA). SBI repositories include data, code, and all relevant collateral artifacts that the science and engineering community need to use and reuse these data sets and surrogates. SBI repositories generate active research from both the participants in SBI and the broad community of AI and domain scientists. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and captures them as surrogate benchmarks with a rich set of metadata covering: Data; Model; Metrics specification; Machine specification; and Science, Speed, and Power Results. We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non-Surrogate benchmarks that have many common features and similar issues as regards FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, Benchmarks have datasets, models, and metadata and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).

97 MATHEMATICS AND COMPUTING↗

A Multi-Model, Multi-Scale Research Program in Stressors, Responses, and Coupled Systems Dynamics at the Energy-Water-Land Nexus and for Concentrated, Interdependent Infrastructures: Toward Next Generation Capabilities in Integrated Impacts, Adaptation, and Vulnerability (I-IAV) Modeling and a Community of Practice

The goal of this research program was to build a next generation integrated suite of science-driven modeling and analytic capabilities, and a more expanded and connected community of practice, for analyses of the stressors, impacts, adaptations and vulnerabilities of global and regional change. The emphasis was on understanding energy-water-land interactions and feedbacks and interdependent infrastructures at appropriate regional and temporal scales. Although the scope spans many complex facets of data, modeling, and analysis, as well as scales appropriate for integrated impacts and adaptation research, the focus of this effort was the development of multi-model, multi-scale capabilities spanning the domains of Multi-Sector Dynamics (MSD) models; Impact, Adaptation, and Vulnerability (IAV) models; and Earth System Models (ESMs).

54 ENVIRONMENTAL SCIENCES↗

Aspen Open Jets: unlocking LHC data for foundation models in particle physics

Foundation models are deep learning models pre-trained on large amounts of data which are capable of generalizing to multiple datasets and/or downstream tasks. This work demonstrates how data collected by the CMS experiment at the Large Hadron Collider can be useful in pre-training foundation models for HEP. Specifically, we introduce the AspenOpenJets (AOJs) dataset, consisting of approximately 178 M high p T jets derived from CMS 2016 Open Data. We show how pre-training the OmniJet-α foundation model on AOJs improves performance on generative tasks with significant domain shift: generating boosted top and QCD jets from the simulated JetClass dataset. In addition to demonstrating the power of pre-training of a jet-based foundation model on actual proton–proton collision data, we provide the ML-ready derived AOJs dataset for further public use.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Bayesian Optimization of Non-Invariant Systems with Constraints Developed for Application to the ECR Ion Source VENUS

In this work, we consider the optimization of non-invariant systems with both safety and control constraints. We present a new approach based on Bayesian optimization for the dynamic, safe and controlled optimization of such systems. Although there are other possible use cases, we focus on the application to the electron cyclotron resonance ion source VENUS. From experimental data, we have observed that VENUS behaves to first order as a non-invariant dynamic system with moving areas of instability. Our novel approach aims at providing a tool that can maintain system optimization in a safe way. This is accomplished by making sure the objective function, the beam current in the case of VENUS, does not fall under an operational minimum, while simultaneously requiring the optimization to avoid areas where VENUS is unstable. We compare the result of our approach on synthetic data modeled to mimic the behavior of VENUS with two methods from the literature, a standard Bayesian optimizer and a safe Bayesian optimizer, both adapted to deal with dynamic systems. A cross Student T-test is conducted to show the significance of the improvement given by the new method we introduce here, regarding the two preexisting methods we compared to. The results of the tests conducted on synthetic data show that the proposed method succeeds at maintaining the system optimized and obeys the predefined constraints better than the literature methods explored.

Bayesian optimization↗

sup3ruhi (Super Resolution for Renewable Resource Data and Urban Heat Islands) [SWR-25-05]

Urban heat is a growing concern, particularly in dense metropolitan areas where high temperatures increase the risk of heat-related illness and drive energy expenses for cooling. Estimating the effects of urban heat remains a challenge due to limitations in describing the built environment, computational constraints, and the need for high-resolution data. This software presents open-source, computationally efficient machine learning methods that enhance the accuracy of urban temperature estimates compared to historical reanalysis data. Models trained using this software have been applied to urban microclimates in Los Angeles and Seattle showing greater accuracy and less bias when compared to low-resolution reanalysis datasets like ERA5 and even when compared to high-resolution mesoscale numerical weather models like WRF with an urban canopy model. Initial findings highlight how machine learning can support urban heat resilience planning by enabling improved assessments of local heat islands, mitigation strategies, and their energy implications. This software is an extension of (sup3r). This software supports the following publication: Buster, Grant, et al. Tackling Extreme Urban Heat: A Machine Learning Approach to Assess the Impacts of Climate Change and the Efficacy of Climate Adaptation Strategies in Urban Microclimates. arXiv:2411.05952, arXiv, 8 Nov. 2024. arXiv.org, https://doi.org/10.48550/arXiv.2411.05952. And has related public data records available at: Buster, Grant, Cox, Jordan, Benton, Brandon, and King, Ryan. Super-Resolution for Renewable Resource Data and Urban Heat Islands (Sup3rUHI). United States: N.p., 16 Oct, 2024. Web. https://data.openei.org/submissions/6220.

Buster, Grant [National Renewable Energy Laborator↗

Leveraging large language models to address data scarcity in machine learning for graphene synthesis

Machine learning in experimental materials science faces significant challenges due to the scarcity of data, which are costly and time-consuming to generate, particularly when relying on in-house experiments. Literature data mining offers a potential solution but introduces issues like mixed data quality, inconsistent formats, and non-uniform reporting of synthesis parameters, resulting in partially missing and heterogeneous features across the dataset. Here, we propose data imputation and feature engineering methods that employ pre-trained large language models (LLMs) to enhance machine learning performance on scarce, heterogeneous datasets, demonstrated on graphene CVD synthesis data and the ML-HydPARK hydrogen storage dataset. GPT models perform data imputation via tailored prompting and semantic normalization of inconsistently reported features through embeddings, for example, to harmonize the complex nomenclature of CVD substrates. Beyond yielding more diverse and richer feature representations than traditional methods such as K-nearest neighbors (KNN) and Multivariate Imputation by Chained Equations (MICE), LLM-based data imputation is evaluated against dataset characteristics and prompting strategies. We vary the level of autonomy granted to the LLM, from generic prompting that leverages pre-trained knowledge for autonomous data generation to data-informed prompting that constrains outputs using target-specific information, and demonstrate which level of autonomy yields superior imputation performance across datasets and feature types. The proposed data engineering methods markedly improve downstream performance; for example, in graphene layer number classification using a support vector machine (SVM), binary accuracy increases from 39% to 65% and ternary accuracy from 52% to 72%. Fine-tuning experiments on both datasets show that combining our proposed LLM-based data imputation and feature encoding methods with numerical machine learning predictors outperforms standalone fine-tuned LLM predictors in data-scarce settings. The proposed strategies emphasize data enhancement techniques rather than refining learning architectures or regularizing loss functions, offering a broadly applicable framework for improving machine learning performance on scarce, inhomogeneous datasets.

Chemical vapor deposition↗