Search NASA⌕ Search

SEARCH · Search NASA

Results for “data platform”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

R&D GREET Battery Carbon Footprint Calculator

The Battery Carbon Footprint (CF) Calculator was developed to help U.S. battery manufacturers meet the carbon footprint reporting requirements of the EU Battery Regulation (EU) 2023/1542. The calculator incorporates several major battery carbon footprint frameworks, including the Joint Research Centre's Rules for the Calculation of the Carbon Footprint of Electric Vehicle Batteries (CFB-EV), RECHARGE's Product Environmental Footprint Category Rules for High Specific Energy Rechargeable Batteries for Mobile Applications (PEFCR), the Catena-X Product Carbon Footprint Rulebook (CX-PCF Rules), Battery Pass's Battery Carbon Footprint: Rules for Calculating the Carbon Footprint of the "Distribution" and "End-of-Life and Recycling" Life Cycle Stages, the Global Battery Alliance's Greenhouse Gas Rulebook: Generic Rules, Version 2.1, and the Ministry of Economy, Trade and Industry's draft Carbon Footprint Calculation Method for Automotive Batteries. The tool pairs these frameworks with foreground data from Argonne's R&D GREET models and integrates user-supplied background data covering battery manufacturing and supply chain activities. By bringing multiple international methodologies together in a single platform, the calculator enables manufacturers to evaluate product carbon footprints, improve data consistency, and prepare for evolving regulatory compliance and global market reporting requirements.

Zhang, Jingyi↗

Initial Characterization of the NREL Large-Amplitude Motion Platform

The Large Amplitude Motion Platform (LAMP) at NREL represents a significant advancement in the controlled testing of Wave Energy Converters (WECs) under laboratory conditions. Originally designed by E2M as a six-degree-of-freedom (DOF) Stewart platform for flight simulation, LAMP has been adapted by NREL to facilitate the mounting and evaluation of WECs. This adaptation enables dry testing of WECs using motion profiles similar to the ocean, facilitating the iterative design, testing, and validation of WEC performance prior to ocean deployments. This report presents the initial work completed to characterize LAMP, with particular emphasis on its stability and operational capabilities across various single and multi-degree-of-freedom (DOF) motion profiles. The report includes planned comparisons at three distinct mass payloads, aimed at assessing the platform's positional accuracy, frequency response, and endurance over extended runtime periods. These experimental tests are critical for establishing the platform's limitations and ensuring that the data generated during WEC validation is both accurate and reproducible. The outcomes of this study not only contribute to a deeper understanding of LAMP's capabilities but also lay the groundwork for future advancements in WEC testing methodologies. By providing robust and reliable performance data within a controlled laboratory setting, the findings are expected to significantly enhance the development and commercialization of marine energy technologies. This report presents the initial findings of the LAMP Characterization work and proposed steps to further understand and characterize LAMP. Data collected during this work can be found on MHKDR at: https://mhkdr.openei.org/submissions/602 Note that this report shares the measured/found instantaneous maximum operating range of LAMP. For most applications the maximum operating range cannot be used for system health and longevity. The operating range and capabilities of LAMP will be evaluated on a case-by-case basis, single-DOF position, velocity, and acceleration values presented in Table 4a-c, and Table 5 should be taken as instantaneous absolute maximum values. Future use of LAMP will likely be limited to smaller values.

16 TIDAL AND WAVE POWER↗

Automated Label‐Free Assay for Viral Detection and Inhibitor Screening via Biomembrane‐Functionalized Microelectrode Arrays

Most virus infection assays have indirect readout such as virus number following entry (e.g., PCR, cell lysis). While effective, these technologies are labor‐intensive, require specialized environments (e.g., sterile or RNA‐free), and detect later‐stage viral events like lysis or cell death, lacking sensitivity to early fusion events. To address these limitations, we present biologically relevant 2D membrane materials, host‐cell‐derived supported lipid bilayers (hcd‐SLBs), integrated with organic microelectrode arrays (OMEAs) for detection of severe acute respiratory syndrome coronavirus 2 (SARS‐CoV‐2) fusion. By overexpressing angiotensin‐converting enzyme 2 (ACE2) receptors on the native membranes, the platform functions as a viral sensor capable of detecting virus pseudo particles (VPPs) through the late pathway. Additionally, hcd‐SLBs extracted from human lung epithelium expressing native ACE2 detect fusion events through the early pathway. The platform's utility as a drug‐screening tool is demonstrated by testing antibodies targeting either the ACE2 on the host membrane or the viral spike (S) proteins. To enhance the throughput, microfluidics are integrated for automation and OMEAs are incorporated within each channel, miniaturizing the testing units. This system supports high‐throughput data generation, automation, and scalability, providing an efficient platform for viral fusion detection that advances the study of pathogen‐host interactions and accelerates antiviral drug discovery.

Biology↗

CHESS 2025: Discrete-return LiDAR point clouds from NEON AOP surveys

This dataset provides Level 1 (L1) discrete-return light detection and ranging (LiDAR) point cloud data collected for the 2025 Colorado Headwaters Ecological Spectroscopy Study (CHESS). These data were acquired to enable characterization of vegetation structure and other three-dimensional features of the land surface, and to evaluate structural changes that may have occurred between a prior LiDAR acquisition in 2018 and the 2025 overflight. The data were acquired over three study domains in the Upper Gunnison river basin: the upper East River watershed (CRBU); Almont Triangle and Taylor Canyon (ALMO); and Upper Taylor River watershed (UPTA) between 2025-06-13 and 2025-07-15. LiDAR data were acquired using the Optech Galaxy Prime Airborne LiDAR Terrain Mapper onboard the National Ecological Observatory Network (NEON) Airborne Observation Platform (AOP). These are the primary unclassified discrete-return LiDAR data delivered by NEON and are provided per flightline as LASzip (LAZ) 1.4 Format 6 files. Data were processed following the workflow described in the NEON L0-to-L1 Discrete Return LiDAR Algorithm Theoretical Basis Document (Krause and Goulden 2022). Each record in the unclassified point clouds represents a geolocated laser target/return recorded by the LiDAR system, with values for X, Y, Z position and return intensity. All point coordinates are provided in meters. Horizontal coordinates are referenced in Universal Transverse Mercator (UTM) zone 13N and the World Geodetic System (WGS) 1984 ensemble datum. Elevations are referenced to Geoid12A. Flight metadata describing flightline boundaries and positional uncertainty by point are also included. CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgement: Field and remote-sensing data acquisition was performed under a grant from the National Aeronautics and Space Administration (80NSSC24K1005). This work was also supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS Campaigns↗

Data for Discovery, Characterization, and Application of Chromosomal Integration Sites for Stable Heterologous Gene Expression in Rhodotorula toruloides

Rhodotorula toruloides is a non-model, oleaginous yeast uniquely suited to produce acetyl-CoA-derived chemicals. However, the lack of well-characterized genomic integration sites has impeded the metabolic engineering of this organism. Here we report a set of computationally predicted and experimentally validated chromosomal integration sites in R. toruloides . We first implemented an in silico platform by integrating essential gene information and transcriptomic data to identify candidate sites that meet stringent criteria. We then conducted a full experimental characterization of these sites, assessing integration efficiency, gene expression levels, impact on cell growth, and long-term expression stability. Among the identified sites, 12 exhibited integration efficiencies of 50% or higher, making them sufficient for most metabolic engineering applications. Using selected high-efficiency sites, we achieved simultaneous double and triple integrations and efficiently integrated long functional pathways (up to 14.7 kb). Additionally, we developed a new inducible marker recycling system that allows multiple rounds of integration at our characterized sites. We validated this system by performing five sequential rounds of GFP integration and three sequential rounds of MaFAR integration for fatty alcohol production, demonstrating, for the first time, precise gene copy number tuning in R. toruloides . These characterized integration sites should significantly advance metabolic engineering efforts and future genetic tool development in R. toruloides .

Conversion↗

Multi-Scale Integrated Monitoring System for Enhancing Methane Emission Detection, Quantification & Prediction

This report details the progress and findings of a comprehensive study on reviewing existing solutions, identifying technology gaps, and formulating an “all-in-one” integrated strategy for developing the next-generation multiscale methane monitoring and modeling platform, conducted under grant number DE-FE0032292. Co-led by Dr. David Ebert, Dr. Binbin Weng, and Dr. Chenghao Wang at the University of Oklahoma, the project’s goal was to develop an integrated approach for building this engineering platform to detect, quantify, and mitigate methane emissions across various temporal scale, spatial scales, and sectors. The planning grant study began with an extensive review of various methane sensing and monitoring technologies and systems, surveying over 100 technology providers globally. This review revealed the prevalence of optical methods over chemical methods in commercially available sensors, with Non-Dispersive Infrared (NDIR), Tunable Diode Laser Absorption Spectroscopy (TDLAS), and Optical Gas Imaging (OGI) cameras being the most prevalent options. A trend towards more advanced optical techniques was observed, driven by increased regulatory focus and technological advancements. The technical evaluation of these sensing technologies provided crucial insights into their capabilities and limitations. The study examined emerging technologies such as Differential Absorption LiDAR (DIAL), which show promise for high-precision and long-range detection. The team then investigated the features and application bandwidth of various sensing platforms, including handheld, fixed/stationary, mobile, aerials, and spaceborne monitors. Pilot field studies were conducted to assess the capabilities of solutions for different emission scenarios. Field work with sensor deployments was conducted at three distinct site types: an oil & gas industry site, a cattle ranching operation, and a waste processing facility. The team also conducted a thorough review of methane flux inverse modeling approaches, focused on physically based methods. These approaches were categorized into simple, intermediate, and advanced methods. A realtime WRF-GHG (Weather Research and Forecasting-Greenhouse Gas) modeling system was developed and applied, incorporating multiple data sources to guide field experiments and inform methane plume detection. The project identified and analyzed numerous categories of methane data sources, including satellite measurements, ground-based sensors, and inventory databases. Key platforms examined include EDGAR, EPA GHGI, NASA TROPOMI, Carbon Mapper, and Climate TRACE, among others. The team proposed an architecture for a comprehensive methane monitoring platform. This system incorporates multi-source data acquisition, advanced data processing and assimilation, interactive visualization tools, and analytical capabilities for emissions forecasting and scenario analysis. The proposed platform aims to provide a user-friendly interface catering to various stakeholders, from researchers to policymakers. The architecture includes sophisticated data ingestion methods, a centralized data warehouse, and advanced analytical tools for data fusion and interpretation. To ensure the relevance and effectiveness of the proposed system, a comprehensive survey was conducted to gather stakeholder input on system requirements. Key findings include a strong need for integrating various data types and formats, a preference for real-time data updates and advanced visualization tools, and a demand for user-friendly interfaces catering to different expertise levels.

03 NATURAL GAS↗

Predicting Pulsed-Laser Deposition SrTiO 3 Homoepitaxy Growth Dynamics Using High-Speed Reflection High-Energy Electron Diffraction

Pulsed-laser deposition (PLD) is a powerful technique for growing complex oxides with controlled stoichiometry. To understand growth dynamics therein, it is common to leverage in situ spectroscopies, such as reflection high-energy electron diffraction (RHEED), to monitor surface crystallinity. Most commercial systems rely on video-rate cameras operating at 60-120 Hz that lack sufficient temporal resolution to capture growth dynamics at practical deposition frequencies. Here, a high-speed platform to record in situ dynamics via RHEED at >500 Hz is implemented. An open-source analysis package is designed to fit diffraction spots to 2D Gaussians, allowing single-pulse surface reconstruction kinetics extraction. Using homoepitaxially deposited (001)-oriented SrTiO 3 as a model system, we demonstrate how high-speed RHEED can provide real-time insight into growth processes obscured by slower acquisition systems. By fitting the single-pulse intensity to a set of exponential functions, we observe changes in the characteristic decay time and mechanism correlated to the substrate step width and surface termination. We observe distinct surface effects, with diffraction intensity decaying on lower-energy TiO 2 -terminated surfaces and stabilizing on SrO- or mixed-terminated surfaces. Similarly, using an exponential model, the extracted characteristic time of adatom deposition decreases with increased density of bonding sites associated with mixed termination and narrower step widths. Ultimately, this work shows how increasing RHEED temporal resolution can uncover new insights into growth processes, with practical implications for the design and control of PLD processes. This experimental platform provides new capabilities to enable data-driven machine learning analysis and autonomous control systems to enhance the complexity and fecundity of PLD.

(SrO)↗

A parcel-level evaluation of distributed wind opportunity in the contiguous United States

This study examines the potential for distributed wind (DW) energy across the contiguous United States, leveraging advancements in the National Renewable Energy Laboratory's distributed wind model, dWind. The novel modeling approach described here utilizes a high-resolution dataset and analyzes over 150 million parcels, a significant improvement from prior methods that extrapolated results from a smaller random sample. This achievement is enabled through key model performance improvements, such as transitioning to multiprocessing, which reduces runtime by 97 %. This optimized, high-resolution approach allows the inspection of technology deployment potential and impact on a variety of scales tailored to individual properties and regions. The results here align with prior work showing substantial opportunity for energy generation using DW technologies. Key findings reveal a substantial increase from prior results in estimated technical and economic potential for DW. Metrics tuned to highlight economic potential also show increased incentives supporting rural adoption. Results are spatially aggregated for usability and published via the U.S. Department of Energy Wind Data Portal and a custom scenario visualization platform, aiding policymakers, industry, and property owners in assessing DW viability across various scenarios and spatial scales.

17 WIND ENERGY↗

Development of the ARM Lagrangian Large-Scale Forcing Data (ARMLAGTRAJ) Value-Added Product Based on the lagtraj Framework

The Atmospheric Radiation Measurement (ARM) large-scale forcing data developed based on the constrained variational analysis (VARANAL) value-added product (VAP) (Zhang and Lin 1997, Zhang et al. 2001, Xie et al. 2004, Tang et al. 2019) has been widely used for single-column models (SCMs), cloud-resolving models (CRMs), and large-eddy simulation models (LESs) to understand and improve physical processes in models. Recently, the U.S. Department of Energy (DOE) ARM user facility conducted several major field campaigns using ship-based moving observational platforms. For example, the Marine ARM GPCI Investigation of Clouds (MAGIC) field campaign focused on the role of subtropical marine-boundary layer (MBL) clouds, and the Multidisciplinary Drifting Observatory for the Study of Arctic Climate (MOSAiC) field campaign aimed to improve understanding of the coupled climate systems in the Arctic. Observations from moving platforms are critical to provide a comprehensive characterization of coupled-system processes associated with all stages of the cloud and/or sea-ice life cycle. Traditional ARM large-scale forcing data have been developed at fixed locations. They need to be extended to include these moving platforms to address data needs for ship-based field campaigns or to support LES modeling in a Lagrangian framework. With these considerations in mind, we develop ARM-type Lagrangian large-scale forcing data sets based on the lagtraj framework (Boeing et al. 2020) with notable enhancements in generating forcings that are more suitable for ARM field campaigns. The lagtraj is a novel tool that generates forcings for LES and SCM simulation in both Lagrangian and Eulerian perspective. This technical report focuses on the major changes we performed on the lagtraj algorithm and provides an overview of the ARM Lagrangian Large-Scale Forcing Data (ARMLAGTRAJ) value-added products.

54 ENVIRONMENTAL SCIENCES↗

Towards a Public Event Display for DUNE

The Deep Underground Neutrino Experiment (DUNE) is a next generation long baseline neutrino experiment based at Fermilab, with a near detector near the beam target and a Far Detector (FD) in South Dakota. As the experiment prepares for its first data runs, creating pathways for public engagement and data transparency is essential. We present the first-ever DUNE event display designed for public outreach and education. Developed using data from the ProtoDUNE detectors at the CERN Neutrino Platform, this tool provides an intuitive and interactive interface that allows non-experts to visualise and explore particle interactions in a Liquid Argon Time Projection Chamber (LArTPC). By translating raw experimental data into a browser-accessible format, we establish the essential infrastructure for DUNE’s pathway to open data. This talk will detail the technical development of the display, its current implementation with ProtoDUNE data, and the strategic roadmap for integrating it into DUNE’s long-term open-access framework.

Sabater, Eva [U. Sussex (main)] (ORCID:00090001748↗

Rapid Detection and Quick Characterization of African Swine Fever Virus Using the VolTRAX Automated Library Preparation Platform

African swine fever virus (ASFV) is the causative agent of a severe and highly contagious viral disease affecting domestic and wild swine. The current ASFV pandemic strain has a high mortality rate, severely impacting pig production and, for countries suffering outbreaks, preventing the export of their pig products for international trade. Early detection and diagnosis of ASFV is necessary to control new outbreaks before the disease spreads rapidly. One of the rate-limiting steps to identify ASFV by next-generation sequencing platforms is library preparation. Here, we investigated the capability of the Oxford Nanopore Technologies’ VolTRAX platform for automated DNA library preparation with downstream sequencing on Nanopore sequencing platforms as a proof-of-concept study to rapidly identify the strain of ASFV. Within minutes, DNA libraries prepared using VolTRAX generated near-full genome sequences of ASFV. Thus, our data highlight the use of the VolTRAX as a platform for automated library preparation, coupled with sequencing on the MinION Mk1C for field sequencing or GridION within a laboratory setting. These results suggest a proof-of-concept study that VolTRAX is an effective tool for library preparation that can be used for the rapid and real-time detection of ASFV.

60 APPLIED LIFE SCIENCES↗

Techno-Economic Assessment of Data Center Load Demand Powered by Small Modular Reactors and Distributed Energy Resources

The rapid increase in data center energy demand, driven by AI and large-scale data processing, poses significant challenges to global energy infrastructure. Data centers require substantial and reliable energy for continuous operations and high-performance computing. Current electrical grids face issues such as transmission bottlenecks and aging infrastructure, making it difficult to meet these demands. Integrating inverter-based-resources (IBRs) like solar and wind presents both opportunities and challenges due to their intermittent nature. Small Modular Reactors (SMRs) offer a promising solution with their enhanced safety, modularity, reliability, and scalability, providing consistent base load power ideal for data center operations. This study presents a comprehensive techno-economic assessment of powering data center load demand using a combination of SMRs and IBRs with grid-connected and islanded mode. This study utilized Idaho National Laboratory’s (INL) HPC data center hourly load profiles and Xendee microgrid optimization platform to conduct the analysis. In this configuration, SMRs serves as the primary base load power source, consistently providing a steady supply of electricity necessary to meet the minimum load demand of the data center with support from the IBRs. Key performance indicators such as Levelized Cost of Electricity (LCOE), Net Present Value (NPV) has been calculated to assess the economic feasibility. The findings from this research will underscore the strategic benefits of integrating SMR plant with DERs – particularly for critical infrastructure load such as data centers.

14 - SOLAR ENERGY↗

GeoThermalCloud: Cloud Fusion of Big Data and Multi-Physics Models using Machine Learning for Discovery, Exploration, and Development of Hidden Geothermal Resources

The primary goals of this project are exploring hidden geothermal resources in the U.S.A. and designing profitable enhanced geothermal systems (EGS). Many processes and parameters control geothermal exploration and energy production from geothermal fields. Diverse datasets (e.g., geology, geochemistry, geophysics, satellite, airborne geophysics) are available to help characterize subsurface geothermal conditions. Sparse and multi-scale characteristics of these datasets prohibit properly leveraging these datasets for geothermal exploration and profitable EGS design. Recent advancements in machine learning (ML) promise to resolve these issues. The tremendous challenges and risks of geothermal exploration and production bring the demand for novel ML methods and tools that can (1) analyze large field datasets, (2) assimilate model simulations (large inputs and outputs), (3) process sparse datasets, (4) perform transfer learning (between sites with different exploratory levels), (5) extract hidden geothermal signatures in the field and simulation data, (6) label geothermal resources and processes, (7) identify high-value data acquisition targets, and (8) guide geothermal exploration and production by selecting optimal exploration, production, and drilling strategies. To address these necessities, ML-based geothermal resources exploration and enhanced geothermal systems (EGS) design tools have been developed. The exploration tool is called GeoThermalCloud and EGS design tool is called GeoDT-ML. GeoThermalCloud (https://github.com/SmartTensors/GeoThermalCloud.jl) utilizes a LANL unsupervised ML platform called SmartTensors (https://tensors.lanl.gov/) to automate data analyses and interpretations by extracting hidden signatures to identify geothermal prospects. Also, it enables the identification of critical measurements needed to identify geothermal resource signatures. Alternatively, GeoDT-ML (https://github.com/SmartTensors/GeoThermalCloud.jl/tree/master/EGS) is an ML-based alternative to GeoDT (https://github.com/GeoDesignTool/GeoDT.git), a fast, simplified multi-physics solver to evaluate EGS project designs in uncertain geologic systems. GeoDT-ML leverages recent advances in deep learning and high-performance computing. It is a faster and simpler version of GeoDT. To make this project a success, we used capabilities of LANL, PNNL, Google, Stanford, and Julia Computing. We analyzed eight datasets of the U.S.A. using GeothermalCloud and demonstrated potential highly prospective geothermal resources and identified key factors defining highly prospective sites. The first data set includes 44 locations in southwest New Mexico and 18 geological, hydrogeological, geophysical, geothermal, geochemical attributes. We defined low- and medium-temperature hydrothermal systems and discovered a new highly prospective site. The second data set analyzed 18 shallow water chemistry attributes at 14,342 locations in the Great Basin. It demarcated modestly, moderately, and highly prospective sites including key attributes for each type of prospectivity. The third data set analyzed Utah FORGE data including satellite (InSAR), geophysical (gravity, seismic), geochemical, and geothermal attributes. Here, we performed prospectivity analysis to identify future drilling locations using geological, geochemical, and geophysical attributes. Maps of temperature at depth and heat flow are constructed based on the available data. Prospectivity maps were generated, and drilling locations were proposed for future geothermal field exploration. The fourth data set analyzed 21 attributes at 120 locations in Tularosa Basin, New Mexico; data comes from past play fairway analyses in this region. ML analyses identified geothermal signatures associated with modestly, moderately, and highly hydrothermal systems. We also defined dominant attributes and spatial distribution of the geothermal signatures. The fifth, sixth, seventh, and eighth datasets include Tohatchi Springs, New Mexico, Hawaii, Brady site, Nevada, and EGS Collab, respectively. Moreover, we coupled GeothermalCloud and magnetotellurics data to pinpoint drilling locations for developing geothermal projects in the Tularosa Basin, New Mexico. GeothermalCloud found potential prospective locations for geothermal resources near White Sands Missile Range and McGregor Range at Fort Bliss. Magnetotellurics data determined the potential depth (~1800m) of geothermal prospects at McGregor Range based on apparent resistivity structures/layers in the subsurface. The McGregor Range consists of three resistivity layers and two resistivity structures. Magnetotellurics data also helps identify that the western portion of the McGregor Range has thick and low-resistivity earth materials. The low resistivity to the west is most likely for a fault system. Assuming temperature is consistent with a geothermal reservoir, the west-central part of the McGregor Range has the highest geothermal potential because of the increase in porosity and associated permeability attributed to the interpreted fault system. Also, we devised a coupling strategy between a process model and GeothermalCloud to characterize hydrogeological conditions and geothermal conditions, respectively. The process model characterizes hydrogeological and geothermal conditions on highly prospective geothermal sites provided by GeothermalCloud. We developed a physics-informed neural network (PINN) version of the Burns equation that can be easily coupled with GeothermalCloud. Furthermore, we performed an optimal design decision maximizing the economic value of an EGS power plant. This study optimized the range of well spacing between injection and production wells maximizing net present value in dollars (NPV). For this task, we used the GeoDT to simulate the Utah FORGE EGS development cycle from the initial well design to the end of production. Next, we accomplished another crucial task, which is predicting permeability of geothermal reservoirs. Predicting permeability of geothermal reservoirs is a non-trivial task because of huge computational runtime of simulation and lack of measurements. To avoid these limitations, we used easy-to-measure chemical concentrations in the subsurface as measurement data and convolutional neural network based ML model of a high-fidelity model. Next, we predicted permeability using Markov chain Monte Carlo simulation. We found that Markov chain Monte Carlo simulation predicts permeability with a high certainty if the prediction zone in the simulation area has chemical concentration data. Finally, we analyzed the DOE funded INGENIOUS and GeoDAWN projects data. For discovering hidden geothermal systems in the Great Basin, the INGENIOUS project accumulated old data, collected new data, and released them in 2022. The dataset includes a total of 24 geological, geophysical, and geochemical attributes. Data resolution and scale significantly vary prohibiting an appropriate usage. To avoid such limitations, we brought all data in the same resolution and scale by applying the inverse distance weighting interpolation technique for predicting data in unsampled locations. Subsequently, we analyzed LiDAR data of the GeoDAWN project. We received data in tiles format. The DOE’s overarching goal is to use ML on LiDAR data for finding favorable geological structures (e.g., step up faults in Brady, Nevada). To serve the purpose, we need to label favorable geologic structures that correspond to LiDAR data. We wrote an algorithm to label the LiDAR data with the favorable geologic structures.

15 GEOTHERMAL ENERGY↗

MDLoader: A Hybrid Model-Driven Data Loader for Distributed Graph Neural Network Training

Scalable data management is essential for processing large scientific dataset on HPC platforms for distributed deep learning. In-memory distributed storage is preferred for its speed, enabling rapid, random, and frequent data access required by stochastic optimizers. Processes use one-sided or collective communication to fetch remote data, with optimal performance depending on (i) dataset characteristics, (ii) training scale, and (iii) interconnection network. Empirical analysis shows collective communication excels with larger mini-batch sizes and/or fewer processes, whereas one-sided communication outperforms at larger scales. We propose MDLoader, a hybrid in-memory data loader for distributed graph neural network training. MDLoader features a model-driven performance estimator that dynamically selects between one-sided and collective communication at the beginning of training using Tree of Parzen Estimators (TPE). Evaluations on NERSC Perlmutter and OLCF Summit show MDLoader outperforms single-backend loaders by up to 2.83 × and predicts the suitable communication method with 96.3% (Perlmutter) and 94.3% (Summit) success rate.

Bae, Jonghyun↗

Cloud Fusion of Big Data and Multi-Physics Models using Machine Learning for Discovery, Exploration, and Development of Hidden Geothermal Resources

The primary goals of this project are identifying hidden geothermal resources in the USA and designing profitable enhanced geothermal systems (EGS). Many non-obvious processes and parameters could characterize geothermal resources and could control the ultimate energy potential of geothermal fields. Diverse datasets (e.g., geology, geochemistry, geophysics, satellite, airborne geophysics) are available to help characterize geothermal resources, but this data is sparse and multi-scale. This has hindered attempts to leverage the datasets for geothermal exploration and profitable EGS design. Recent advancements in machine learning (ML) give promise to overcome these issues. Modern ML methods and tools can (1) analyze large datasets, (2) assimilate model ensembles that include a multitude of inputs and outputs, (3) process sparse datasets, (4) perform transfer learning between sites with different data quality, (5) extract hidden geothermal signatures from field and simulation data, (6) label geothermal resources and processes, (7) identify high-value data acquisition targets, and (8) guide geothermal exploration and production by selecting optimal exploration, production, and drilling strategies. In this work, we implement ML-based geothermal exploration and an enhanced geothermal systems (EGS) design tool to achieve the above goals. Our exploration tool is GeoThermalCloud (GTC) EGS design tool is GeoDT-ML. GTC (github.com/SmartTensors/GeoThermalCloud.jl) utilizes a LANL unsupervised ML platform called SmartTensors (https://tensors.lanl.gov/) to automate data analyses and interpretations by extracting hidden signatures to identify geothermal prospects. It enables the identification of critical measurements needed to identify geothermal resource signatures. GeoDT-ML (github.com/SmartTensors/GeoThermalCloud.jl/tree/master/) adds coupling to GeoDT (https://github.com/GeoDesignTool/GeoDT.git) for stochastic EGS design optimization and performance prediction. GeoDT-ML leverages recent advances in deep learning and high-performance computing. Contributors to this effort include LANL, PNNL, Google, Stanford, and Julia Computing.

15 GEOTHERMAL ENERGY↗

Characterization of a CMOS camera based film digitization platform for gated x-ray imaging diagnostics at the National Ignition Facility

Hardened gated x-ray detectors use photographic film as the data recording medium due to its low sensitivity to the high-yield neutron environments at the National Ignition Facility (NIF). The photographic film is digitized with a Photometric Data Systems (PDS) microdensitometer, which measures the film’s optical density. The PDS scanner is able to measure a dynamic range of 0–5 OD; however, raster scanning the film is time consuming and maintenance of the instrument is challenging due to legacy technology. Since film usage at NIF is expected to continue in the foreseeable future, a digitization platform that is faster and more maintainable would benefit the NIF’s current and future operations. Here, this work presents the characterization of the digital transitions (DT) atom, a CMOS camera-based digitization platform that records film data in a single image capture very quickly and has widely available user support. The preliminary results suggest that the DT atom is able to reconstruct exposures accurately enough to be a competitive alternative to the PDS Scanner.

47 OTHER INSTRUMENTATION↗

Ocean surface radiation measurement best practices

Ocean surface radiation measurement best practices have been developed as a first step to support the interoperability of radiation measurements across multiple ocean platforms and between land and ocean networks. This document describes the consensus by a working group of radiation measurement experts from land, ocean, and aircraft communities. The scope was limited to broadband shortwave (solar) and longwave (terrestrial infrared) surface irradiance measurements for quantification of the surface radiation budget. Best practices for spectral measurements for biological purposes like photosynthetically active radiation and ocean color are only mentioned briefly to motivate future interactions between the physical surface flux and biological radiation measurement communities. Topics discussed in these best practices include instrument selection, handling of sensors and installation, data quality monitoring, data processing, and calibration. It is recognized that platform and resource limitations may prohibit incorporating all best practices into all measurements and that spatial coverage is also an important motivator for expanding current networks. Thus, one of the key recommendations is to perform interoperability experiments that can help quantify the uncertainty of different practices and lay the groundwork for a multi-tiered global network with a mix of high-accuracy reference stations and lower-cost platforms and practices that can fill in spatial gaps.

54 ENVIRONMENTAL SCIENCES↗

LinkML: an open data modeling framework

Background Scientific research relies on well-structured, standardized data; however, much of it is stored in formats such as free-text lab notebooks, nonstandardized spreadsheets, or data repositories. This lack of structure challenges interoperability, making data integration, validation, and reuse difficult. Findings LinkML (Linked Data Modeling Language) is an open framework that simplifies the process of authoring, validating, and sharing data. LinkML can describe a range of data structures, from flat, list-based models to complex, interrelated, and normalized models that utilize polymorphism and compound inheritance. It offers an approachable syntax that is not tied to any one technical architecture and can be integrated seamlessly with many existing frameworks. The LinkML syntax provides a standard way to describe schemas, classes, and relationships, allowing modelers to build well-defined, stable, and optionally ontology-aligned data structures. Once defined, LinkML schemas may be imported into other LinkML schemas. These key features make LinkML an accessible platform for interdisciplinary collaboration and a reliable way to define and share data semantics. Conclusions LinkML helps reduce heterogeneity, complexity, and the proliferation of single-use data models while simultaneously enabling compliance with FAIR (Findable, Accessible, Interoperable, and Reusable) data standards. LinkML has seen increasing adoption in various fields, including biology, chemistry, biomedicine, microbiome research, finance, electrical engineering, transportation, and commercial software development. In short, LinkML makes implicit models explicitly computable and allows data to be standardized at their origin. LinkML documentation and code are available at https://linkml.io/.

AI-ready data↗