Search NASA⌕ Search

SEARCH · Search NASA

Results for “INTERPRETATION”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Utah FORGE: RESMAN Well 16A(78)-32 and 16B(78)-32 Stimulation and Circulation Tracer Test Results - 2024

This dataset contains tracer test results from stimulation and circulation experiments conducted on the Utah FORGE wells 16A(78)-32 and 16B(78)-32 during 2024. The data was collected by RESMAN Energy Technology and includes detailed tracer analysis from flowback, short- and extended-duration circulation tests, and reinjection sampling. Sampling included analysis of tracers during different stages of testing in April, August, and September 2024. The dataset is accompanied by an interpretation report and contains time-series tracer concentration data with identification of test phases and sampling conditions. It includes results for flowback from well 16A, commingling effects with water from well 16B, tracer data from short and extended circulation tests, and reinjection tracer corrections for the August/September test. Users should be aware that proprietary tracer methodologies were applied, and they should consult the interpretation report for insights into experimental procedures and data contextualization.

15 GEOTHERMAL ENERGY↗

Feedback, physics, and forecasts: The emerging paradigm of machine learning-driven battery research

Machine learning (ML) is reshaping how we understand, predict, and optimize electrochemical systems. In batteries, ML accelerates discovery across chemistry, design, and operation by transforming massive experimental and simulated datasets into predictive, interpretable models. This review consolidates a decade of progress in ML-driven battery innovation, from early-cycle feature extraction to operando image analysis and physics-informed modeling. We categorize approaches by data domain and physical fidelity, emphasizing interpretable ML for diagnostics, reinforcement learning for control, and multi-objective optimization for lifetime extension strategies. Additionally, we demonstrate how integrated models accelerate discovery, reduce testing time, and guide sustainable design. Economic analyses furthermore illustrate how these advances can lower cost per cycle and improve circularity. Together, these developments chart a path toward self-optimizing, sustainable battery technologies.

artificial intelligence↗

Characterization of Pliocene and Miocene Formations in the Wilmington Graben, Offshore Los Angeles, for Large-Scale Geologic Storage of CO2

The project Characterization of Pliocene and Miocene Formations in the Wilmington Graben, Offshore Los Angeles, for Large-Scale Geologic Storage of CO2 is one of 9 site characterization projects that were implemented as part of ARRA (American Recovery and Reinvestment Act). Data from this project was used to improve resolution of data in NATCARB in the area of study. Data related to this study has already been incorporated in NATCARB Atlas. The Los Angeles Basin presents an opportunity for large-scale geologic CO2 storage. Due to its large population and historical and geologic setting as one of the most prolific oil and gas producing basins in the United States, the region is home to more than 12 major power plants and oil refineries that produce more than 5 million metric tons of fossil fuel-related CO2 emissions each year. GeoMechanics Technologies worked to characterize the Pliocene and Miocene sediments in the Wilmington Graben, offshore of Los Angeles, California, for high-volume CO2 storage. The Graben is located offshore of the Los Angeles and Long Beach Harbor area, making it accessible yet geologically isolated from the nearby Wilmington oilfield and onshore areas. These sediments span more than 5,000 feet of vertical interval with an estimated storage resource of more than 100 million metric tons of CO2. The project team analyzed and interpreted existing geologic data within the region, including detailed exploration well log data and 2-D and 3-D seismic data. New seismic lines were acquired to fill in current data gap areas and two new characterization wells were drilled and logged. This information was integrated with existing geologic interpretations for adjacent onshore areas to help characterize optimal areas for CO2 storage and seals to safely store CO2. Integrated 3-D geologic and geomechanical models for the Wilmington Graben were developed to simulate the fate and transport of injected CO2 in the subsurface and to assess risks. This project contributed to the understanding of injectivity, containment mechanisms, rate of dissolution and mineralization, and storage capacity of the Wilmington Graben and associated analogous basins. This effort also provided greater insight into the potential for offshore geologic formations to safely and permanently store CO2.

.las↗

Higher-form symmetry and chiral transport in real-time Abelian lattice gauge theory

We study classical lattice simulations of theories of electrodynamics coupled to charged matter at finite temperature, interpreting them using the higher-form symmetry formulation of magnetohydrodynamics (MHD). We compute transport coefficients using classical Kubo formulas on the lattice and show that the properties of the simulated plasma are in complete agreement with the predictions from effective field theories. In particular, the higher-form formulation allows us to understand from hydrodynamic considerations the relaxation rate of axial charge in the chiral plasma observed in previous simulations. A key point is that the resistivity of the plasma – defined in terms of Kubo formulas for the electric field in the 1-form formulation of MHD – remains a well-defined and predictive quantity at strong electromagnetic coupling. However, the Kubo formulas used to define the conventional conductivity vanish at low frequencies due to electrodynamic fluctuations, and thus the concept of the conductivity of a gauged electric current must be interpreted with care.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

X-ray Spectral-Timing Pipeline to Investigate an Electron-Scattering Time Delay in Black Hole Accretion Disks [Slides]

The soft lag in black hole X-ray binaries (BHXRBs) refers to the time delay for the soft, thermal disk to respond to hard, variable coronal irradiation. This time lag increases from less than 1 ms to ~10 ms during the hard-to-soft state transition. Interpretations of soft lag trends appeal to changing the light-travel path via an evolving coronal height and/or inner accretion disk radius. Both interpretations neglect a time delay contribution from the reprocessing of irradiation inside the disk, where electron-scattering opacity dominates. A new theory considering a thermalization (electron scattering) time delay in the disk can plausibly produce ~10 ms time delays in the intermediate state. To further investigate the impacts of adding a time delay component from electron scattering in the disk, we are creating a spectral-timing pipeline that can analyze NICER (Neutron Star Interior Composition Explorer) X-ray observations of BHXRBs in outburst. In the near future, we will apply this spectral timing pipeline to develop a reverberation lag model from simulations that include the thermalization time delay from electron scattering in the disk.

79 ASTRONOMY AND ASTROPHYSICS↗

Manganese-rich sandstones as an indicator of ancient oxic lake water conditions in Gale crater, Mars

Manganese has been observed on Mars by the NASA Curiosity rover in a variety of contexts and is an important indicator of redox processes in hydrologic systems on Earth. Within the Murray formation, an ancient primarily fine-grained lacustrine sedimentary deposit in Gale crater, Mars, have observed up to 45× enrichment in manganese and up to 1.5× enrichment in iron within coarser grained bedrock targets compared to the mean Murray sediment composition. This enrichment in manganese coincides with the transition between two stratigraphic units within the Murray: Sutton Island, interpreted as a lake margin environment, and Blunts Point, interpreted as a lake environment. On Earth, lacustrine environments are common locations of manganese precipitation due to highly oxidizing conditions in the lakes. Here, we explore three mechanisms for ferromanganese oxide precipitation at this location: authigenic precipitation from lake water along a lake shore, authigenic precipitation from reduced groundwater discharging through porous sands along a lake shore, and early diagenetic precipitation from groundwater through porous sands. All three scenarios require highly oxidizing conditions and we discuss oxidants that may be responsible for the oxidation and precipitation of manganese oxides. This work has important implications for the habitability of Mars to microbes that could have used Mn redox reactions, owing to its multiple redox states, as an energy source for metabolism.

58 GEOSCIENCES↗

Specifications of Cladding Diameter Measurements Conducted at AGHCF

Fuel element diameter data was collected on-site at the Alpha-Gamma Hot Cell Facility (AGHCF) before and after out-of-pile furnace transient tests of fuel elements. Available data records have been collected and preserved in the Out-of-Pile Transient Database (OPTD). This data is used to determine the transient-induced changes in fuel element diameter, or cladding strain, for the Whole Pin Furnace (WPF) tests. Fuel element diameter was measured by contact profilometry along the length of the fuel pin at specified rotational orientations and/or by a manually operated micrometer at several discrete axial locations along the pin. The Alpha-Gamma Hot Cell Facility Operations Manual details the way the facility operated, organizational and oversight responsibilities, and procedures for examination of samples. The working version of the manual at the time of the whole pin furnace tests and the test pin examinations is Doc. No. IPS-2-00-00, dated June 1989. This specification was developed using the operations manual and recovered measurement records in consultation with subject matter experts (SMEs). The measurement methods, format of the available post-test examination (PTE) data, and recommended methods for use and interpretation of that data are summarized. Section 2 describes the available diameter measurement data for the WPF tests with guidance for interpretation and usage, and Section 3 describes the instrument measurement procedures and calibration methods.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Elucidating Photoinduced Processes of Photosystem I Via Multidimensional Electronic and Vibrational Spectroscopies

This project was motivated by an overarching goal to elucidate the mechanism of energy and electron transfer that governs the efficient charge separation in photosystem I (PSI) complexes. PSI is a natural light harvesting complex that drives oxygenic photosynthesis in plants, algae, and cyanobacteria. It uses ~300 tightly packed chlorophylls (Chls) to absorb photons, transfer the excitation energy to the reaction center (RC), and generate a charge separated state with near unity quantum efficiency (QE). A better understanding of the mechanism of energy transfer and charge separation in PSI is required for understanding the high QE of natural light harvesting complexes, and it could lead to the further development of artificial photosynthetic systems for solar energy conversion and modification of light harvesting complexes to improve crop yields. We applied two-dimensional optical spectroscopies to different cyanobacterial photosystem I complexes, including PSI complexes that contain Chl f molecules, to map energy transfer pathways and gain insight into the efficient light harvesting of PSI. We used two-dimensional electronic spectroscopies (2DES) to map energy transfer in Chl a and Chl f containing PSI complexes. To investigate the Chl f PSI complexes, we modified our spectrometer to probe the lower energy states associated with Chl f molecules. We interpreted the 2DES spectra through global analysis procedures to generate maps of energy transfer. We also constructed a two-dimensional electronic vibrational (2DEV) spectrometer that will be used to investigate charge transfer transitions and dynamics within PSI complexes. Measurements were performed on model systems to establish general data analysis procedures for interpreting 2D spectra and gain insight into protein cofactor interactions.

14 SOLAR ENERGY↗

Data-Enabled Fusion Technology (Final Scientific/Technical Report)

Advancing Scientific Understanding in Fusion Energy and Machine Learning This research represented a significant step forward in machine learning (ML) applications for fusion energy experiments. The project integrated advanced data-driven modeling, optimization techniques, and artificial intelligence to enhance the predictive capabilities and operational efficiency of plasma-based fusion systems. Specifically, tasks focused on ML-enhanced diagnostics, operator guidance tools, and predictive modeling helped improve the ability to interpret complex fusion experiments. Key areas of advancement included: 1) data-driven plasma control, i.e., using ML algorithms to optimize experimental conditions and classify plasma behaviors based on historical data; 2) spectroscopy and diagnostics, i.e., applying AI models to extract previously inaccessible insights from experimental spectroscopy data; and 3) configuration mapping and operator guidance, i.e., developing a predictive framework to assist scientists in identifying the most effective experimental parameters, reducing reliance on manual adjustments. By refining these ML-driven techniques, the project contributed to the broader scientific community’s understanding of plasma dynamics and fusion energy viability. Technical Effectiveness and Economic Feasibility The methods investigated demonstrated high technical effectiveness, as reflected in milestones assessing the predictive accuracy, performance, and optimization of fusion configurations. The development of an Operator Guidance Tool (OGT), for example, led to more precise control of plasma conditions by learning from experimental data and offering real-time adjustments. From an economic standpoint, DeFT provided: 1) the ability to reduce trial-and-error experimentation, which lowered operational costs; 2) improved data interpretation methods, which enabled more efficient resource allocation in large-scale fusion research projects; and 3) the automation of key diagnostic tasks, which reduced manual labor and human error, increasing overall efficiency. 13 The final assessments of predictive models and optimization strategies demonstrated that these approaches were scalable and could be implemented across multiple fusion energy research programs. Public Benefit and Societal Impact This project contributed directly to the broader goal of achieving sustainable and commercially viable fusion energy, which had profound implications for clean energy production and climate change mitigation. The integration of AI-driven solutions into fusion research: 1) sped up scientific discovery, accelerating progress towards achieving energy breakthroughs; 2) reduced the cost of experimentation, making fusion research more accessible; and 3) provided a framework for future AI applications in high-energy physics, benefiting adjacent fields like space exploration, material science, and renewable energy. Additionally, by fostering collaborations between AI researchers and plasma physicists, this project promoted interdisciplinary innovation that could lead to broader applications beyond fusion research.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Creation of a Weather Drivers Test Suite for Inclusion in ASHRAE Standard 140

Weather conditions are an important boundary condition for building performance simulation (BPS) calculations. For existing test cases in ASHRAE Standard 140 "Method of Test for Evaluating Building Performance Simulation Software" (ANSI/ASHRAE 2020), it was assumed that the software being tested could adequately read and interpret the weather data in the provided standard weather files. As differences between the programs have been reduced and as more programs have shifted to sub-hourly time steps this assumption has become more stretched. To address these concerns a new test suite testing a program's ability to read and interpret the data from a standard weather file was developed. The purpose of the test suite is to test the use of the typical data used from standard weather files.

54 ENVIRONMENTAL SCIENCES↗

TRANSP Workshop Summary - September 27-28, 2024, Princeton Plasma Physics Laboratory, NJ

The TRANSP Code Workshop provided a platform for in-depth discussions on advancing the capabilities of the TRANSP code, focusing on key areas such as predictive capabilities, interpretive frameworks, core-edge coupling, and integration with engineering components. More than 25 scientists from PPPL and around the world contributed to the workshop by making presentations and participating in discussions. The workshop covered a range of topics including: (1) Current Status of TRANSP; (2) Code Infrastructure, Core-Edge Coupling, and Engineering Integration; (3) Enhancing Interpretive Capabilities, and (4) Predictive Capabilities for Discharge Optimization

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Yes, No, Maybe So: Human Factors Considerations for Fostering Calibrated Trust in Foundation Models Under Uncertainty

High-stakes analytical environments require analysts to evaluate evidence and generate conclusions to inform critical decisions often under conditions of uncertainty. Probabilistic decision-making based on incomplete or inaccurate information can reduce productivity, compromise national interests, and endanger public safety. Researchers are developing expert systems built on foundation models (FMs) to support analysts’ decision-making processes by enabling human-artificial intelligence (AI) teaming, in part through the quantification and expression of uncertainty information. As FMs continue to mature, it is imperative to correspondingly consider analysts’ needs for appropriately interpreting and using uncertainty information. However, prior research indicates that it remains unclear how analysts engage with FM-generated uncertainty information and the extent to which these interactions influence trust in, and reliance on, expert systems. We plan to review the state of the science and conduct an exploratory, qualitative study to (a) understand how properly communicated uncertainty can foster calibrated trust and appropriate reliance and (b) identify approaches for effectively conveying FM-generated uncertainty information during analytical workflows. We will administer semi-structured interviews with analysts from a specific high-stakes analytical environment to collect their current experiences with job-related uncertainty and their impressions when viewing FM-generated uncertainty information. During the interview protocol, participants will be presented with several different FM outputs and invited to discuss their thoughts and beliefs about the uncertainty information displayed. Participants may provide insights into how trust and reliance may be influenced by uncertainty. The results of this study will help us to better understand how analysts currently interpret and use uncertainty information. Our findings may inform human factors recommendations for effectively conveying uncertainty information to foster calibrated trust in, and appropriate reliance on, expert systems. Interaction designers and FM developers can use this knowledge to enhance human-AI teaming and ensure the responsible deployment of FM-based expert systems in analytical workflows.

97 MATHEMATICS AND COMPUTING↗

Computational Prediction of Infrasound Arrival Times and Directions from Stationary and Moving Impulsive Sources

This report addresses the need to predict infrasound signal arrival times and back azimuths at monitoring stations, enabling more focused and efficient searches within recorded waveform data. The primary challenge is estimating expected signal arrival windows for stationary and moving acoustic sources, such as chemical explosions, volcanic eruptions, meteoroids, and spacecraft re-entry events. To address this challenge, a reproducible methodology is described that uses simplified propagation speeds for boundary layer, tropospheric, stratospheric, and thermospheric atmospheric waveguides. While the Python source code itself is not freely available, this document provides detailed, step-by-step instructions, and equations enabling users to replicate and adapt the method independently. The method reliably predicts signal arrival intervals and back azimuths, thereby supporting rapid detection and accurate interpretation of infrasound events. Results demonstrate that this method effectively identifies plausible signal arrival intervals and directions, facilitating faster event detection and more reliable interpretation. This methodology directly supports atmospheric monitoring, planetary defense, and forensic analysis of explosive atmospheric events.

47 OTHER INSTRUMENTATION↗

NeuroSymbolic Approaches as a Vector for Assured Artificial Intelligence

The deployment of artificial intelligence systems in critical applications requires higher levels of assurance for safety, security, and interpretability. While neurosymbolic (NESY) approaches combining neural networks with symbolic reasoning offer potential advantages for assured AI, existing differentiable neurosymbolic frameworks face significant limitations including computational overhead and performance constraints. This report investigates the ISED (InferSampleEstimateDescend) framework as an alternative approach that enables neurosymbolic learning without requiring endtoend differentiability. We evaluate ISED’s utility for geointelligence applications by comparing neurosymbolic models against standard neural networks on aircraft classification tasks using the RarePlanes and MTARSI imagery datasets. Our results demonstrate that while ISEDbased models achieve slightly lower accuracy (89.7% vs 92.1% on RarePlanes; 91.1% vs 92.5% on MTARSI), they provide critical explainability capabilities that enable tracing incorrect predictions back to specific attribute misclassifications. We also present an automated pipeline that generates both attributeclass mappings and neurosymbolic model architectures from natural language descriptions, significantly reducing the manual effort required for NESY model deployment. These findings suggest that ISED offers a promising direction for developing assured AI systems where interpretability and reasoning transparency are prioritized alongside performance.

97 MATHEMATICS AND COMPUTING↗

GeoThermalCloud: Cloud Fusion of Big Data and Multi-Physics Models using Machine Learning for Discovery, Exploration, and Development of Hidden Geothermal Resources

The primary goals of this project are exploring hidden geothermal resources in the U.S.A. and designing profitable enhanced geothermal systems (EGS). Many processes and parameters control geothermal exploration and energy production from geothermal fields. Diverse datasets (e.g., geology, geochemistry, geophysics, satellite, airborne geophysics) are available to help characterize subsurface geothermal conditions. Sparse and multi-scale characteristics of these datasets prohibit properly leveraging these datasets for geothermal exploration and profitable EGS design. Recent advancements in machine learning (ML) promise to resolve these issues. The tremendous challenges and risks of geothermal exploration and production bring the demand for novel ML methods and tools that can (1) analyze large field datasets, (2) assimilate model simulations (large inputs and outputs), (3) process sparse datasets, (4) perform transfer learning (between sites with different exploratory levels), (5) extract hidden geothermal signatures in the field and simulation data, (6) label geothermal resources and processes, (7) identify high-value data acquisition targets, and (8) guide geothermal exploration and production by selecting optimal exploration, production, and drilling strategies. To address these necessities, ML-based geothermal resources exploration and enhanced geothermal systems (EGS) design tools have been developed. The exploration tool is called GeoThermalCloud and EGS design tool is called GeoDT-ML. GeoThermalCloud (https://github.com/SmartTensors/GeoThermalCloud.jl) utilizes a LANL unsupervised ML platform called SmartTensors (https://tensors.lanl.gov/) to automate data analyses and interpretations by extracting hidden signatures to identify geothermal prospects. Also, it enables the identification of critical measurements needed to identify geothermal resource signatures. Alternatively, GeoDT-ML (https://github.com/SmartTensors/GeoThermalCloud.jl/tree/master/EGS) is an ML-based alternative to GeoDT (https://github.com/GeoDesignTool/GeoDT.git), a fast, simplified multi-physics solver to evaluate EGS project designs in uncertain geologic systems. GeoDT-ML leverages recent advances in deep learning and high-performance computing. It is a faster and simpler version of GeoDT. To make this project a success, we used capabilities of LANL, PNNL, Google, Stanford, and Julia Computing. We analyzed eight datasets of the U.S.A. using GeothermalCloud and demonstrated potential highly prospective geothermal resources and identified key factors defining highly prospective sites. The first data set includes 44 locations in southwest New Mexico and 18 geological, hydrogeological, geophysical, geothermal, geochemical attributes. We defined low- and medium-temperature hydrothermal systems and discovered a new highly prospective site. The second data set analyzed 18 shallow water chemistry attributes at 14,342 locations in the Great Basin. It demarcated modestly, moderately, and highly prospective sites including key attributes for each type of prospectivity. The third data set analyzed Utah FORGE data including satellite (InSAR), geophysical (gravity, seismic), geochemical, and geothermal attributes. Here, we performed prospectivity analysis to identify future drilling locations using geological, geochemical, and geophysical attributes. Maps of temperature at depth and heat flow are constructed based on the available data. Prospectivity maps were generated, and drilling locations were proposed for future geothermal field exploration. The fourth data set analyzed 21 attributes at 120 locations in Tularosa Basin, New Mexico; data comes from past play fairway analyses in this region. ML analyses identified geothermal signatures associated with modestly, moderately, and highly hydrothermal systems. We also defined dominant attributes and spatial distribution of the geothermal signatures. The fifth, sixth, seventh, and eighth datasets include Tohatchi Springs, New Mexico, Hawaii, Brady site, Nevada, and EGS Collab, respectively. Moreover, we coupled GeothermalCloud and magnetotellurics data to pinpoint drilling locations for developing geothermal projects in the Tularosa Basin, New Mexico. GeothermalCloud found potential prospective locations for geothermal resources near White Sands Missile Range and McGregor Range at Fort Bliss. Magnetotellurics data determined the potential depth (~1800m) of geothermal prospects at McGregor Range based on apparent resistivity structures/layers in the subsurface. The McGregor Range consists of three resistivity layers and two resistivity structures. Magnetotellurics data also helps identify that the western portion of the McGregor Range has thick and low-resistivity earth materials. The low resistivity to the west is most likely for a fault system. Assuming temperature is consistent with a geothermal reservoir, the west-central part of the McGregor Range has the highest geothermal potential because of the increase in porosity and associated permeability attributed to the interpreted fault system. Also, we devised a coupling strategy between a process model and GeothermalCloud to characterize hydrogeological conditions and geothermal conditions, respectively. The process model characterizes hydrogeological and geothermal conditions on highly prospective geothermal sites provided by GeothermalCloud. We developed a physics-informed neural network (PINN) version of the Burns equation that can be easily coupled with GeothermalCloud. Furthermore, we performed an optimal design decision maximizing the economic value of an EGS power plant. This study optimized the range of well spacing between injection and production wells maximizing net present value in dollars (NPV). For this task, we used the GeoDT to simulate the Utah FORGE EGS development cycle from the initial well design to the end of production. Next, we accomplished another crucial task, which is predicting permeability of geothermal reservoirs. Predicting permeability of geothermal reservoirs is a non-trivial task because of huge computational runtime of simulation and lack of measurements. To avoid these limitations, we used easy-to-measure chemical concentrations in the subsurface as measurement data and convolutional neural network based ML model of a high-fidelity model. Next, we predicted permeability using Markov chain Monte Carlo simulation. We found that Markov chain Monte Carlo simulation predicts permeability with a high certainty if the prediction zone in the simulation area has chemical concentration data. Finally, we analyzed the DOE funded INGENIOUS and GeoDAWN projects data. For discovering hidden geothermal systems in the Great Basin, the INGENIOUS project accumulated old data, collected new data, and released them in 2022. The dataset includes a total of 24 geological, geophysical, and geochemical attributes. Data resolution and scale significantly vary prohibiting an appropriate usage. To avoid such limitations, we brought all data in the same resolution and scale by applying the inverse distance weighting interpolation technique for predicting data in unsampled locations. Subsequently, we analyzed LiDAR data of the GeoDAWN project. We received data in tiles format. The DOE’s overarching goal is to use ML on LiDAR data for finding favorable geological structures (e.g., step up faults in Brady, Nevada). To serve the purpose, we need to label favorable geologic structures that correspond to LiDAR data. We wrote an algorithm to label the LiDAR data with the favorable geologic structures.

15 GEOTHERMAL ENERGY↗

Reliable and Efficient Machine Learning (Final Technical Report)

Modern scientific experiments generate massive amounts of data at a pace much faster than humans can manually analyze. While machine learning has revolutionized commercial data analysis (such as recommending movies or recognizing faces), applying these tools to complex scientific discovery is challenging because scientific answers must be precise, interpretable, and adhere to physical laws. The research under this project aims to develop new mathematical tools and computer algorithms specifically designed for scientific applications. Major progress has been made in automatically cleaning and deconstructing messy experimental data, analyzing the visual information of physical phenomena, determining the underlying physical variables, and providing rig orous mathematical analysis of interesting algorithms and concepts widely used in machine learning. This project addressed the critical gap between our ability to generate massive scientific data and our ability to extract interpretable information from it. We established mathematical foundations for Scientific Machine Learning (SciML) aimed at effective data analytics and automated discovery. Our work focused on three core objectives: (1) developing reliable feature extraction methods for dynamic high-dimensional data, (2) establishing mathematical foundations for discovering dynamics via neural networks, and (3) creating rigorous optimization techniques for these models. Key outcomes come from two fronts. On the practical side, they include the development of algorithms that significantly enhance the extraction of signals from field data, as well as the capability to handle situations that exhibit smooth variations or physical stretching due to temperature changes. They also include the creation of an automated framework for discovering fundamental state variables from raw experimental data, demonstrating the ability to identify intrinsic physical dimensions without prior knowledge of the governing laws. On the theoretical front, the research results in theoretical advances in Optimal Transport, a widely used notion in SciML, specifically regarding functions with fixed-size nodal sets, provide sharp bounds relevant to uncertainty quantification. Meanwhile, the outcomes also include the establishment of convergence theories for nonlocal gradient descent methods, enabling robust optimization with noisy data in high-dimensional settings commonly encountered in scientific modeling. The project also helps creating opportunities to train the next generation of researchers, equipping them with the necessary technical skills for today’s workplace and preparing them for future advances.

97 MATHEMATICS AND COMPUTING↗

New Particle Formation and Growth in the Houston Atmosphere During TRACER (Final Report)

From 2020-2025, researchers from UC Irvine, UC Riverside, and Colorado State University collaborated on a Department of Energy-funded project to understand how airborne particles form and grow in urban atmospheres, conducting an intensive field campaign in Houston, Texas during summer 2022. Using advanced instruments to measure gas-phase chemicals, particle composition, and a specialized chamber to study particle growth, the team discovered that sulfur-containing compounds from industrial and power plant emissions are the dominant driver of new particle formation in Houston, with particles typically forming locally in the city and growing as air moves away in the urban plume. The research revealed an important methodological insight: measurements from fixed ground stations can be misleading when interpreting how particles actually evolve as air masses move, which has significant implications for how scientists worldwide interpret atmospheric observations. These findings improve understanding of urban air quality and help reduce uncertainties in climate models, since these particles play critical roles in cloud formation and Earth's radiation balance, while also providing detailed information about ultrafine particle composition relevant to public health. The project trained three doctoral students, developed enhanced computer models for urban particle formation, and made all data publicly available through the DOE Atmospheric Radiation Measurement data archive for use by the broader scientific community.

54 ENVIRONMENTAL SCIENCES↗

Data Summarization and Inference at Scale

This is the final report for the DOE ASCR grant SC-0022260, Data Summarization and Inference at Scale, PI: Alex Pothen, Purdue University. The goal of the project was to solve data-intensive and compute-intensive problems in the physical sciences, engineering, information science, data science, etc. by designing and implementing new algorithms that could work with a subset of the data. The four subgoals were: (a) The solution of problems where the data is too large to be stored in the memory of a computer. In this streaming model of computation, the data arrives as a stream of elements to the computer, each element is processed as it arrives, and a decision is made to discard the data or to store it; only a small subset of the data proportional to the size of the output solution is stored, and when all the data has been streamed, a solution to the problem is computed from the stored subset. (b) The use of machine learning methods to compute solutions to data-intensive problems. The use of GPUs is critical to obtain high performance on machine learning tasks, but their memory sizes are smaller relative to that of CPUs. For large-scale problems, the data is sampled many times, and small samples are used with repetition, for robustness, to compute solutions to inference tasks. This sampling reduces the memory required to solve the problem, but attention is needed to avoid slow convergence to the solutions, and reduced accuracy of inference. We propose submodular optimization, Large Language Models, and physics-informed neural networks to enable GPU computations here. (c) Modeling and visualization of high-dimensional data using interpretable features. Clinical proteomic data sets from immunology for the detection of cancer and other diseases are temporal and high-dimensional, and algorithms for visualizing these data sets using clinically interpretable features are lacking. We propose methods that compute distances based on the optimal transportation problem and graph edit distances to address this problem. We also propose the use of optimal transport-based distances, spatial statistics, and network structure to classify image data sets, We apply these algorithms to electron micrographs of the peripheral nervous system in the digestive tract. (d) The design of data-intensive algorithms on emerging architectures, specifically, noisy, intermediate-scale quantum (NISQ) devices. Quantum computers offer the possibility of exploring large solution spaces due to the principle of superposition, but current quantum computers are limited by few qubits, short coherence times due to noise, poor interconections among the qubits, etc. We propose the use of the divide and conquer paradigm to solve large-scale problems, wherein collections of small subproblems are solved on the quantum devices, and the solutions to the subproblems are integrated into a solution for the original problem on a classical computer.

97 MATHEMATICS AND COMPUTING↗