Search NASA⌕ Search

SEARCH · Search NASA

Results for “task analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

Efficient Anomaly Detection Driven By Different Machine Learning Architectures And Models

The rapid growth and ubiquitous adoption of the internet and cyber-physical systems (CPS) have fundamentally transformed modern communication, work, and human-system interactions. While networks now form the backbone of critical digital ecosystems, enabling seamless data transmission across diverse, interconnected systems, this increased connectivity also expands the attack surface, making real-time detection of network intrusions and anomalies a pressing challenge. Detecting unusual activities within network infrastructure requires advanced data traffic analysis to differentiate between legitimate and malicious interactions. Traditional approaches to network anomaly detectionâ??such as rule-based and signature-based systemsâ??often depend on predefined patterns to identify known anomalies, limiting their effectiveness against emerging, stealthy, or previously unseen threats. These conventional methods suffer from high false alarm rates and fail to adapt to the ever-evolving nature of network traffic, particularly in large-scale, decentralized environments where data volume, velocity, and variety are constantly increasing. This dissertation presents artificial intelligence (AI)-driven approaches to anomaly detection that leverage graphics processing unit (GPU)-enabled high-performance computing (HPC) platforms for processing massive network traffic data and monitoring the components of cyber-physical systems (CPS) for potentially hazardous conditions. The research advances several key contributions: (1) Designing efficient machine learning techniques for CPS condition monitoring and anomaly detection; (2) enabling federated learning (FL) frameworks that enable distributed detection while preserving data privacy and system resilience; (3) exploring graph-based methodologies combining graph neural networks (GNN) and graph machine learning (ML) approaches for the Internet of Things (IoT) and automotive network security, and (4) performing distributed edge computing optimizations that integrate FL with scalable technologies for reduced communication overhead. Through extensive experiments, these methodologies demonstrate that complex anomaly detection and condition monitoring tasks can be achieved while balancing computational efficiency and detection accuracy through fine-grained network information processing. The frameworks developed in this research establish a robust foundation for network anomaly detection, providing scalable, adaptive, and privacy-preserving solutions for safeguarding CPS and IoT networks in an increasingly interconnected digital landscape. The practical implications of these research findings are significant, as they can inform the development of next-generation network security systems and contribute to the protection of critical infrastructure against sophisticated cyber attacks.

Marfo, William↗

Surrogate-driven design optimization with uncertainty constraints in Monte Carlo simulations

In multi-objective design tasks, the computational cost increases rapidly when high-fidelity simulations are used to evaluate objective functions. Surrogate models help mitigate this cost by approximating the simulation output, simplifying the design process. However, under high uncertainty, surrogate models trained on noisy data can produce inaccurate predictions, as their performance depends heavily on the quality of training data. This study investigates the impact of data uncertainty on two multi-objective design problems modelled using Monte Carlo transport simulations: a neutron moderator and an ion-to-neutron converter. For each, a grid search was performed using five different tally uncertainty levels to generate training data for neural network surrogate models. These models were then optimized using NSGA-III. The recovered Pareto-fronts were analyzed across uncertainty levels: in the moderator problem, normalized hypervolume dropped from 0.886 at 1.0% uncertainty to 0.748 at 10% uncertainty, while in the converter problem it remained near 0.50 for all cases. Average simulation times were also compared to evaluate the trade-off between accuracy and computational cost. Results show that the influence of simulation uncertainty is strongly problem-dependent. In the neutron moderator case, higher uncertainties led to exaggerated objective sensitivities and distorted Pareto-fronts, reducing normalized hypervolume. In contrast, the ion-to-neutron converter task was less affected—low-fidelity simulations produced results similar to those from high-fidelity data. These findings suggest that a fixed-fidelity approach is not optimal. Surrogate models can recover the Pareto-front under noisy conditions, and multi-fidelity studies help identify suitable uncertainty levels for each problem to balance efficiency and accuracy.

07 ISOTOPE AND RADIATION SOURCES↗

Effect of realistic routing on the social burden metric

The distance people travel to reach critical services is a key input to the Social Burden metric used by Sandia’s Resilient Node Cluster Analysis Tool (ReNCAT) in the optimization’s objective function. By default, ReNCAT utilizes Euclidian distances between population blocks and critical facilities when calculating Social Burden. However, these straight-line distances do not reflect how most residents or goods would travel throughout the area. As distance is a vital input to the burden calculation, a more realistic distance calculation will yield more realistic burden values. This work uses real road networks and calculates the shortest distance path between population centers and critical facilities using a standard graph theory approach. These realistic route distances are then used to compute Social Burden for four areas of study. It was found that distances using real road routes are generally, but not always, longer than the Euclidean distance. The increased length increases the final Social Burden metric, however, the overall burden percent change ranged between 17% and 52%, which means the impact of realistic routes relies heavily upon the area’s road topology. It was found that rural locations within an area may have larger burden increases than urban areas as more dense road networks allow routes to more closely follow a straight-line path. Additionally, using the most straight forward routing algorithms requires high computational effort for areas with large road networks. While it is believed this process can be made more performant, that task is beyond this scope of work.

99 GENERAL AND MISCELLANEOUS↗

Innovative Advanced Hydrogen Mobile Fueler (Final Technical Report)

The US Department of Energy (DOE) funded a project to design, develop, deploy and analyze the economic viability of an innovative Advanced Hydrogen Mobile Fueler (AHMF). As part of the design activity the project team defined specifications based upon vehicle requirements and compliance with specific fueling performance criteria. The AHMF was originally designed to fuel 10-20 fuel cell vehicles (FCV) per day, consistent with the requirements of the H70 fueling category. The AHMF is able to operate without remote power connections, is modular for easy transport and deployment, and can provide expanded daily capacity and multi-day operations using delivered gaseous hydrogen. Toward the end of the project, due to request from the hydrogen industry and some necessary modifications of the AHMF, the system is now modified to fuel heavy duty vehicles. The project was conducted over two primary phases, each including several key tasks, subtasks, and project milestones. Phase 1 involved the design, development, and construction of the AHMF, moving from a conceptual design through to completion of assembly and testing. In addition to reviewing several different approaches, the team chose to take a conventional station design and modify it for mobile fueling. There were several innovative components included in the design. The high pressure storage was the first to achieve a US Department of Energy (DOE) Special Permit (SP 20391) to transport high pressure hydrogen (95 MPa) in a composite cylinder. The second novel system was a liquid nitrogen (LIN) cooling system with a compact heat exchanger. Without this change, the cooling system would not be able to fit into the AHMF. Phase 2 demonstrated the AHMF by fueling fuel cell buses at a fleet in Pomona, CA. The site was chosen because a temporary fueling solution was required while a fixed station was being installed. The existing permits for the fixed station and non-public access also was a determining factor. Over two months, the buses were fueled 320 times with over 5000 kg of hydrogen from tube trailers. Existing shore power was used, and the average electrical efficiency was 0.13 kWh/kg. The average liquid nitrogen (LIN) consumption was 90.68 scf/kg. The fueling and consumption data was provided to the National Renewable Energy Laboratory (NREL) for analysis. An economic analysis was performed and will be provided in a separate report. The project was successful, but not without challenges and lessons learned. The DOT special permit led the way for the use of high pressure, composite cylinders and is being by multiple other systems and applications. The pandemic along with time for the DOT special permit approval delayed the project for years. The team also believes that hydrogen mobile fueling has a use for the industry, especially during this upcoming phase of expansion. However, it does have its limitations due to high cost to build and operate, and the same permitting challenges as a fixed station. A fully capable system at the speeds and pressures of the AHMF may not be necessary for most applications and would help reduce the cost and increase storage capacity. The on-board generator can be easily replaced with shore power or the wide range of power generation solutions in the marketplace. One of the major results of the project was the development of new code language in National Fire Protection Agency (NFPA) 2 and the International Fire Code (IFC) for on-demand mobile fueling. This will provide guidance for Authorities Having Jurisdiction (AHJ) and user on how to permit temporary fueling sites across the nation.

08 HYDROGEN↗

MOOSE ProbML: Parallelizable Probabilistic Machine Learning and Uncertainty Quantification Capabilities

The Multiphysics Object Oriented Simulation Environment (MOOSE) is a widely used open- source finite element software for performing multiphysics multiscale simulations in a massively parallel fashion. Recently, the computational team at Idaho National Laboratory (INL) has implemented Probabilistic Machine Learning (ProbML) capabilities in MOOSE—in a parallelized fashion—and enable active learning with large-scale computational models for tasks such as surrogate model development, scale bridging, forward/inverse uncertainty quantification (UQ), Bayesian optimization, etc. This presentation summarizes these developments in MOOSE along with demonstrations on several real applications relevant to nuclear energy. At the fundamental level, samplers like Monte Carlo/Latin Hypercube, variance reduction, parallelized Markov Chain Monte Carlo (MCMC) support uncertainty propagation in both forward and inverse settings. These samplers can be integrated with the Gaussian processes (GP) suite in MOOSE, which offer several variants like scalar GPs, multi-output GPs, and deep GPs, to enable active learning. These GPs can be tuned using gradient-based optimization methods like Adam and its variants or gradient-free methods like the elliptical slice sampler (a variant of MCMC adept under Gaussian settings) for more complex covariance kernels or likelihoods whose gradient computations can be cumbersome. A variety of batch acquisition functions permit parallelized evaluation of the computational model and support different learning objectives with high efficiency like Bayesian inference, global surrogate development, optimization, etc. Furthermore, libtorch integration supports training, evaluation, and re-training of neural networks and other complex machine learning models in active learning settings. The impacts of these developments are shown on several real applications: (1) nuclear fuel inverse UQ and model inadequacy assessment using the Kennedy O’Hagan framework; (2) uncertainty aware surrogate modeling for additive manufacturing to predict field quantities; (3) nuclear reactor rare events analysis; and (4) complex fluid flow prediction using a global surrogate with quantified prediction uncertainty. Finally, the outlook of MOOSE ProbML is discussed for both outer-loop and inner-loop computations in the broad view to accelerate fuels and materials qualification, address gaps in knowledge and data, and assess new reactor/fuel systems.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

"Soiling, Cleaning, and Abrasion: The Results of the Five-Year Photovoltaic Glass Coating Field Study" [Slides]

External contamination ("soiling") of the incident surface is a major limiting factor for solar technologies. A 5-year field glass coupon study was conducted to better understand external contamination and its effects; compare cleaning methods and the use of preventative coatings; and explore the abrasion resulting from cleaning to advise on accelerated abrasion testing. Test sites included the cities of Dubai (UAE), Kuwait City (Kuwait), Mesa (AZ), Mumbai (India), and Sacramento (CA). Through the 5-year cumulative study, dry brush, water spray, and wet sponge and squeegee cleaning methods were compared to no cleaning. Optical microscopy was used to obtain images, including representative color images, grayscale images for object analysis, and oblique images for coating integrity assessment. A thresholding protocol was developed to analyze and distinguish specimens using the ImageJ software. Optical performance was quantified using a spectrophotometer, including comprehensive optical characterization (transmittance, reflectance, and absorptance in addition to forward- and back-scattering). Atomic force microscopy was used to verify the abrasion damage morphology, including the width and depth of surface scratches. Analysis of the results included correlation of optical performance and particle area coverage, rank order (by coating or location), and the acceleration factor for abrasion damage. The efficacy of external cleaning was more readily distinguished from the effectiveness of antisoiling coatings. The acceleration factor for dry brush cleaning of a porous silica coating was found to be on the order of unity.

14 SOLAR ENERGY↗

RAG for FLAG: AI Assistance for a Physics Code

Artificial intelligence (AI) has quickly become an important tool in scientific research, where significant efforts are underway to develop tools that will expedite the research process. One area of particular impact is scientific software, which can be particularly complex, and therefore time consuming to learn and use effectively. AI assistants are increasingly helping to streamline the process by performing tasks such as interactively answering user questions or suggesting solutions. Los Alamos National Laboratory (LANL) develops several advanced scientific codes, such as FLAG, which can be used to run multiphysics simulations. With this study, our goal was to develop an AI assistant for FLAG that could help make the process of understanding the software and running physics simulations more efficient. To develop an AI assistant for FLAG, we used a method called retrieval-augmented generation (RAG), which is a technique that uses information from relevant data sources to enhance the accuracy of large language models (LLMs). We used the FLAG user manual and other FLAG documentation as the knowledge base for the RAG system. When a user provides a query, RAG retrieves relevant sections from the knowledge base in response, then uses those excerpts to generate grounded and contextually rich answers. We found that our AI assistant was able to provide context aware answers and source references to user queries. To evaluate performance, we developed a set of 40 benchmark questions and compared the accuracy of the responses to those of two standard LLMs without retrieval. Our AI assistant significantly outperformed the standard LLMs at answering FLAG-related questions, with an 82.5% accuracy rate, compared to 47.5% for both of the standard LLMs. This has the potential to make the process of learning and using FLAG much easier, especially for new users. Ultimately, it supports LANL’s broader mission by empowering scientists and engineers to focus more on discovery and analysis rather than on navigating complex software systems.

97 MATHEMATICS AND COMPUTING↗

The HTAP_v3.2 emission mosaic: merging regional and global monthly emissions (2000–2020) to support air quality modelling and policies

This study, performed under the umbrella of the Task Force on Hemispheric Transport of Air Pollution (TF-HTAP), responds to the need of the global and regional atmospheric modelling community of having a mosaic emission inventory of air pollutants that conforms to specific requirements: global coverage, long time series, spatially distributed emissions with high time resolution, and a high sectoral resolution. The mosaic approach of integrating official regional emission inventories based on locally reported data, with a global inventory based on a globally consistent methodology, allows modellers to perform simulations of a high scientific quality while also ensuring that the results remain relevant to policymakers. HTAP_v3.2, an ad-hoc global mosaic of anthropogenic inventories, is an update to the HTAP_v3 global mosaic inventory and has been developed by integrating official inventories over specific areas (North America, Europe, Asia including China, Japan and Korea) with the independent Emissions Database for Global Atmospheric Research (EDGAR) inventory for the remaining world regions. The results are spatially and temporally distributed emissions of SO 2 , NO x , CO, NMVOC, NH 3 , PM 10 , PM 2.5 , Black Carbon (BC), and Organic Carbon (OC), with a spatial resolution of 0.1 × 0.1° and time intervals of months and years covering the period 2000–2020 (https://doi.org/10.5281/zenodo.17086684, Crippa, 2025, https://edgar.jrc.ec.europa.eu/dataset_htap_v32, last access: 27 October 2025). The emissions are further disaggregated to 16 anthropogenic emitting sectors. This paper describes the methodology applied to develop such an emission mosaic, reports on source allocation, differences among existing inventories, and best practices for the mosaic compilation. One of the key strengths of the HTAP_v3.2 emission mosaic is its temporal coverage, enabling the analysis of emission trends over the past two decades. The development of a global emission mosaic over such long time series represents a unique product for global air quality modelling and for better-informed policy making, reflecting the community effort expended by the TF-HTAP to disentangle the complexity of transboundary transport of air pollution.

Guizzardi, Diego [European Commission, Ispra (Ital↗

Hierarchical Network Partitioning for Solution of Potential-Driven, Steady-State Nonlinear Network Flow Equations

The solution of potential-driven steady-state flow in large networks is a task which manifests in various engineering applications, such as transport of natural gas or water through pipeline networks. The resultant system of nonlinear equations depends on the network topology, and in general, there is no numerical algorithm that offers guaranteed convergence to the solution (assuming a solution exists). Some methods offer guarantees in cases where the network topology satisfies certain assumptions, but these methods fail for larger networks. On the other hand, the Newton-Raphson algorithm offers a convergence guarantee if the starting point lies close to the (unknown) solution. It would be advantageous to compute the solution of the large nonlinear system through the solution of smaller nonlinear sub-systems wherein the solution algorithms (Newton-Raphson or otherwise) are more likely to succeed. Here, this letter proposes and describes such a procedure, a hierarchical network partitioning algorithm that enables the solution of large nonlinear systems corresponding to potential-driven steady-state network flow equations.

42 ENGINEERING↗

Automation of Vulnerability and Patch Management: Information Extraction, Association, and Optimization

Vulnerability and patch management is an integral part of a robust cybersecurity program, yet it grows increasingly complex due to the sheer amount of data that must be analyzed. Particularly in Operational Technology (OT) environments, analysis must be done manually because of the lack of automated solutions. Additionally, there are many steps in this process, from the initial discovery of the vulnerability to the implementation of its remediation, and each step in the process requires different data in order to be performed effectively. In this work, we provide approaches and strategies to assist operators in industrial or OT environments throughout the vulnerability management cycle. Security advisories provide key information about mitigation strategies, or actions that can be taken when a patch is unavailable or cannot be installed. Details of these strategies are not shared in public vulnerability databases and must be found manually. We approach this problem by designing a solution to automatically identify that information within vendor security advisories and retrieve it for operator use. We start with an approach that requires domain-specific knowledge of certain frequently-seen reference websites. Next, an approach that can work on an arbitrary website but relies on certain keywords. Finally, an approach that uses Natural Language Processing (NLP) methods and does not require specific knowledge or keywords. Each of these approaches is more general than its predecessor; we demonstrate high accuracy for all approaches Advisories also often contain details of affected products in non-standard or natural language formats. While this information can be easily understood when read by an operator, the non-standard format acts as a barrier to effective automation. We provide an approach for the first step in this process: identifying vendors in security advisories and mapping them to a standard framework for representing digital assets and software products. We evaluate five established string similarity algorithms, plus one of our own design that combines string similarity and information theory, on the task of mapping vendors to their corresponding entries in the Common Platform Enumeration (CPE) repository. Our results show that our proposed metric outperforms all others. Due to the constraints on time, finances, and personnel for organizations, Large Language Models (LLMs) may seem like attractive opportunities for security operators to speed up information gathering; however, it is still not clear whether LLMs can handle vulnerability management tasks well. To answer this question, we perform an empirical study of LLMs’ ability to provide consistent, accurate information about vulnerabilities in order to guide organizations in their adoption of LLMs. We observe poor performance for all models tested, suggesting that these models are not well-suited to the consistent retrieval of accurate vulnerability information. Finally, once vulnerabilities have been identified and any additional information has been obtained, operators must decide which remediation actions to implement based on their available resources. This already-complex problem becomes even more so when we consider that a vulnerability may have multiple avenues for remediation. We formulate this scenario as two knapsack problems and provide solutions, which we then compare against several existing strategies for vulnerability prioritization seen in real operational environments.

McClanahan, Kylie↗

Artificial Intelligence for Event Reconstruction and Higgs Physics at CMS and Future Colliders

This dissertation charts a trajectory in which advances in artificial intelligence (AI) play a central role in pushing the high-energy physics frontier, complementing progress driven by higher collision energies and larger colliders. The discovery potential of the LHC and future colliders relies on accurate reconstruction of increasingly complex particle collision events. In the CMS experiment, this task is performed by the particle-flow (PF) algorithm. This dissertation presents the first implementation of a machine-learning-based particle-flow (MLPF) reconstruction in the CMS detector based on transformer architectures. In simulated top quark--antiquark pair (ttbar) events under LHC Run~3 (2023--2024) conditions, MLPF improves jet energy resolution by 10--20\% compared to standard PF for jets with transverse momentum between 30--100\GeV. Runtime performance is evaluated using simulated multijet events, with a median inference time of 20\unit{ms} per event on an NVIDIA L4 GPU, compa red to approximately 110\unit{ms} for standard PF. The MLPF algorithm is also validated on Run~3 collision data, representing the first data-validated ML-based reconstruction pipeline at any LHC experiment. We then extend MLPF toward future electron--positron colliders and introduce the first full-simulation cross-detector transfer learning workflow for PF reconstruction. The model is pre-trained on simulated events from the Compact Linear Collider detector (CLICdet) and fine-tuned on the CLIC-like detector (CLD) proposed for the Future Circular Collider (FCC). This approach achieves up to a 40\% improvement in jet energy resolution over rule-based reconstruction while reducing the required training dataset size by an order of magnitude, demonstrating the potential of AI to accelerate detector development and optimization. This dissertation also demonstrates how modern AI techniques enhance the sensitivity of LHC physics analyses. A CMS search for highly Lorentz-boosted Higgs bosons decaying to \textrm{W} boson pairs is presented, focusing on the single-lepton final state. A dedicated fine-tuning strategy for \ParT yields an approximately 70\% increase in expected sensitivity relative to the baseline model. The analysis uses proton--proton collision data at a center-of-mass energy of \ensuremath{\sqrt{s}=13\TeV} collected by CMS between 2016 and 2018, corresponding to an integrated luminosity of 138\ensuremath{\ \mathrm{fb}^{-1}}. The expected significance of the search is $1.86\sigma$, with an observed signal strength of $-0.19^{+0.48}_{-0.46}$. Finally, explainable AI techniques are applied to the MLPF and \ParticleNet algorithms using layerwise relevance propagation, showing that both models base their predictions on physically meaningful features consistent with our physics intuition. Together, these results demonstrate how advanced AI methods can enhance reconstruction, analysis sensitivity, and interpretability, shaping the next era of experimental parti cle physics.

Mokhtar, Farouk [UC, San Diego]↗

Deep nonparametric estimation of operators between infinite dimensional spaces

Learning operators between infinitely dimensional spaces is an important learning task arising in machine learning, imaging science, mathematical modeling and simulations, etc. This paper studies the nonparametric estimation of Lipschitz operators using deep neural networks. Non-asymptotic upper bounds are derived for the generalization error of the empirical risk minimizer over a properly chosen network class. Under the assumption that the target operator exhibits a low dimensional structure, our error bounds decay as the training sample size increases, with an attractive fast rate depending on the intrinsic dimension in our estimation. Our assumptions cover most scenarios in real applications and our results give rise to fast rates by exploiting low dimensional structures of data in operator estimation. We also investigate the influence of network structures (e.g., network width, depth, and sparsity) on the generalization error of the neural network estimator and propose a general suggestion on the choice of network structures to maximize the learning efficiency quantitatively.

97 MATHEMATICS AND COMPUTING↗

Biases in preconstruction estimates of wind plant annual energy production

Estimating the energy yield of a wind plant during the preconstruction phase is a historically difficult task, even with industry improvements in these estimations. We build on prior research comparing the realized energy production of wind plants and their estimated annual energy production P50 values (median energy production), using owner-provided energy production and losses. We produced similar results to prior studies but with a slightly increasing bias of overestimating median energy production (a bias between realized and estimated energy production of −7.4 % to −6.6 %, depending on the scenario, as opposed to −6.7 % to −5.5 % from earlier studies). In addition to assessing annual energy production P50 bias, we compared both the 1-year and the long-term annual energy production P90 and uncertainty energy yield assessment estimates to the observed long-term-corrected energy production. We found that neither the energy yield assessment uncertainty nor the P90 is conservative enough compared to the observed distribution of prediction errors, suggesting significant room for improvement in the energy yield assessment process.

17 WIND ENERGY↗

Chemical Process Safety at TRISO-Based, Metal-Based, and Salt-Based Fuel Fabrication Facilities: Technical Assessment and Guidance Assessment

As part of efforts to prepare for potential and ongoing safety reviews for licensing of advanced non-light-water reactor fuel cycles, the U.S. Nuclear Regulatory Commission (NRC) tasked Pacific Northwest National Laboratory to prepare an assessment on the state of knowledge of potential chemical processes at fuel cycle facilities supporting the front end of these fuel cycles, and to assess the associated regulatory guidance. This report provides a technical assessment of chemical process safety considerations to support NRC licensing reviews of fabrication processes for tri-structural isotropic (TRISO) based, metallic-based, and salt-based fuels. The assessments involved collecting publicly available information on the fuel fabrication processes to (i) identify the operational process steps, characteristics and chemicals involved, (ii) identify the physical safety considerations and health safety considerations during licensing reviews of the various process steps, and (iii) collect information to support assessments of severity of accidents and potential mitigative measures to be implemented. The assessment provides a foundational basis on chemical process safety considerations for advanced fuel fabrication activities, although it is recognized that licensing reviews may necessitate design-specific considerations. The specific conditions under which chemical hazards emerge will require process-specific considerations, highlighting the importance of process-informed interpretation. The assessment also determined that exposure guidelines and limits to assess the consequences of acute exposures are limited for some chemicals, although alternative limits and supplementary information from databases or safety data sheets provide sufficient information to evaluate consequences of acute exposures. In addition, it was identified that metallic and salt fuel fabrication processes may involve beryllium, which is an exposure hazard. The regulatory framework for the licensing of advanced fuel cycle facilities, per 10 CFR Part 70 Domestic Licensing of Special Nuclear Material, is deemed robust and flexible to address the chemical safety considerations in this report. A review was conducted on various regulatory guidance and technical basis documents. This included reviewing NUREG-1520, Revision 2, Standard Review Plan for Fuel Cycle Facilities License Applications – Final Report and the process descriptions in Appendix A of NUREG/CR-6410, Nuclear Fuel Cycle Facility Accident Analysis Handbook, to address advanced fuel types. As new fuels will involve process-specific chemical uses, process-specific considerations are provided in this report. Additionally, it is noted that the U.S. Department of Energy protective action criteria database includes Temporary Emergency Exposure Limits (TEELs) for process-specific chemicals. This report provides technical information to support chemical safety assessments of new advanced fuel cycle facilities and identifies technical and safety information to support licensing reviews. No regulatory barriers were identified for the licensing of advanced fuel cycle facilities.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Machine-Learning-Based Mapping and Modeling of Solar Energy with Ultra-High Spatiotemporal Granularity

Despite the rapid growth of solar energy, we still lack a dynamic, high-fidelity database that tracks the spatiotemporal variations of solar PVs and their associated infrastructures across different places at a spatially resolved scale. The absence of such data presents a barrier to various applications such as solar PV growth projection, solar energy integration, solar incentive design, and climate risk assessment. In this project, we aim to bridge this gap by developing AI-based algorithms to extract granular information about solar PV installations and their associated infrastructures (i.e., distribution grids) from widely available unstructured data like remote sensing images and street views. As a result, we have built the Solar Energy Atlas, a fine-grained, large-scale geospatial overlay of distributed solar PVs and distribution grids. On top of it, we have advanced the understanding of solar adoption and distribution grid vulnerability to climate-induced extremes. Our major contributions can be summarized as follow: (1) By developing new AI algorithms, we have built the most comprehensive solar PV spatiotemporal database covering the entire US. This is the first time we obtained the exact GPS locations, size, subtype, and installation year information for rooftop solar PVs across the US. This database can be used for solar PV growth projection, solar energy integration, solar energy policy analysis and design, and spatially-resolved climate risk assessment. (2) Leveraging this database, we have uncovered the socioeconomic driving factors that are correlated with earlier onset of solar adoption and higher saturated adoption levels. We have identified the heterogeneity in the effects of different types of financial incentives on solar adoption and provided implications for tailoring incentive design based on local income levels to promote equitable solar adoption. (3) We have developed a distribution grid GIS mapping algorithm which can obtain granular geospatial and topology information about distribution grids using multi-modal open data, reducing the dependency on hard-to-obtain smart meter data of conventional approaches. It shows effectiveness in both the U.S. and Sub-Saharan Africa. Using this algorithm, we have uncovered the non-uniform vulnerability of distribution grids to wildfires in California in the aspects of undergrounding protection and Distributed Energy Resources (DER) preparedness. This has provided important implications for improving the affordability and equity of grid adaptation approaches. (3) We have made our produced database publicly available and provided user-friendly interface to enable various stakeholders and the general public to interact with the data. We have also integrated the produced data into the Data Commons platform to enable the public to access the data and correlate it with other location-specific characteristics simply using natural language as queries. The impact of our project is three-fold: (1) New algorithms for mapping solar PVs and distribution grids across space and time, which are open source to facilitate researchers and industry; (2) New databases of solar PVs and distribution grids that have been made publicly available for engineering, social, and policy applications; (3) New understandings and actionable insights on the potential approaches to promoting solar adoption and reducing energy infrastructure vulnerabilities. In this report, we start by discussing the project background and motivation (section 5), followed by the overview of project objectives (section 6). Results and discussion for each task are presented in section 7. Significant accomplishments are summarized in section 8. This report will be concluded by discussing the paths forwards (section 9), products (section 10), and team roles (section 11).

14 SOLAR ENERGY↗

GeoThermalCloud: Cloud Fusion of Big Data and Multi-Physics Models using Machine Learning for Discovery, Exploration, and Development of Hidden Geothermal Resources

The primary goals of this project are exploring hidden geothermal resources in the U.S.A. and designing profitable enhanced geothermal systems (EGS). Many processes and parameters control geothermal exploration and energy production from geothermal fields. Diverse datasets (e.g., geology, geochemistry, geophysics, satellite, airborne geophysics) are available to help characterize subsurface geothermal conditions. Sparse and multi-scale characteristics of these datasets prohibit properly leveraging these datasets for geothermal exploration and profitable EGS design. Recent advancements in machine learning (ML) promise to resolve these issues. The tremendous challenges and risks of geothermal exploration and production bring the demand for novel ML methods and tools that can (1) analyze large field datasets, (2) assimilate model simulations (large inputs and outputs), (3) process sparse datasets, (4) perform transfer learning (between sites with different exploratory levels), (5) extract hidden geothermal signatures in the field and simulation data, (6) label geothermal resources and processes, (7) identify high-value data acquisition targets, and (8) guide geothermal exploration and production by selecting optimal exploration, production, and drilling strategies. To address these necessities, ML-based geothermal resources exploration and enhanced geothermal systems (EGS) design tools have been developed. The exploration tool is called GeoThermalCloud and EGS design tool is called GeoDT-ML. GeoThermalCloud (https://github.com/SmartTensors/GeoThermalCloud.jl) utilizes a LANL unsupervised ML platform called SmartTensors (https://tensors.lanl.gov/) to automate data analyses and interpretations by extracting hidden signatures to identify geothermal prospects. Also, it enables the identification of critical measurements needed to identify geothermal resource signatures. Alternatively, GeoDT-ML (https://github.com/SmartTensors/GeoThermalCloud.jl/tree/master/EGS) is an ML-based alternative to GeoDT (https://github.com/GeoDesignTool/GeoDT.git), a fast, simplified multi-physics solver to evaluate EGS project designs in uncertain geologic systems. GeoDT-ML leverages recent advances in deep learning and high-performance computing. It is a faster and simpler version of GeoDT. To make this project a success, we used capabilities of LANL, PNNL, Google, Stanford, and Julia Computing. We analyzed eight datasets of the U.S.A. using GeothermalCloud and demonstrated potential highly prospective geothermal resources and identified key factors defining highly prospective sites. The first data set includes 44 locations in southwest New Mexico and 18 geological, hydrogeological, geophysical, geothermal, geochemical attributes. We defined low- and medium-temperature hydrothermal systems and discovered a new highly prospective site. The second data set analyzed 18 shallow water chemistry attributes at 14,342 locations in the Great Basin. It demarcated modestly, moderately, and highly prospective sites including key attributes for each type of prospectivity. The third data set analyzed Utah FORGE data including satellite (InSAR), geophysical (gravity, seismic), geochemical, and geothermal attributes. Here, we performed prospectivity analysis to identify future drilling locations using geological, geochemical, and geophysical attributes. Maps of temperature at depth and heat flow are constructed based on the available data. Prospectivity maps were generated, and drilling locations were proposed for future geothermal field exploration. The fourth data set analyzed 21 attributes at 120 locations in Tularosa Basin, New Mexico; data comes from past play fairway analyses in this region. ML analyses identified geothermal signatures associated with modestly, moderately, and highly hydrothermal systems. We also defined dominant attributes and spatial distribution of the geothermal signatures. The fifth, sixth, seventh, and eighth datasets include Tohatchi Springs, New Mexico, Hawaii, Brady site, Nevada, and EGS Collab, respectively. Moreover, we coupled GeothermalCloud and magnetotellurics data to pinpoint drilling locations for developing geothermal projects in the Tularosa Basin, New Mexico. GeothermalCloud found potential prospective locations for geothermal resources near White Sands Missile Range and McGregor Range at Fort Bliss. Magnetotellurics data determined the potential depth (~1800m) of geothermal prospects at McGregor Range based on apparent resistivity structures/layers in the subsurface. The McGregor Range consists of three resistivity layers and two resistivity structures. Magnetotellurics data also helps identify that the western portion of the McGregor Range has thick and low-resistivity earth materials. The low resistivity to the west is most likely for a fault system. Assuming temperature is consistent with a geothermal reservoir, the west-central part of the McGregor Range has the highest geothermal potential because of the increase in porosity and associated permeability attributed to the interpreted fault system. Also, we devised a coupling strategy between a process model and GeothermalCloud to characterize hydrogeological conditions and geothermal conditions, respectively. The process model characterizes hydrogeological and geothermal conditions on highly prospective geothermal sites provided by GeothermalCloud. We developed a physics-informed neural network (PINN) version of the Burns equation that can be easily coupled with GeothermalCloud. Furthermore, we performed an optimal design decision maximizing the economic value of an EGS power plant. This study optimized the range of well spacing between injection and production wells maximizing net present value in dollars (NPV). For this task, we used the GeoDT to simulate the Utah FORGE EGS development cycle from the initial well design to the end of production. Next, we accomplished another crucial task, which is predicting permeability of geothermal reservoirs. Predicting permeability of geothermal reservoirs is a non-trivial task because of huge computational runtime of simulation and lack of measurements. To avoid these limitations, we used easy-to-measure chemical concentrations in the subsurface as measurement data and convolutional neural network based ML model of a high-fidelity model. Next, we predicted permeability using Markov chain Monte Carlo simulation. We found that Markov chain Monte Carlo simulation predicts permeability with a high certainty if the prediction zone in the simulation area has chemical concentration data. Finally, we analyzed the DOE funded INGENIOUS and GeoDAWN projects data. For discovering hidden geothermal systems in the Great Basin, the INGENIOUS project accumulated old data, collected new data, and released them in 2022. The dataset includes a total of 24 geological, geophysical, and geochemical attributes. Data resolution and scale significantly vary prohibiting an appropriate usage. To avoid such limitations, we brought all data in the same resolution and scale by applying the inverse distance weighting interpolation technique for predicting data in unsampled locations. Subsequently, we analyzed LiDAR data of the GeoDAWN project. We received data in tiles format. The DOE’s overarching goal is to use ML on LiDAR data for finding favorable geological structures (e.g., step up faults in Brady, Nevada). To serve the purpose, we need to label favorable geologic structures that correspond to LiDAR data. We wrote an algorithm to label the LiDAR data with the favorable geologic structures.

15 GEOTHERMAL ENERGY↗

Energy metric prediction for double insertion mutants via the RoseNet deep learning framework

Studying the structural and functional implications of protein mutations is an important task in computational biology and bioinformatics. We leverage our previously proposed RoseNet neural network architecture to predict energy metrics of proteins with double amino acid insertions or deletions (InDels). We train models on previously generated benchmark datasets containing the exhaustive double InDel mutations for three proteins, as well as an additional three proteins for which ∼145k random mutants, each with two InDels, have been generated. We expand on our previous work by evaluating three additional proteins and analyzing domain features that impact the prediction capabilities of RoseNet. These features include InDels into secondary structures and the solvent accessible surface area (SASA) scores of the residues. We uncover further evidence to support that RoseNet has a higher proficiency of generalizing to unseen residue combinations than unseen insertion positions. We also observe that RoseNet produces higher-quality predictions when inserting into a β-sheet over an α-helix. Additionally, when the insertions fall in an area of high SASA, RoseNet often displays better performance than inserting into areas of low SASA.

59 BASIC BIOLOGICAL SCIENCES↗