Search NASASearch

SEARCH · Search NASA

Results for “successive training”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

ReLIC: Full-Scale Realization of Reinforcement Learning for Infrastructure Control

Prior efforts have shown that deep reinforcement learning (DRL) may provide a new method for controlling networked power systems. Though successful, prior approaches have not yet demonstrated their behavior on systems of realistic scale. This effort examined multiple theoretical and technical approaches to allow a DRL model to operate over a system of 2,000 buses or more. We find that allowing the DRL models to run training episodes in parallel provides near limitless efficiency gains, allowing us to train successful agents to behave on our Kuramoto transmission model of up to 4,000 buses. We further show that we can expand our PowerWorld DRL implementation to systems of up to 25 buses but struggle to go beyond this limit due to PowerWorld’s inability to run multiple instances at once. Finally, we examine a multi-agent approach and find that it performs as well if not better than our existing centralized approach.

97 MATHEMATICS AND COMPUTING

Combining High-Throughput Experiments and Active Learning to Characterize Deep Eutectic Solvents

The high tunability of deep eutectic solvents (DESs) stems from the ease of changing their precursors and relative compositions. However, measuring the physicochemical properties across large composition and temperature ranges, necessary to properly design target-specific DESs, is tedious and error-prone and represents a bottleneck in the advancement and scalability of DES-based applications. As such, active learning (AL) methodologies based on Gaussian processes (GPs) were developed in this work to minimize the experimental effort necessary to characterize DESs. Owing to its importance for large-scale applications, the reduction of DES viscosity through the addition of a low-molecular-weight solvent was explored as a case study. A high-throughput experimental screening was initially performed on nine different ternary DESs. Then, GPs were successfully trained to predict DES viscosity from its composition and temperature, showcasing the ability of these stochastic, nonparametric models to accurately describe the physicochemical properties of complex mixtures. Finally, the ability of GPs to provide estimates of their own uncertainty was leveraged through an AL framework to minimize the number of data points necessary to obtain accurate viscosity modes. This led to a significant reduction in data requirements, with many systems requiring only five independent viscosity data points to be properly described.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Identification of carbohydrate gene clusters obtained from in vitro fermentations as predictive biomarkers of prebiotic responses

Prebiotic fibers are non-digestible substrates that modulate the gut microbiome by promoting expansion of microbes having the genetic and physiological potential to utilize those molecules. Although several prebiotic substrates have been consistently shown to provide health benefits in human clinical trials, responder and non-responder phenotypes are often reported. These observations had led to interest in identifying, a priori, prebiotic responders and non-responders as a basis for personalized nutrition. In this study, we conducted in vitro fecal enrichments and applied shotgun metagenomics and machine learning tools to identify microbial gene signatures from adult subjects that could be used to predict prebiotic responders and non-responders. Using short chain fatty acids as a targeted response, we identified genetic features, consisting of carbohydrate active enzymes, transcription factors and sugar transporters, from metagenomic sequencing of in vitro fermentations for three prebiotic substrates: xylooligosacharides, fructooligosacharides, and inulin. A machine learning approach was then used to select substrate-specific gene signatures as predictive features. These features were found to be predictive for XOS responders with respect to SCFA production in an in vivo trial. Our results confirm the bifidogenic effect of commonly used prebiotic substrates along with inter-individual microbial responses towards these substrates. We successfully trained classifiers for the prediction of prebiotic responders towards XOS and inulin with robust accuracy (≥ AUC 0.9) and demonstrated its utility in a human feeding trial. Overall, the findings from this study highlight the practical implementation of pre-intervention targeted profiling of individual microbiomes to stratify responders and non-responders.

59 BASIC BIOLOGICAL SCIENCES

Building collaboration to advance our understanding of regional climate impacts of dust in California's San Joaquin Valley

This project successfully achieved its central objective of building collaborative research capabilities at UC Merced, a Hispanic-Serving Institution, to advance understanding of the regional climate impacts of dust in California's San Joaquin Valley. Through strategic partnerships with three DOE national laboratories (PNNL, LLNL, and LBNL), we developed critical expertise in the Energy Exascale Earth System Model (E3SM) and Atmospheric Radiation Measurement (ARM) facilities. Among the project's scientific contributions, one key publication includes demonstrating that fallowed agricultural lands are the primary source of anthropogenic dust in California's Central Valley, with dust activities increasing substantially between 2008 and 2022, in correlation with drought severity and expanded fallowed land coverage. This finding suggests that current climate models, including E3SM, likely underestimate the dust burden due to inadequate representation of agricultural land-use changes. Beyond the scientific contributions, the project successfully trained a PhD student, established ongoing collaborations resulting in multiple manuscripts in preparation, and positioned UC Merced to participate in the DUSTIEAIM campaign for 2026-2027, thereby building sustainable research capacity while addressing climate science questions directly relevant to the California Central Valley.

54 ENVIRONMENTAL SCIENCES

Gravitational Lenses in UNIONS and Euclid (GLUE). I. A Search for Strong Gravitational Lenses in UNIONS with Subaru, CFHT, and Pan-STARRS Data

We present the results of our pipeline for discovering strong gravitational lenses in the ongoing Ultraviolet Near-Infrared Optical Northern Survey (UNIONS). We successfully train a deep residual convolutional neural network based on CMU-Deeplens architecture, which is designed to detect strong lenses in ground-based imaging surveys. We train on images of real strong lenses and deploy on a sample of 8 million galaxies in areas with full coverage in the g, r, and i filters—the first multiband search for strong gravitational lenses in UNIONS. Following human inspection and grading, we report the discovery of a total of 1346 new strong-lens candidates, of which 146 are Grade A, 199 are Grade B, and 1001 are Grade C. Of these candidates, 283 have lens galaxy spectroscopic redshifts from the Sloan Digital Sky Survey, and an additional 297 have them from the Dark Energy Spectroscopic Instrument Data Release 1. We find 15 of these systems display evidence of both lens and source galaxy redshifts in spectral superposition. We also report the spectroscopic confirmation of seven lensed sources in high-quality systems, all with z > 2.1, using the Keck Near Infrared Echellette Spectrograph and the Gemini Near-Infrared Spectrograph.

Storfer, Christopher J. [University of Hawaii, Hon

Stacked networks improve physics-informed training: Applications to neural networks and deep operator networks

Physics-informed neural networks and operator networks have shown promise for effectively solving equations modeling physical systems. However, these networks can happen to be difficult or impossible to train accurately. Here, we present a novel multifidelity framework for stacking physics-informed neural networks and operator networks that facilitates training. We successively build a chain of networks, where the output at one step can act as a low-fidelity input for training a longer chain, gradually increasing the expressivity of the learnt model. The equations imposed at each step of the iterative process can be the same or different (akin to simulated annealing). The iterative (stacking) nature of the proposed method allows us to learn progressively features of a solution which could have been hard to learn directly. Through benchmark problems including a nonlinear pendulum, the wave equation, and the viscous Burgers equation, we show how stacking can be used to improve the accuracy and reduce the required size of physics-informed neural networks and operator networks.

97 MATHEMATICS AND COMPUTING

Remote quantification of Cm(III) and HNO 3 by fluorescence spectroscopy and chemometrics

A unique approach to remotely quantify Cm(III) (0–100 µg mL −1 ) in HNO 3 (1–12 M) using steady-state laser fluorescence spectroscopy and multivariate regression models was developed. Photoluminescence is amenable to remote measurements using fiber-optic cables and is sensitive to numerous lanthanide and actinide species. In-line measurements can provide feedback to support complex processing in harsh environments (e.g., hot cells) to help guide and optimize radiochemical separations. In this work, Cm(III) spectra were acquired remotely in a glove box as a function of HNO 3 concentration to better understand spectral characteristics and evaluate the utility of multivariate regression models in this system. Furthermore, the Cm(III) fluorescence peak shape, width, position, and intensity changed significantly as a function of HNO 3 concentration, likely because of the displacement of emission quenching inner-sphere water molecules and complexation with nitrate ions. Despite significant covariance and nonlinearity in the data, a D-optimal design strategy successfully minimized training set sample size and was used to build effective partial least squares regression models for Cm(III) and HNO 3 concentrations without a priori knowledge of solution conditions. Chemometrics for modeling complex fluorescence spectra are promising and may find widespread applicability for online analysis in numerous chemical systems found in the nuclear field.

Actinide

Two-Scale Neural Networks for Partial Differential Equations with Small Parameters

We propose a two-scale neural network method for solving partial differential equations (PDEs) with small parameters using physics-informed neural networks (PINNs). We directly incorporate the small parameters into the architecture of neural networks. The proposed method enables solving PDEs with small parameters in a simple fashion, without adding Fourier features or other computationally taxing searches of truncation parameters. Various numerical examples demonstrate reasonable accuracy in capturing features of large derivatives in the solutions caused by small parameters.

97 MATHEMATICS AND COMPUTING

Anomaly Detection In DUNE Using AI/ML

We designed and built an AI/ML model to detect anomalies in the DUNE far detector data. The model has been trained on simulated radiological background (rbkg) data, which is the major background for supernova burst neutrinos. The trained model was evaluated on both new samples of radiological backgrounds and supernova burst neutrino events in the elastic scattering and charged current interaction channels. We found that the trained model can successfully identify supernova burst neutrino events as anomalies while identifying radiological backgrounds as nominal events.

Novello, Eric [Unlisted, US]

Neural Network Prediction of Strong Lensing Systems with Domain Adaptation and Uncertainty Quantification

Modeling strong gravitational lenses is computationally expensive for the complex data from modern and next-generation cosmic surveys. Deep learning has emerged as a promising approach for finding lenses and predicting lensing parameters, such as the Einstein radius. Mean-variance Estimators (MVEs) are a common approach for obtaining aleatoric (data) uncertainties from a neural network prediction. However, neural networks have not been demonstrated to perform well on out-of-domain target data successfully - e.g., when trained on simulated data and applied to real, observational data. In this work, we perform the first study of the efficacy of MVEs in combination with unsupervised domain adaptation (UDA) on strong lensing data. The source domain data is noiseless, and the target domain data has noise mimicking modern cosmology surveys. We find that adding UDA to MVE increases the accuracy on the target data by a factor of about two over an MVE model without UDA. Including UDA also permits much more well-calibrated aleatoric uncertainty predictions. Advancements in this approach may enable future applications of MVE models to real observational data.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Finding Real Uncertainties From Physical Simulations

Modeling strong gravitational lenses is computationally expensive for the complex data from modern and next-generation cosmic surveys. Deep learning has emerged as a promising approach for finding lenses and predicting lensing parameters, such as the Einstein radius. Mean-variance Estimators (MVEs) are a common approach for obtaining aleatoric (data) uncertainties from a neural network prediction. However, neural networks have not been demonstrated to perform well on out-of-domain target data successfully - e.g., when trained on simulated data and applied to real, observational data. In this work, we perform the first study of the efficacy of MVEs in combination with unsupervised domain adaptation (UDA) on strong lensing data. The source domain data is noiseless, and the target domain data has noise mimicking modern cosmology surveys. We find that adding UDA to MVE increases the accuracy on the target data by a factor of about two over an MVE model without UDA. Including UDA also permits much more well-calibrated aleatoric uncertainty predictions. Advancements in this approach may enable future applications of MVE models to real observational data.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Anomaly Detection in DUNE FD LArTPC Readouts

We designed and built an AI/ML model to detect anomalies in the DUNE far detector data. The model has been trained on simulated radiological background (rbkg) data, which is the major background for supernova burst neutrinos that we want to detect as anomaly in thios work. The trained model was evaluated on both new samples of radiological backgrounds and supernova burst neutrino events in the elastic scattering and charged current interaction channels. We found that the trained model can successfully identify supernova burst neutrino events as anomalies while identifying radiological backgrounds as nominal events.

Novello, Eric [Fermilab]

Quantum Computing Strategy 2026

Quantum computing (QC) is a rapidly maturing technology with the potential for revolutionary impacts on stockpile stewardship science and national security. Recent developments in fault-tolerant architectures have compressed vendor roadmaps, and predictions of a production-ready quantum computer by the mid-2030s are becoming increasingly credible. This strategy provides a roadmap for integrating QC into the Advanced Simulation and Computing (ASC) program by investing in four strategic focus areas: 1. Develop Capabilities in Mission-Relevant Quantum Applications: ASC will prioritize developing quantum-ready applications in mission areas that have shown significant promise for quantum advantage, including simulations of materials in extreme environments, nuclear dynamics, solving linear and nonlinear partial differential equations, and uncertainty quantification. These applications directly support stockpile stewardship science and modernization objectives. 2. Conduct R&D in Algorithms, Software, and Hardware: Sustained research into quantum algorithms, robust software tools, and quantum hardware is essential. ASC will develop efficient quantum algorithms; invest in quantum compilers, debuggers, and performance tools; and explore specialized quantum hardware tailored to NNSA’s unique requirements. 3. Engage with Vendors and Partners: Early and active collaboration with commercial quantum hardware vendors and academic partners is critical. Through testbeds, co-design agreements, and quantum demonstration facilities, ASC will influence hardware design, gain early access to emerging technologies, and ensure that quantum platforms evolve to meet mission needs. 4. Build Knowledge, Experience, and Workforce: Expanding and upskilling the quantum-trained workforce is essential to long-term success. This includes hiring, internal training, university outreach, and postdoctoral support to ensure ASC maintains the expertise required to operate, program, and integrate quantum systems as they become available. While quantum computing will never replace classical computing, it has the potential to solve certain problems with speed and accuracy that would be unachievable using any conceivable classical high-performance computing (HPC) system. By investing strategically in QC, ASC will help propel the emergent QC industry, maintain U.S. technological leadership, ensure mission readiness, and position itself to rapidly adopt quantum technologies as they mature.

97 MATHEMATICS AND COMPUTING

Tula: Optimizing Time, Cost, and Generalization in Distributed Large-Batch Training

Distributed training increases the number of batches processed per iteration either by scaling-out (adding more nodes) or scaling-up (increasing the batch-size). However, the largest configuration does not necessarily yield the best performance. Horizontal scaling introduces additional communication overhead, while vertical scaling is constrained by computation cost and device memory limits. Thus, simply increasing the batch-size leads to diminishing returns: training time and cost decrease initially but eventually plateaus, creating a knee-point in the time/cost vs. batch-size pareto curve. The optimal batch-size therefore depends on the underlying model, data and available compute resources. Large batches also suffer from worse model quality due to the well-known “generalization gap”. In this paper, we present Tula, an online service that automatically optimizes time, cost, and convergence quality for large-batch training of convolutional models. It combines parallel-systems modeling with statistical performance prediction to identify the optimal batchsize. Tula predicts training time and cost within 7.5−14% error across multiple models, and achieves up to 20× overall speedup and improves test accuracy by ≈9% on average over standard large-batch training on various vision tasks, thus successfully mitigating the generalization gap and accelerating training at the same time.

Tyagi, Sahil [ORNL] (ORCID:0009000783144745)

The Importance of Being Adaptable: An Exploration of the Power and Limitations of Domain Adaptation for Simulation-Based Inference with Galaxy Clusters

The application of deep machine learning methods in astronomy has exploded in the last decade, with new models showing remarkably improved performance on benchmark tasks. Not nearly enough attention is given to understanding the models' robustness, especially when the test data are systematically different from the training data, or "out of domain." Domain shift poses a significant challenge for simulation-based inference, where models are trained on simulated data but applied to real observational data. In this paper, we explore domain shift and test domain adaptation methods for a specific scientific case: simulation-based inference for estimating galaxy cluster masses from X-ray profiles. We build datasets to mimic simulation-based inference: a training set from the Magneticum simulation, a scatter-augmented training set to capture uncertainties in scaling relations, and a test set derived from the IllustrisTNG simulation. We demonstrate that the Test Set is out of domain in subtle ways that would be difficult to detect without careful analysis. We apply three deep learning methods: a standard neural network (NN), a neural network trained on the scatter-augmented input catalogs, and a Deep Reconstruction-Regression Network (DRRN), a semi-supervised deep model engineered to address domain shift. Although the NN improves results by 17% in the Training Data, it performs 40% worse on the out-of-domain Test Set. Surprisingly, the Scatter-Augmented Neural Network (SANN) performs similarly. While the DRRN is successful in mapping the training and Test Data onto the same latent space, it consistently underperforms compared to a straightforward Yx scaling relation. These results serve as a warning that simulation-based inference must be handled with extreme care, as subtle differences between training simulations and observational data can lead to unforeseen biases creeping into the results.

Ntampaka, Michelle [Baltimore, Space Telescope Sci

Fine-tuning machine-learned particle-flow reconstruction for new detector geometries in future colliders

We demonstrate transfer learning capabilities in a machine-learned algorithm trained for particle-flow reconstruction in high energy particle colliders. This paper presents a cross-detector fine-tuning study, where we initially pretrain the model on a large full simulation dataset from one detector design, and subsequently fine-tune the model on a sample with a different collider and detector design. Specifically, we use the Compact Linear Collider detector (CLICdet) model for the initial training set and demonstrate successful knowledge transfer to the CLIC-like detector (CLD) proposed for the Future Circular Collider in electron-positron mode. We show that with an order of magnitude less samples from the second dataset, we can achieve the same performance as a costly training from scratch, across particle-level and event-level performance metrics, including jet and missing transverse momentum resolution. Furthermore, we find that the fine-tuned model achieves comparable performance to the traditional rule-based particle-flow approach on event-level metrics after training on 100,000 CLD events, whereas a model trained from scratch requires at least 1 million CLD events to achieve similar reconstruction performance. To our knowledge, this represents the first full-simulation cross-detector transfer learning study for particle-flow reconstruction. These findings offer valuable insights towards building large foundation models that can be fine-tuned across different detector designs and geometries, helping to accelerate the development cycle for new detectors and opening the door to rapid detector design and optimization using machine learning.

43 PARTICLE ACCELERATORS

Multilabel proportion prediction and out-of-distribution detection on gamma spectra of short-lived fission products

In the machine learning problem of multilabel classification, the objective is to determine for each test instance which classes the instance belongs to. In this work, we consider an extension of multilabel classification, called multilabel proportion prediction, in the context of radioisotope identification (RIID) using gamma spectra data. We aim to not only predict radioisotope proportions, but also identify out-of-distribution (OOD) spectra. We achieve this goal by viewing gamma spectra as discrete probability distributions, and based on this perspective, we develop a custom semi-supervised loss function that combines a traditional supervised loss with an unsupervised reconstruction error function. Our approach was motivated by its application to the analysis of short-lived fission products from spent nuclear fuel. In particular, we demonstrate that a neural network model trained with our loss function can successfully predict the relative proportions of 37 radioisotopes simultaneously. The model trained with synthetic data was then applied to measurements taken by Pacific Northwest National Laboratory (PNNL) to conduct analysis typically done by subject-matter experts. Here, we also extend our approach to successfully identify when measurements are OOD, and thus should not be trusted, whether due to the presence of a novel source or novel proportions.

Anomaly detection

Efficient online quantum circuit learning with no upfront training

Optimization is a promising candidate for studying the utility of variational quantum algorithms (VQAs). However, evaluating cost functions using quantum hardware introduces runtime overheads that limit exploration. Surrogate-based methods can reduce calls to a quantum computer, yet existing approaches require hyperparameter pre-training and have been tested only on small problems. Here, we show that surrogate-based methods can enable successful optimization at scale, without pre-training, by using radial basis function interpolation (RBF) to construct an adaptive, hyperparameter-free surrogate. Using the surrogate as an acquisition function drives hardware queries to the vicinity of the true optima. For 16-qubit random 3-regular Max-Cut instances with the Quantum Approximate Optimization Algorithm (QAOA), our method outperforms state-of-the-art approaches, without considering their upfront training costs. Furthermore, we successfully optimize QAOA circuits for 127-qubit random Ising models on an IBM processor using 10 4 −10 5 measurements. Strong empirical performance demonstrates the promise of automated surrogate-based learning for large-scale VQA applications.

97 MATHEMATICS AND COMPUTING