Search NASA⌕ Search

SEARCH · Search NASA

Results for “Loss function”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Automated RF Phase Adjustment for Beam Stabilization in the Fermilab Linac

The Fermilab Linac experiences longitudinal beam phase drift, leading to increased particle loss, conventionally cor- rected through labor-intensive manual RF adjustments. This project explores machine learning-based automation for drift correction, employing a prototype-based classification approach. Our model utilizes a 34-dimensional feature set (RF settings and BPM readings) and leverages a 7x27 response matrix for system modeling. To overcome limited real-world data, we generate synthetic data, enhancing model training and generalizability. Custom loss functions, including a sur- rogate energy-consistent loss and a temporal smoothness constraint, ensure physically plausible drift predictions. The goal is a robust system for autonomous phase adjustments, ensuring stable beam acceleration and reduced manual intervention.

Chichili, R. R. [U. Illinois, Chicago]↗

On the Stability Analysis of Astrophysical Cooling Functions

To model the temperature evolution of optically thin astrophysical environments at MHD scales, radiative and collisional cooling rates are typically either pretabulated or fit into a functional form and then input into MHD codes as a radiative loss function. Thermal balance requires estimates of the analogous heating rates, which are harder to calculate, and due to uncertainties in the underlying dissipative heating processes these rates are often simply parameterized. The resulting net cooling function defines an equilibrium curve that varies with density and temperature. Such cooling functions can make the gas prone to thermal instability (TI), which will cause departures from equilibrium. There has been no systematic study of thermally unstable parameter space for nonequilibrium states. Motivated by our recent finding that there is a related linear instability, catastrophic cooling instability, that can dominate over TI, here we carry out such a study. We show that Balbus instability criteria for TI can be used to define a critical cooling rate, Λ c , that permits a nonequilibrium analysis of cooling functions through the mapping of TI zones. We furthermore illustrate how thermal conduction modifies the shape of TI zones. Upon applying a Λ c -based stability analysis to coronal loop simulations, we find that loops undergoing periodic episodes of coronal rain formation are linearly unstable to catastrophic cooling instability, while TI is stabilized by thermal conduction.

79 ASTRONOMY AND ASTROPHYSICS↗

Prediction of laser beam spatial profiles in a high-energy laser facility by use of deep learning

We adapt the significant advances achieved recently in the field of generative artificial intelligence/machine-learning to laser performance modeling in multipass, high-energy laser systems with application to high-shot-rate facilities relevant to inertial fusion energy. Advantages of neural-network architectures include rapid prediction capability, data-driven processing, and the possibility to implement such architectures within future low-latency, low-power consumption photonic networks. Four models were investigated that differed in their generator loss functions and utilized the U-Net encoder/decoder architecture with either a reconstruction loss alone or combined with an adversarial network loss. We achieved inference times of 1.3 ms for a 256 × 256 pixel near-field beam with errors in predicted energy of the order of 1% over most of the energy range. It is shown that prediction errors are significantly reduced by ensemble averaging the models with different weight initializations. These results suggest that including the temporal dimension in such models may provide accurate, real-time spatiotemporal predictions of laser performance in high-shot-rate laser systems.

47 OTHER INSTRUMENTATION↗

Imaging nanoscale carrier, thermal, and structural dynamics with time-resolved and ultrafast electron energy-loss spectroscopy

Time-resolved and ultrafast electron energy-loss spectroscopy (EELS) is an emerging technique for measuring photoexcited carriers, lattice dynamics, and near-fields across femtosecond to microsecond timescales. When performed in either a specialized scanning transmission electron microscope or ultrafast electron microscope (UEM), time-resolved and ultrafast EELS can directly image charge carriers, lattice vibrations, and heat dissipation following photoexcitation or applied bias. Yet, recent advances in theoretical calculations and electron optics are often required to realize the full potential of ultrafast EEL spectrum imaging. Here, in this review, we present a comprehensive overview of the recent progress in the theory and instrumentation of time-resolved and ultrafast EELS. We begin with an introduction to the technique, followed by a physical description of the loss function. We outline approaches for calculating and interpreting ground-state and transient EEL spectra spanning low-loss plasmons to core-level excitations analogous to x-ray absorption. We then survey the current state of time-resolved and ultrafast EELS techniques beyond photon-induced near-field electron microscopy, highlighting abilities to image carrier and thermal dynamics. Finally, we examine future directions enabled by emerging technologies, including electron beam monochromation, in situ and operando cells, laser-free UEM, and high-speed direct electron detectors. These advances position time-resolved and ultrafast EELS as a critical tool for uncovering nanoscale dynamic processes in quantum materials and solar energy conversion devices.

Computational methods↗

Deep-learning methods for contrast enhancement and artifact reduction in cryo-electron tomography: a systematic analysis of the state of the art and proposed improvements

Cryo-electron tomography (cryo-ET) has emerged as the preferred technique for visualizing the organization of macromolecular complexes in situ and resolving their structures at subnanometre resolution [Tegunov et al. (2021)View full citation, Nat. Methods, 18, 186–193]. Despite improvements in data quality as a result of advances in detector technology, microscope stability and stage precision, the analysis and interpretation of tomograms remains challenging due to a low signal-to-noise ratio and reconstruction artifacts stemming from experimental constraints in specimen tilt during data collection resulting in a missing wedge in the Fourier space. Recently, self-supervised deep-learning methods have been proposed for contrast enhancement and reduction of resolution anisotropy in reconstructed tomograms. Here, we evaluate several state-of-the-art deep-learning methods which aim to improve the interpretability of cryo-ET reconstructions, with a focus on their performance on downstream tasks of template matching, sub­tomogram averaging and segmentation. We propose new training architectures and a loss function based on Fourier shell correlation that show improved performance over the standard U-Net with L1/L2 losses. We demonstrate our analysis on four diverse experimental datasets: purified 80S ribosomes, in situ Chlamydomonas reinhardtii, immature HIV-1 virus-like particles and INS-1E cells.

contrast enhancement↗

Generative Vulnerability Assessment for Cyber-Physical Systems

Cyber-physical systems (CPS) are highly susceptible to malicious attacks due to their complex dynamics and interconnectivity. A comprehensive understanding of their vulnerabilities is essential for designing effective resilience measures. This paper presents a data-driven attack generative system for evaluating the vulnerability of CPS. The proposed approach formulates the vulnerability assessment problem as determining the feasibility of a specific attack set based on two boundary functions that represent the effectiveness and stealthiness of attacks. The attack generative model is trained using a custom loss function, with two universal approximators designed to learn the effectiveness and stealthiness functions simultaneously. Theoretical results for successful generation and asymptotic convergence of the resulting training algorithm are given. As a result, the proposed approach is evaluated via numerical simulation of an IEEE 14-bus system and gas pipeline systems, demonstrating its viability in learning how to attack nonlinear CPS and identify potential vulnerabilities.

Computer systems organization↗

Collective excitations and low-energy ionization signatures of relativistic particles in silicon detectors

Abstract Solid-state detectors with a low energy threshold have several applications, including searches of non-relativistic halo dark-matter particles with sub-GeV masses. When searching for relativistic, beyond-the-Standard-Model particles with enhanced cross sections for small energy transfers, a small detector with a low energy threshold may have better sensitivity than a larger detector with a higher energy threshold. In this paper, we calculate the low-energy ionization spectrum from high-velocity particles scattering in a dielectric material. We consider the full material response including the excitation of bulk plasmons. We generalize the energy-loss function to relativistic kinematics, and benchmark existing tools used for halo dark-matter scattering against electron energy-loss spectroscopy data. Compared to calculations commonly used in the literature, such as the Photo-Absorption-Ionization model or the free-electron model, including collective effects shifts the recoil ionization spectrum towards higher energies, typically peaking around 4–6 electron-hole pairs. We apply our results to the three benchmark examples: millicharged particles produced in a beam, neutrinos with a magnetic dipole moment produced in a reactor, and upscattered dark-matter particles. Our results show that the proper inclusion of collective effects typically enhances a detector’s sensitivity to these particles, since detector backgrounds, such as dark counts, peak at lower energies.

Physics↗

Reducing the Parameter Dependency of Phase-Picking Neural Networks with Dice Loss

Training a neural network for picking seismic phase arrivals has been commonly posed as a segmentation problem. It is a highly imbalanced segmentation problem in the sense that the background vastly dominates the foreground because we are trying to pick the optimal single sample point that represents the arrival of a seismic phase in a many seconds long time window. Here, we test the Dice loss, which is a preferred loss function for highly imbalanced image segmentation problems. We show that phase-picking neural networks trained on the Dice loss behave in a binary fashion for which the prediction output is almost always either nearly 1 or nearly 0. This feature removes the strong dependence of data processing workflows on the prediction score threshold, which is an otherwise critical parameter to determine when using neural networks trained on the cross-entropy loss. When strategically used, models trained on the Dice loss can reduce the parameter dependency of machine learning-based seismic monitoring.

58 GEOSCIENCES↗

Efficient Training of Deep Neural Operator Networks via Randomized Sampling

Neural operators (NOs) employ deep neural networks to learn the mappings between infinitedimensional function spaces. Deep operator network (DeepONet), a popular NO architecture, has demonstrated success in the real-time prediction of complex dynamics across various scientific and engineering applications. In this work, we introduce a random sampling technique to be adopted during the training of DeepONet, aimed at improving the generalization ability of the model, while significantly reducing the computational time. The proposed approach targets the trunk network of the DeepONet model that outputs the basis functions corresponding to the spatiotemporal locations of the bounded domain on which the physical system is defined. While constructing the loss function, DeepONet training traditionally considers a uniform grid of spatiotemporal points at which all the output functions are evaluated for each iteration. This approach leads to a larger batch size, resulting in poor generalization and increased memory demands, due to the limitations of the stochastic gradient descent (SGD) optimizer. The proposed random sampling over the inputs of the trunk net mitigates these challenges, improving generalization and reducing the memory requirements during training, resulting in significant computational gains. We validate our hypothesis through three benchmark examples, demonstrating substantial reductions in training time while achieving comparable or lower overall test errors relative to the traditional training approach. Our results indicate that incorporating randomization in the trunk network inputs during training enhances the efficiency and robustness of DeepONet, offering a promising avenue for improving the framework’s performance in modeling complex physical systems.

Karumuri, Sharmila [Department of Civil & Systems ↗

FunDiff: diffusion models over function spaces for physics-informed generative modeling

Recent advances in generative modeling-particularly diffusion models and flow matching-have been widely used for synthesizing discrete data such as images and videos. However, adapting these models to physical applications remains challenging, as the quantities of interest are continuous functions governed by complex physical laws. To address this, we introduce FunDiff, an efficient and robust framework for generative modeling in function spaces. FunDiff combines a latent diffusion process with a function autoencoder architecture to handle input functions with varying discretizations, generates continuous functions that can be evaluated at arbitrary locations, and seamlessly incorporate physical priors. These priors are enforced through architectural constraints or physics-informed loss functions, ensuring that generated samples satisfy fundamental physical laws. We theoretically establish minimax optimality guarantees for density estimation in function spaces, demonstrating that diffusion-based estimators achieve optimal convergence rates under suitable regularity conditions. We further demonstrate the practical effectiveness of FunDiff across diverse applications in fluid dynamics and solid mechanics. Empirical results indicate that our method can generate physically consistent samples with high fidelity to the target distribution, and exhibit robustness to noisy and low-resolution data.

Wang, Sifan [Yale University, New Haven, CT (Unite↗

Transforming the $v$ World: A New Multivariate Transformer Energy Estimator for NOvA

The NOvA Transformer Energy Estimator (Transformer_EE) is a universal machine learning tool currently used to infer the incoming beam neutrino energy and the outgoing lepton energy in both near andfar detectors. It uses a unique, highly flexible framework for simultaneous multivariate prediction that supports many possible loss functions. A spectral reweighting and flattening scheme lessens training bias. A feature noising subroutine enables adversarial-like training, mitigating sensitivities to certain systematic effects at marginal resolution loss at inference time. The state of the Transformer_EE will be reviewed, and its robustness with respect to several NOvA Near and Far Detector systematics highlighted.

Tong, Leon [Minnesota U.] (ORCID:0000000231625965)↗

Triangle Method for Dense ReLU Layers [SWR-25-72]

This software is an implementation of the methods for initializing and training neural networks to be more efficient per parameter, described more fully below and in the related publication: In theory, depth should make a ReLU network EXPONENTIALLY more efficient by enabling it to produce an exponential number of piecewise linear sections in its output. This reasoning is largely based on the work of mathematicians that have hand-constructed networks that make good use of depth. In practice however, even very deep ReLU networks that have been randomly initialized will behave identically to their shallow counterparts - missing an entire exponential dimension of efficiency. The triangle method is a first attempt at realizing the exponential potential of deep networks. Instead of randomly setting weights, we force pairs of neurons in each layer learn to build triangles (i.e. functions from [0,1] -> [0,1] that look like triangles). This is a very efficient pattern for generating lots of linear pieces because composing two triangular functions doubles the number of pieces with each composition. The triangle method is more than just a different initialization, it is a new paradigm of training. Instead of making direct updates to the matrix weights, we do an extra step of backpropagation to collect the derivatives of the loss function with respect to the shapes of the triangles, training them to tilt left or right. This process essentially holds the networks hand throughout the loss landscape and forces it to always use depth effectively by producing triangular shapes internally. This can produce several orders of magnitude of improvement on convex one-dimensional regression problems. Much more theoretical work is needed to realize its full potential beyond this context, but the implementation in this repository will still work in arbitrary numbers of dimensions. The file Triangle_Method.py is a generalized form of the method that will build each neuron its own custom 1-d convex activation function (with exponential efficiency). Example usage on one dimensional problems can be found in Example_Usage.ipynb and an example of using this in a real neural network can be found in Example_VGG16_CIFAR10.ipynb.

Milkert, Max [National Renewable Energy Laboratory↗

Chapter 4: Physically informed deep learning networks for simulating microstructure evolution of 3D polycrystals

As discussed in the previous chapter, high energy diffraction microscopy (HEDM) is used to study the micromechanical evolution of a material during in situ loading. HEDM experiments have been used to verify crystal plasticity (CP) simulations [119, 91, 90, 120], for experimental planning, material design, and to further analyze experimental results. However, Fast Fourier transform-based CP (CP-FFT) or finite element-based CP (CP-FE) methods are often too slow to be used in real-time during an experiment. CP-FFT is faster than CP-FE simulations due to the absence of meshing, but can still take hours to simulate the response of a single volume depending on the size and number of strain steps [127]. Reducing computation time would create a larger exploration space in planning and design, and enable faster analysis of experimental results and real-time feedback during an experiment. This research expands upon previous works to develop a workflow for predicting the full-field evolution of a 3D polycrystal. The workflow is simplified from previous works to predict only orientation and elastic strain tensors (from which stress tensors are calculated). The network is physically informed through loss functions and network architecture for a more robust model. The orientation predictions are informed about the cubic crystal symmetry of the material by incorporating disorientation and misorientation information into the network architecture and loss. The Von Mises stress is used to enforce the correct stress-strain trends in the strain tensor predictions. Additional total strain steps from the elastic and elastoplastic region are included to better capture the stress-strain evolution at smaller total strain steps. Material and hardening parameters are additional inputs into the networks to further inform the network and to study the network’s ability to predict different materials other than those used for training.

36 MATERIALS SCIENCE↗

Exact enforcement of temporal continuity in sequential physics-informed neural networks

The use of deep learning methods in scientific computing represents a potential paradigm shift in engineering problem solving. One of the most prominent developments is Physics-Informed Neural Networks (PINNs), in which neural networks are trained to satisfy partial differential equations (PDEs). While this method shows promise, the standard version has been shown to struggle in accurately predicting the dynamic behavior of time-dependent problems. To address this challenge, methods have been proposed that decompose the time domain into multiple segments, employing a distinct neural network in each segment and directly incorporating continuity between them in the loss function of the minimization problem. In this work we introduce a method to exactly enforce continuity between successive time segments via a solution ansatz. This hard constrained sequential PINN (HCS-PINN) method is simple to implement and eliminates the need for any loss terms associated with temporal continuity. The method is tested for a number of benchmark problems involving both linear and non-linear PDEs. Examples include various first order time dependent problems in which traditional PINNs struggle, namely advection, Allen–Cahn, and Korteweg–de Vries equations. Furthermore, second and third order time-dependent problems are demonstrated via wave and Jerky dynamics examples, respectively. Notably, the Jerky dynamics problem is chaotic, making the problem especially sensitive to temporal accuracy. Finally, the numerical experiments conducted with the proposed method demonstrated superior convergence and accuracy over both traditional PINNs and the soft-constrained counterparts.

42 ENGINEERING↗

SchrödingerNet: A Universal Neural Network Solver for the Schrödinger Equation

Recent advances in machine learning have facilitated numerically accurate solution of the electronic Schrödinger equation (SE) by integrating various neural network (NN)-based wave function ansatzes with variational Monte Carlo methods. Nevertheless, such NN-based methods are all based on the Born–Oppenheimer approximation (BOA) and require computationally expensive training for each nuclear configuration. In this work, we propose a novel NN architecture, SchrödingerNet, to solve the full electronic-nuclear SE by defining a loss function designed to equalize local energies across the system. This approach is based on a translationally, rotationally and permutationally symmetry-adapted total wave function ansatz that includes both nuclear and electronic coordinates. Furthermore, this strategy not only allows for an efficient and accurate generation of a continuous potential energy surface at any geometry within the well-sampled nuclear configuration space, but also incorporates non-BOA corrections, through a single training process. Comparison with benchmarks of atomic and small molecular systems demonstrates its accuracy and efficiency.

Chemical calculations↗

Sparsified time-dependent Fourier neural operators for fusion simulations

This paper presents a sparsified Fourier neural operator for coupled time-dependent partial differential equations (ST-FNO) as an efficient machine learning surrogate for fluid and particle-based fusion codes such as NIMROD (Non-Ideal Magnetohydrodynamics with Rotation - Open Discussion) and GTC (Gyrokinetic Toroidal Code). ST-FNO leverages the structures in the governing equations and utilizes neural operators to represent Green's function-like numerical operators in the corresponding numerical solvers. Once trained, ST-FNO can rapidly and accurately predict dynamics in fusion devices compared with first-principle numerical algorithms. In general, ST-FNO represents an efficient and accurate machine learning surrogate for numerical simulators for multi-variable nonlinear time-dependent partial differential equations, with the proposed architectures and loss functions. The efficacy of ST-FNO has been demonstrated using quiescent H-mode simulation data from NIMROD and kink-mode simulation data from GTC. The ST-FNO H-mode results show orders of magnitude reduction in memory and central processing unit usage in comparison with the numerical solvers in NIMROD when computing fields over a selected poloidal plane. The ST-FNO kink-mode results achieve a factor of 2 reduction in the number of parameters compared to baseline FNO models without accuracy loss.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Enhancing high-fidelity neural network potentials through low-fidelity sampling

The efficacy of neural network potentials (NNPs) critically depends on the quality of the configurational datasets used for training. Prior research using empirical potentials has shown that well-selected liquid–solid transitional configurations of a metallic system can be translated to other metallic systems. This study demonstrates that such validated configurations can be relabeled using density functional theory (DFT) calculations, thereby enhancing the development of high-fidelity NNPs. Training strategies and sampling approaches are efficiently assessed using empirical potentials and subsequently relabeled via DFT in a highly parallelized fashion for high-fidelity NNP training. Our results reveal that relying solely on energy and force for NNP training is inadequate to prevent overfitting, highlighting the necessity of incorporating stress terms into the loss functions. To optimize training involving force and stress terms, we propose employing transfer learning to fine-tune the weights, ensuring that the potential surface is smooth for these quantities composed of energy derivatives. This approach markedly improves the accuracy of elastic constants derived from simulations in both empirical potential-based NNPs and relabeled DFT-based NNPs. Overall, this study offers significant insights into leveraging empirical potentials to expedite the development of reliable and robust NNPs at the DFT level.

97 MATHEMATICS AND COMPUTING↗

4D Multimodal Co-attention Fusion Network with Latent Contrastive Alignment for Alzheimer’s Diagnosis

Multimodal neuroimaging provides complementary structural and functional insights into both human brain organization and disease-related dynamics. Recent studies demonstrate enhanced diagnostic sensitivity for Alzheimer’s disease (AD) through synergistic integration of neuroimaging data (e.g., sMRI, fMRI) with tabular data (e.g., behavioral and cognitive tests). However, the intrinsic heterogeneity across modalities (e.g., 4D spatiotemporal fMRI dynamics vs. 3D anatomical sMRI structure) presents critical challenges for discriminative feature fusion, often leading to information loss or biased fusion. To bridge this gap, we propose M2M-AlignNet: a multimodal co-attention network with latent alignment for early AD diagnosis using sMRI and fMRI. At the core of our approach is a multi-patch-to-multi-patch (M2M) contrastive loss function that quantifies and reduces representational discrepancies via weighted patch correspondence, explicitly aligning fMRI components across brain regions with their sMRI structural substrates without one-to-one constraints. Additionally, we propose a latent-as-query co-attention module to autonomously discover fusion patterns, circumventing modality prioritization biases while minimizing feature redundancy. We conduct extensive experiments to confirm the effectiveness of our method and highlight the correspondence between fMRI and sMRI as AD biomarkers.

Wei, Yuxiang [Georgia Institute of Technology]↗