Search NASA⌕ Search

SEARCH · Search NASA

Results for “Loss function”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Temperature and flow velocity of the interplanetary gases along solar radii

The velocity distributions along solar radii for hydrogen and helium in interplanetary space are calculated by using the Danby-Camm formula modified with a loss function. From these distributions the radial temperature and radial flow velocity of the interplanetary gases are determined. The effects of solar gravitation and ionization loss, due to charge exchange and photoionization, on the gas temperature and velocity are described.

Wu, F. M.↗

Transforming the $v$ World: A New Multivariate Transformer Energy Estimator for NOvA

The NOvA Transformer Energy Estimator (Transformer_EE) is a universal machine learning tool currently used to infer the incoming beam neutrino energy and the outgoing lepton energy in both near andfar detectors. It uses a unique, highly flexible framework for simultaneous multivariate prediction that supports many possible loss functions. A spectral reweighting and flattening scheme lessens training bias. A feature noising subroutine enables adversarial-like training, mitigating sensitivities to certain systematic effects at marginal resolution loss at inference time. The state of the Transformer_EE will be reviewed, and its robustness with respect to several NOvA Near and Far Detector systematics highlighted.

Tong, Leon [Minnesota U.] (ORCID:0000000231625965)↗

Triangle Method for Dense ReLU Layers [SWR-25-72]

This software is an implementation of the methods for initializing and training neural networks to be more efficient per parameter, described more fully below and in the related publication: In theory, depth should make a ReLU network EXPONENTIALLY more efficient by enabling it to produce an exponential number of piecewise linear sections in its output. This reasoning is largely based on the work of mathematicians that have hand-constructed networks that make good use of depth. In practice however, even very deep ReLU networks that have been randomly initialized will behave identically to their shallow counterparts - missing an entire exponential dimension of efficiency. The triangle method is a first attempt at realizing the exponential potential of deep networks. Instead of randomly setting weights, we force pairs of neurons in each layer learn to build triangles (i.e. functions from [0,1] -> [0,1] that look like triangles). This is a very efficient pattern for generating lots of linear pieces because composing two triangular functions doubles the number of pieces with each composition. The triangle method is more than just a different initialization, it is a new paradigm of training. Instead of making direct updates to the matrix weights, we do an extra step of backpropagation to collect the derivatives of the loss function with respect to the shapes of the triangles, training them to tilt left or right. This process essentially holds the networks hand throughout the loss landscape and forces it to always use depth effectively by producing triangular shapes internally. This can produce several orders of magnitude of improvement on convex one-dimensional regression problems. Much more theoretical work is needed to realize its full potential beyond this context, but the implementation in this repository will still work in arbitrary numbers of dimensions. The file Triangle_Method.py is a generalized form of the method that will build each neuron its own custom 1-d convex activation function (with exponential efficiency). Example usage on one dimensional problems can be found in Example_Usage.ipynb and an example of using this in a real neural network can be found in Example_VGG16_CIFAR10.ipynb.

Milkert, Max [National Renewable Energy Laboratory↗

Chapter 4: Physically informed deep learning networks for simulating microstructure evolution of 3D polycrystals

As discussed in the previous chapter, high energy diffraction microscopy (HEDM) is used to study the micromechanical evolution of a material during in situ loading. HEDM experiments have been used to verify crystal plasticity (CP) simulations [119, 91, 90, 120], for experimental planning, material design, and to further analyze experimental results. However, Fast Fourier transform-based CP (CP-FFT) or finite element-based CP (CP-FE) methods are often too slow to be used in real-time during an experiment. CP-FFT is faster than CP-FE simulations due to the absence of meshing, but can still take hours to simulate the response of a single volume depending on the size and number of strain steps [127]. Reducing computation time would create a larger exploration space in planning and design, and enable faster analysis of experimental results and real-time feedback during an experiment. This research expands upon previous works to develop a workflow for predicting the full-field evolution of a 3D polycrystal. The workflow is simplified from previous works to predict only orientation and elastic strain tensors (from which stress tensors are calculated). The network is physically informed through loss functions and network architecture for a more robust model. The orientation predictions are informed about the cubic crystal symmetry of the material by incorporating disorientation and misorientation information into the network architecture and loss. The Von Mises stress is used to enforce the correct stress-strain trends in the strain tensor predictions. Additional total strain steps from the elastic and elastoplastic region are included to better capture the stress-strain evolution at smaller total strain steps. Material and hardening parameters are additional inputs into the networks to further inform the network and to study the network’s ability to predict different materials other than those used for training.

36 MATERIALS SCIENCE↗

Exact enforcement of temporal continuity in sequential physics-informed neural networks

The use of deep learning methods in scientific computing represents a potential paradigm shift in engineering problem solving. One of the most prominent developments is Physics-Informed Neural Networks (PINNs), in which neural networks are trained to satisfy partial differential equations (PDEs). While this method shows promise, the standard version has been shown to struggle in accurately predicting the dynamic behavior of time-dependent problems. To address this challenge, methods have been proposed that decompose the time domain into multiple segments, employing a distinct neural network in each segment and directly incorporating continuity between them in the loss function of the minimization problem. In this work we introduce a method to exactly enforce continuity between successive time segments via a solution ansatz. This hard constrained sequential PINN (HCS-PINN) method is simple to implement and eliminates the need for any loss terms associated with temporal continuity. The method is tested for a number of benchmark problems involving both linear and non-linear PDEs. Examples include various first order time dependent problems in which traditional PINNs struggle, namely advection, Allen–Cahn, and Korteweg–de Vries equations. Furthermore, second and third order time-dependent problems are demonstrated via wave and Jerky dynamics examples, respectively. Notably, the Jerky dynamics problem is chaotic, making the problem especially sensitive to temporal accuracy. Finally, the numerical experiments conducted with the proposed method demonstrated superior convergence and accuracy over both traditional PINNs and the soft-constrained counterparts.

42 ENGINEERING↗

SchrödingerNet: A Universal Neural Network Solver for the Schrödinger Equation

Recent advances in machine learning have facilitated numerically accurate solution of the electronic Schrödinger equation (SE) by integrating various neural network (NN)-based wave function ansatzes with variational Monte Carlo methods. Nevertheless, such NN-based methods are all based on the Born–Oppenheimer approximation (BOA) and require computationally expensive training for each nuclear configuration. In this work, we propose a novel NN architecture, SchrödingerNet, to solve the full electronic-nuclear SE by defining a loss function designed to equalize local energies across the system. This approach is based on a translationally, rotationally and permutationally symmetry-adapted total wave function ansatz that includes both nuclear and electronic coordinates. Furthermore, this strategy not only allows for an efficient and accurate generation of a continuous potential energy surface at any geometry within the well-sampled nuclear configuration space, but also incorporates non-BOA corrections, through a single training process. Comparison with benchmarks of atomic and small molecular systems demonstrates its accuracy and efficiency.

Chemical calculations↗

Sparsified time-dependent Fourier neural operators for fusion simulations

This paper presents a sparsified Fourier neural operator for coupled time-dependent partial differential equations (ST-FNO) as an efficient machine learning surrogate for fluid and particle-based fusion codes such as NIMROD (Non-Ideal Magnetohydrodynamics with Rotation - Open Discussion) and GTC (Gyrokinetic Toroidal Code). ST-FNO leverages the structures in the governing equations and utilizes neural operators to represent Green's function-like numerical operators in the corresponding numerical solvers. Once trained, ST-FNO can rapidly and accurately predict dynamics in fusion devices compared with first-principle numerical algorithms. In general, ST-FNO represents an efficient and accurate machine learning surrogate for numerical simulators for multi-variable nonlinear time-dependent partial differential equations, with the proposed architectures and loss functions. The efficacy of ST-FNO has been demonstrated using quiescent H-mode simulation data from NIMROD and kink-mode simulation data from GTC. The ST-FNO H-mode results show orders of magnitude reduction in memory and central processing unit usage in comparison with the numerical solvers in NIMROD when computing fields over a selected poloidal plane. The ST-FNO kink-mode results achieve a factor of 2 reduction in the number of parameters compared to baseline FNO models without accuracy loss.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Enhancing high-fidelity neural network potentials through low-fidelity sampling

The efficacy of neural network potentials (NNPs) critically depends on the quality of the configurational datasets used for training. Prior research using empirical potentials has shown that well-selected liquid–solid transitional configurations of a metallic system can be translated to other metallic systems. This study demonstrates that such validated configurations can be relabeled using density functional theory (DFT) calculations, thereby enhancing the development of high-fidelity NNPs. Training strategies and sampling approaches are efficiently assessed using empirical potentials and subsequently relabeled via DFT in a highly parallelized fashion for high-fidelity NNP training. Our results reveal that relying solely on energy and force for NNP training is inadequate to prevent overfitting, highlighting the necessity of incorporating stress terms into the loss functions. To optimize training involving force and stress terms, we propose employing transfer learning to fine-tune the weights, ensuring that the potential surface is smooth for these quantities composed of energy derivatives. This approach markedly improves the accuracy of elastic constants derived from simulations in both empirical potential-based NNPs and relabeled DFT-based NNPs. Overall, this study offers significant insights into leveraging empirical potentials to expedite the development of reliable and robust NNPs at the DFT level.

97 MATHEMATICS AND COMPUTING↗

4D Multimodal Co-attention Fusion Network with Latent Contrastive Alignment for Alzheimer’s Diagnosis

Multimodal neuroimaging provides complementary structural and functional insights into both human brain organization and disease-related dynamics. Recent studies demonstrate enhanced diagnostic sensitivity for Alzheimer’s disease (AD) through synergistic integration of neuroimaging data (e.g., sMRI, fMRI) with tabular data (e.g., behavioral and cognitive tests). However, the intrinsic heterogeneity across modalities (e.g., 4D spatiotemporal fMRI dynamics vs. 3D anatomical sMRI structure) presents critical challenges for discriminative feature fusion, often leading to information loss or biased fusion. To bridge this gap, we propose M2M-AlignNet: a multimodal co-attention network with latent alignment for early AD diagnosis using sMRI and fMRI. At the core of our approach is a multi-patch-to-multi-patch (M2M) contrastive loss function that quantifies and reduces representational discrepancies via weighted patch correspondence, explicitly aligning fMRI components across brain regions with their sMRI structural substrates without one-to-one constraints. Additionally, we propose a latent-as-query co-attention module to autonomously discover fusion patterns, circumventing modality prioritization biases while minimizing feature redundancy. We conduct extensive experiments to confirm the effectiveness of our method and highlight the correspondence between fMRI and sMRI as AD biomarkers.

Wei, Yuxiang [Georgia Institute of Technology]↗

Any Two Learning Algorithms Are (Almost) Exactly Identical

This paper shows that if one is provided with a loss function, it can be used in a natural way to specify a distance measure quantifying the similarity of any two supervised learning algorithms, even non-parametric algorithms. Intuitively, this measure gives the fraction of targets and training sets for which the expected performance of the two algorithms differs significantly. Bounds on the value of this distance are calculated for the case of binary outputs and 0-1 loss, indicating that any two learning algorithms are almost exactly identical for such scenarios. As an example, for any two algorithms A and B, even for small input spaces and training sets, for less than 2e(-50) of all targets will the difference between A's and B's generalization performance of exceed 1%. In particular, this is true if B is bagging applied to A, or boosting applied to A. These bounds can be viewed alternatively as telling us, for example, that the simple English phrase 'I expect that algorithm A will generalize from the training set with an accuracy of at least 75% on the rest of the target' conveys 20,000 bytes of information concerning the target. The paper ends by discussing some of the subtleties of extending the distance measure to give a full (non-parametric) differential geometry of the manifold of learning algorithms.

Wolpert, David H.↗

Challenges in Training PINNs: A Loss Landscape Perspective

This paper explores challenges in training Physics Informed Neural Networks (PINNs), emphasizing the role of the loss landscape in the training process. We examine difficulties in minimizing the PINN loss function, particularly due to ill conditioning caused by differential operators in the residual term. We compare gradient-based optimizers Adam, L-BFGS, and their combination Adam+L-FGS, showing the superiority of Adam+L-BFGS, and introduce a novel secondorder optimizer, NysNewton-CG (NNCG), which significantly improves PINN performance. Theoretically, our work elucidates the connection between ill-conditioned differential operators and ill-conditioning in the PINN loss and shows the benefits of combining first- and second-order optimization methods. Our work presents valuable insights and more powerful optimization strategies for training PINNs, which could improve the utility of PINNs for solving difficult partial differential equations.

Rathore, Pratik↗

Spatial Statistical Data Fusion (SSDF)

As remote sensing for scientific purposes has transitioned from an experimental technology to an operational one, the selection of instruments has become more coordinated, so that the scientific community can exploit complementary measurements. However, tech nological and scientific heterogeneity across devices means that the statistical characteristics of the data they collect are different. The challenge addressed here is how to combine heterogeneous remote sensing data sets in a way that yields optimal statistical estimates of the underlying geophysical field, and provides rigorous uncertainty measures for those estimates. Different remote sensing data sets may have different spatial resolutions, different measurement error biases and variances, and other disparate characteristics. A state-of-the-art spatial statistical model was used to relate the true, but not directly observed, geophysical field to noisy, spatial aggregates observed by remote sensing instruments. The spatial covariances of the true field and the covariances of the true field with the observations were modeled. The observations are spatial averages of the true field values, over pixels, with different measurement noise superimposed. A kriging framework is used to infer optimal (minimum mean squared error and unbiased) estimates of the true field at point locations from pixel-level, noisy observations. A key feature of the spatial statistical model is the spatial mixed effects model that underlies it. The approach models the spatial covariance function of the underlying field using linear combinations of basis functions of fixed size. Approaches based on kriging require the inversion of very large spatial covariance matrices, and this is usually done by making simplifying assumptions about spatial covariance structure that simply do not hold for geophysical variables. In contrast, this method does not require these assumptions, and is also computationally much faster. This method is fundamentally different than other approaches to data fusion for remote sensing data because it is inferential rather than merely descriptive. All approaches combine data in a way that minimizes some specified loss function. Most of these are more or less ad hoc criteria based on what looks good to the eye, or some criteria that relate only to the data at hand.

Braverman, Amy J.↗

Gearbox bearing crack growth prognostics and uncertainty quantification with physics-informed machine learning

This paper introduces the extreme theory of functional connections (X-TFC), a physics-informed machine learning algorithm, and tailors it to estimate the remaining useful life (RUL) of wind turbine gearbox bearings experiencing fatigue crack growth. Unlike purely data-driven methods, X-TFC embeds a physics model, based on Head's theory in this work, into its training objective. The core of X-TFC is a random-projection single-layer neural network trained via an extreme learning machine, which requires only limited damage progression data and solves for output weights with a least-squares optimization algorithm. A composite loss function balances the network's fit to observed degradation data against the residuals of the governing crack growth differential equation, ensuring the learned damage trajectory remains physically plausible. When applied to a vibration-based health-index (HI) dataset measured during the growth of a crack on the inner ring of a high-speed bearing in a wind turbine gearbox (Bechhoefer and Dubé, 2020), X-TFC achieves near-zero prediction bias. Even when trained on only the first 10 %–20 % of the damage progression data, with sufficient physics weighting its predictions remain monotonic and smooth, delivering high prognosability and trendability. To quantify the epistemic uncertainty, we employ a Monte Carlo ensemble of independently initialized X-TFC models trained on noise-perturbed data, which yields confidence intervals around each RUL estimate and captures both model-parameter and epistemic uncertainty. In addition to a vibration-based HI, we demonstrate that the proposed framework can be directly applied to a supervisory control and data acquisition (SCADA) data-based HI (Eftekhari Milani et al., 2026) measured during similar wind turbine gearbox bearing crack faults, preserving its accuracy and interpretability. This extension shows the versatility of our approach, which is applicable to bearings of multiple gearbox manufacturers, models, and ratings using only SCADA data. By integrating domain knowledge with machine learning, X-TFC offers a rapid, reliable tool for crack prognostics. Its adaptability to other bearing failure modes, such as pitch bearing ring cracks, positions X-TFC as a powerful enabler of data-driven, physics-informed asset management in the wind energy sector and beyond.

17 WIND ENERGY↗

Spacecraft automated operations

Trends in automation of planetary spacecraft are examined using data from missions as far back as Mariner '67 and up to the highly sophisticated Galileo. Nine design considerations which influence the degree of automation such as protection against catastrophic failures, highly repetitive functions, loss of spacecraft communications, and the need for near-real-time adaptivity are discussed. Rapid growth of automation is shown in terms of on-board hardware by plots of number of processors on board, the average speed of processors, and total core memory. The number of commands transmitted from the ground has grown to 5 million bits in Voyager, so that increases in mission complexity have increased both in spacecraft automation and ground operations. Achieving greater automation by transferring ground operations to the spacecraft with the current means of controlling missions, are considered noting proposed changes. For the future, improved computer technology, more microprocessors and increased core storage will be used, and the number of automated functions and their complexity will grow. It is concluded that using the growing computational capability of spacecraft will achieve more autonomy thus reversing the trend of increased mission complexity and cost.

Bird, T. H.↗

Physiology of a microgravity environment invited review: microgravity and skeletal muscle

Spaceflight (SF) has been shown to cause skeletal muscle atrophy; a loss in force and power; and, in the first few weeks, a preferential atrophy of extensors over flexors. The atrophy primarily results from a reduced protein synthesis that is likely triggered by the removal of the antigravity load. Contractile proteins are lost out of proportion to other cellular proteins, and the actin thin filament is lost disproportionately to the myosin thick filament. The decline in contractile protein explains the decrease in force per cross-sectional area, whereas the thin-filament loss may explain the observed postflight increase in the maximal velocity of shortening in the type I and IIa fiber types. Importantly, the microgravity-induced decline in peak power is partially offset by the increased fiber velocity. Muscle velocity is further increased by the microgravity-induced expression of fast-type myosin isozymes in slow fibers (hybrid I/II fibers) and by the increased expression of fast type II fiber types. SF increases the susceptibility of skeletal muscle to damage, with the actual damage elicited during postflight reloading. Evidence in rats indicates that SF increases fatigability and reduces the capacity for fat oxidation in skeletal muscles. Future studies will be required to establish the cellular and molecular mechanisms of the SF-induced muscle atrophy and functional loss and to develop effective exercise countermeasures.

short duration↗

Staffing implications of software productivity models

The attributes of software project staffing and productivity implied by equating the effects of two popular software models in a small neighborhood of a given effort-duration point are investigated. The first model presupposes that organizational productivity decreases as a function of the project staff size due to interfacing and intercommunication. The second, the so-called software equation, relates the product size to effort and duration through a power law tradeoff formula. The conclusions that may be reached by assuming that both of these describe project behavior, the former as a global phenomenon and the latter as a localized effect in a small neighborhood of a given effort duration point, are that (1) there is a calculable maximum effective staff level, which, if exceeded, reduces the project production rate, (2) there is a calculable maximum extent to which effort and time may be traded effectively, (3) it becomes ineffective in a practical sense to expend more than an additional 25 to 50% of resources in order to reduce delivery time, and (4) the team production efficiency can be computed directly from the staff level, the slope of the intercommunication loss function, and the ratio of exponents in the software equation.

Tausworthe, R. C.↗

Bandgap engineering of SrZrS 3 chalcogenide perovskite via substitutional doping for photovoltaic applications: a first-principles DFT study

Strontium zirconium sulfide (SrZrS 3 ) has garnered significant attention for photovoltaic (PV) applications due to its excellent optoelectronic properties, high chemical and moisture stability, and non-toxicity. However, the bandgaps of both the α- and β-phases lie outside the optimum ranges for both single-junction solar cells (SJSCs) and tandem solar cells (TSCs), thereby limiting their applications in PV technologies. In this study, we employed hybrid density functional theory to engineer the band gaps of α- and β-SrZrS 3 through substitutional doping. We considered three dopants (Hf, Sn, and Ti) at the Zr site of SrZrS 3 with varying doping concentrations and investigated their effects on the structural, electronic, and optical properties of the materials. We found that Sn and Ti doping effectively lowers the band gaps of both α- and β-SrZrS 3 , whereas Hf doping increases them. For x values up to 0.25, the band gaps of the α-SrZr 1−x Sn x S 3 and α-SrZr 1−x Ti x S 3 are within the optimum range for SJSCs, and those of β-SrZr 1−x Sn x S 3 and β-SrZr 1−x Ti x S 3 lie within the optimum range for Si/perovskite as well as perovskite/perovskite TSCs. The three dopants exhibited significant effects on the optical properties of both α- and β-SrZrS 3 , including the absorption coefficient, energy-loss functions, reflectivity, and refractivity spectra. Thermodynamic stability analysis revealed that for both phases, SrZr 1−x Hf x S 3 can be synthesized via exothermic processes, whereas the formation of SrZr 1−x Ti x S 3 and SrZr1−xSn x S 3 is endothermic and hence, not thermodynamically favorable. Further analysis showed that SrZr 1−x Ti x S 3 in both α- and β-phases are stable under thermodynamic equilibrium conditions, whereas SrZr 1−x Sn x S 3 is prone to dissociation into ternary phases (SrZrS 3 and SrSnS 3 ), especially at higher doping concentrations. These results show that Ti doping is effective in tuning the band gaps of α- and β-SrZrS 3 toward the optimal values for PV applications.

14 SOLAR ENERGY↗