Search NASA⌕ Search

SEARCH · Search NASA

Results for “Loss function”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Entropy-Infused Deep Learning Loss Function for Capturing Extreme Values in Wind Power Forecasting

Extreme scenarios in wind power generation occur with higher frequency and larger magnitude in the recent years due to the ever-increasing extreme meteorological factors. Accurate forecasting of the occurrence of extreme values in wind power generation is of great concern to ensure reliable power system operation. Recently, deep learning models have surged in popularity for wind power forecasting, with the mean squared error (MSE) loss function being commonly used. However, the MSE loss function, being sensitive to extreme values, disproportionately penalizes larger errors, cannot adequately capture the extreme values present in wind energy data, and novel loss functions have seldom been tailored for wind power forecasting. To this end, in this paper, we introduce a novel loss function specifically crafted to capture extreme values in wind power forecasting. The experimental results with four fundamental deep learning methods on open source wind power dataset validate that the new loss function is efficient and superior in all cases compared to MSE in capturing extreme values while maintaining forecasting performance.

17 WIND ENERGY↗

Improved loss functions for machine-learned atomic potentials

Machine learning (ML) has become an invaluable tool across a wide array of domains in science as researchers find new ways to leverage its predictive power. This is especially true in chemistry, where ML is used to fit chemical properties or desirable attributes to the local structure of molecules and materials. In the pursuit of greater accuracy, it is relatively simple to increase the size or complexity of such models, although this often requires simultaneously seeking larger datasets in order to both fit and interpret the larger number of parameters. However, it is equally important to assess the quality and relative importance of the data and how these factors impact the training process. We, therefore, investigate the impact of using different loss functions for training neural network potentials (NNPs), as the loss function defines the error and parameter gradients used to train the NNP. In particular, we test the mean-squared error and Huber loss functions and, using insight from these functions, derive a new loss function based on the Asinh function, which yields significant improvement in the accuracy and generality of NNPs. We show that by discounting/minimizing errors and anomalies in the optimization process, both the Huber and Asinh loss functions improve the training of NNPs, leading to a final potential with a greater effective dimensionality.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Adaptive Quantum Generative Training using an Unbounded Loss Function

We propose a generative quantum learning algorithm using the Adaptive Derivative-Assembled Problem Tailored ansatz (ADAPT) framework in which the loss function to be minimized is the maximal quantum Rényi divergence of order two, an unbounded function that mitigates barren plateaus which inhibit training variational circuits. We benchmark this method against other state-of-the-art adaptive algorithms by learning random two-local thermal states. We perform numerical experiments of up to 12 qubits comparing our method learning algorithms that use linear objective functions and show that Rényi-ADAPT is capable of constructing shallow quantum circuits competitive with existing methods, while the gradients remain favorable resulting from the maximal Rényi divergence loss function.

quantum algorithms, quantum machine learning, quan↗

Effect of coronal elemental abundances on the radiative loss function

The solar photosphere and corona abundances tabulated by Meyer (1985) and the chromospheric abundances given by Murphy (1985) are used here to recalculate radiative loss functions for equilibrium, low-density, optically thin plasmas. Results from a representative standard photospheric abundance set and from coronal and chromospheric abundance sets showing depletions of up to a factor of four in certain elemental abundances are compared. A significant difference is found for both the coronal and chromospheric abundance sets, with the peak of the radiative loss curve shifted closer to 10 to the 6th K than to the standard 2 x 10 to the 5th K found from photospheric abundances. Consequences of these new calculations, in particular for the cool loop model of Antiochos and Noci (1986), are discussed.

Cook, J. W.↗

Jensen–Shannon divergence based novel loss functions for Bayesian neural networks

Bayesian neural networks (BNNs) are state-of-the-art machine learning methods that can naturally regularize and systematically quantify uncertainties using their stochastic parameters. Kullback–Leibler (KL) divergence-based variational inference used in BNNs suffer from unstable optimization and challenges in approximating light-tailed posteriors due to the unbounded nature of the KL divergence. To resolve these issues, we formulate a novel loss function for BNNs based on a new modification to the generalized Jensen–Shannon (JS) divergence, which is bounded. In addition, we propose a Geometric JS divergence-based loss, which is computationally efficient since it can be evaluated analytically. We found that the JS divergence-based variational inference is intractable, and hence employed a constrained optimization framework to formulate these losses. Our theoretical analysis and empirical experiments on multiple regression and classification data sets suggest that the proposed losses perform better than the KL divergence-based loss, especially when the data sets are noisy or biased. Specifically, there are approximately 5% and 8% improvements in accuracy for a noise-added CIFAR-10 dataset and a regression dataset, respectively. There is about 13% reduction in false negative predictions of a biased histopathology dataset. Additionally, we quantify and compare the uncertainty metrics for the regression and classification tasks.

97 MATHEMATICS AND COMPUTING↗

Attitude-Independent Magnetometer Calibration for Spin-Stabilized Spacecraft

The paper describes a three-step estimator to calibrate a Three-Axis Magnetometer (TAM) using TAM and slit Sun or star sensor measurements. In the first step, the Calibration Utility forms a loss function from the residuals of the magnitude of the geomagnetic field. This loss function is minimized with respect to biases, scale factors, and nonorthogonality corrections. The second step minimizes residuals of the projection of the geomagnetic field onto the spin axis under the assumption that spacecraft nutation has been suppressed by a nutation damper. Minimization is done with respect to various directions of the body spin axis in the TAM frame. The direction of the spin axis in the inertial coordinate system required for the residual computation is assumed to be unchanged with time. It is either determined independently using other sensors or included in the estimation parameters. In both cases all estimation parameters can be found using simple analytical formulas derived in the paper. The last step is to minimize a third loss function formed by residuals of the dot product between the geomagnetic field and Sun or star vector with respect to the misalignment angle about the body spin axis. The method is illustrated by calibrating TAM for the Fast Auroral Snapshot Explorer (FAST) using in-flight TAM and Sun sensor data. The estimated parameters include magnetic biases, scale factors, and misalignment angles of the spin axis in the TAM frame. Estimation of the misalignment angle about the spin axis was inconclusive since (at least for the selected time interval) the Sun vector was about 15 degrees from the direction of the spin axis; as a result residuals of the dot product between the geomagnetic field and Sun vectors were to a large extent minimized as a by-product of the second step.

Natanson, Gregory↗

Toward Physics-informed Neural Networks for 3D Multi-layer Cloud Mask Reconstruction

Three-dimensional (3D) cloud retrievals are critical for understanding their impact on climate and other applications such as aviation safety, weather prediction, and remote sensing. However, obtaining high-resolution and accurate vertical representation of clouds remains unsolved due to the limitations imposed by satellite instrumentation, viewing conditions, and the complexity of cloud dynamics. Cloud masks are essential for comprehending various cloud vertical properties, but deriving accurate 3D cloud masks from 2D satellite imagery data is a challenging task. To tackle these challenges, we introduce a physics-informed loss function for training deep learning models that can extend 2D cloud images into 3D cloud masks. The proposed loss, called CloudMask Loss, is composed of two domain knowledge-informed loss terms: one for evaluating cloud position and thickness, and the other for measuring the number of layers. By combining these loss terms, we improve the trainability of the deep learning models for more accurate and meaningful results. We apply the proposed loss function to different neural networks and demonstrate significant improvements in multi-layer cloud mask reconstruction. Utilizing the same neural network architecture, our proposed loss outperforms standard binary crossentropy loss in terms of multi-layer cloud classification accuracy, number of layers accuracy, and thickness mean absolute error (MAE). The proposed loss function can be readily integrated into various neural network architectures, resulting in substantial performance gains in 3D cloud mask generation.

multi-layer clouds↗

LossLens: Diagnostics for Machine Learning Through Loss Landscape Visual Analytics

Modern machine learning often relies on optimizing a neural network's parameters using a loss function to learn complex features. Beyond training, examining the loss function with respect to a network's parameters (i.e., as a loss landscape) can reveal insights into the architecture and learning process. While the local structure of the loss landscape surrounding an individual solution can be characterized using a variety of approaches, the global structure of a loss landscape, which includes potentially many local minima corresponding to different solutions, remains far more difficult to conceptualize and visualize. To address this difficulty, we introduce LossLens, a visual analytics framework that explores loss landscapes at multiple scales. LossLens integrates metrics from global and local scales into a comprehensive visual representation, enhancing model diagnostics. Here we demonstrate LossLens through two case studies: visualizing how residual connections influence a ResNet-20, and visualizing how physical parameters influence a physics-informed neural network (PINN) solving a simple convection problem.

97 MATHEMATICS AND COMPUTING↗

A strong loss-of-function mutation in RAN1 results in constitutive activation of the ethylene response pathway as well as a rosette-lethal phenotype

A recessive mutation was identified that constitutively activated the ethylene response pathway in Arabidopsis and resulted in a rosette-lethal phenotype. Positional cloning of the gene corresponding to this mutation revealed that it was allelic to responsive to antagonist1 (ran1), a mutation that causes seedlings to respond in a positive manner to what is normally a competitive inhibitor of ethylene binding. In contrast to the previously identified ran1-1 and ran1-2 alleles that are morphologically indistinguishable from wild-type plants, this ran1-3 allele results in a rosette-lethal phenotype. The predicted protein encoded by the RAN1 gene is similar to the Wilson and Menkes disease proteins and yeast Ccc2 protein, which are integral membrane cation-transporting P-type ATPases involved in copper trafficking. Genetic epistasis analysis indicated that RAN1 acts upstream of mutations in the ethylene receptor gene family. However, the rosette-lethal phenotype of ran1-3 was not suppressed by ethylene-insensitive mutants, suggesting that this mutation also affects a non-ethylene-dependent pathway regulating cell expansion. The phenotype of ran1-3 mutants is similar to loss-of-function ethylene receptor mutants, suggesting that RAN1 may be required to form functional ethylene receptors. Furthermore, these results suggest that copper is required not only for ethylene binding but also for the signaling function of the ethylene receptors.

NASA Discipline Plant Biology↗

Cooling of solar flares plasmas. 1: Theoretical considerations

Theoretical models of the cooling of flare plasma are reexamined. By assuming that the cooling occurs in two separate phase where conduction and radiation, respectively, dominate, a simple analytic formula for the cooling time of a flare plasma is derived. Unlike earlier order-of-magnitude scalings, this result accounts for the effect of the evolution of the loop plasma parameters on the cooling time. When the conductive cooling leads to an 'evaporation' of chromospheric material, the cooling time scales L(exp 5/6)/p(exp 1/6), where the coronal phase (defined as the time maximum temperature). When the conductive cooling is static, the cooling time scales as L(exp 3/4)n(exp 1/4). In deriving these results, use was made of an important scaling law (T proportional to n(exp 2)) during the radiative cooling phase that was forst noted in one-dimensional hydrodynamic numerical simulations (Serio et al. 1991; Jakimiec et al. 1992). Our own simulations show that this result is restricted to approximately the radiative loss function of Rosner, Tucker, & Vaiana (1978). for different radiative loss functions, other scaling result, with T and n scaling almost linearly when the radiative loss falls off as T(exp -2). It is shown that these scaling laws are part of a class of analytic solutions developed by Antiocos (1980).

Cargill, Peter J.↗

Real-Time Attitude Independent Three Axis Magnetometer Calibration

In this paper new real-time approaches for three-axis magnetometer sensor calibration are derived. These approaches rely on a conversion of the magnetometer-body and geomagnetic-reference vectors into an attitude independent observation by using scalar checking. The goal of the full calibration problem involves the determination of the magnetometer bias vector, scale factors and non-orthogonality corrections. Although the actual solution to this full calibration problem involves the minimization of a quartic loss function, the problem can be converted into a quadratic loss function by a centering approximation. This leads to a simple batch linear least squares solution. In this paper we develop alternative real-time algorithms based on both the extended Kalman filter and Unscented filter. With these real-time algorithms, a full magnetometer calibration can now be performed on-orbit during typical spacecraft mission-mode operations. Simulation results indicate that both algorithms provide accurate integer resolution in real time, but the Unscented filter is more robust to large initial condition errors than the extended Kalman filter. The algorithms are also tested using actual data from the Transition Region and Coronal Explorer (TRACE).

Crassidis, John L.↗

Addressing bias in bagging and boosting regression models

As artificial intelligence (AI) becomes widespread, there is increasing attention on investigating bias in machine learning (ML) models. Previous research concentrated on classification problems, with little emphasis on regression models. This paper presents an easy-to-apply and effective methodology for mitigating bias in bagging and boosting regression models, that is also applicable to any model trained through minimizing a differentiable loss function. Our methodology measures bias rigorously and extends the ML model's loss function with a regularization term to penalize high correlations between model errors and protected attributes. We applied our approach to three popular tree-based ensemble models: a random forest model (RF), a gradient-boosted model (GBT), and an extreme gradient boosting model (XGBoost). We implemented our methodology on a case study for predicting road-level traffic volume, where RF, GBT, and XGBoost models were shown to have high accuracy. Despite high accuracy, the ML models were shown to perform poorly on roads in minority-populated areas. Our bias mitigation approach reduced minority-related bias by over 50%.

97 MATHEMATICS AND COMPUTING↗