Search NASA⌕ Search

SEARCH · Search NASA

Results for “loss landscapes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

LossLens: Diagnostics for Machine Learning Through Loss Landscape Visual Analytics

Modern machine learning often relies on optimizing a neural network's parameters using a loss function to learn complex features. Beyond training, examining the loss function with respect to a network's parameters (i.e., as a loss landscape) can reveal insights into the architecture and learning process. While the local structure of the loss landscape surrounding an individual solution can be characterized using a variety of approaches, the global structure of a loss landscape, which includes potentially many local minima corresponding to different solutions, remains far more difficult to conceptualize and visualize. To address this difficulty, we introduce LossLens, a visual analytics framework that explores loss landscapes at multiple scales. LossLens integrates metrics from global and local scales into a comprehensive visual representation, enhancing model diagnostics. Here we demonstrate LossLens through two case studies: visualizing how residual connections influence a ResNet-20, and visualizing how physical parameters influence a physics-informed neural network (PINN) solving a simple convection problem.

97 MATHEMATICS AND COMPUTING↗

Loss Landscape Analysis for Reliable Quantized ML Models for Scientific Sensing

In this paper, we propose a method to perform empirical analysis of the loss landscape of machine learning (ML) models. The method is applied to two ML models for scientific sensing, which necessitates quantization to be deployed and are subject to noise and perturbations due to experimental conditions. Our method allows assessing the robustness of ML models to such effects as a function of quantization precision and under different regularization techniques -- two crucial concerns that remained underexplored so far. By investigating the interplay between performance, efficiency, and robustness by means of loss landscape analysis, we both established a strong correlation between gently-shaped landscapes and robustness to input and weight perturbations and observed other intriguing and non-obvious phenomena. Our method allows a systematic exploration of such trade-offs a priori, i.e., without training and testing multiple models, leading to more efficient development workflows. This work also highlights the importance of incorporating robustness into the Pareto optimization of ML models, enabling more reliable and adaptive scientific sensing systems.

Baldi, Tommaso [Pisa, Scuola Normale Superiore]↗

Challenges in Training PINNs: A Loss Landscape Perspective

This paper explores challenges in training Physics Informed Neural Networks (PINNs), emphasizing the role of the loss landscape in the training process. We examine difficulties in minimizing the PINN loss function, particularly due to ill conditioning caused by differential operators in the residual term. We compare gradient-based optimizers Adam, L-BFGS, and their combination Adam+L-FGS, showing the superiority of Adam+L-BFGS, and introduce a novel secondorder optimizer, NysNewton-CG (NNCG), which significantly improves PINN performance. Theoretically, our work elucidates the connection between ill-conditioned differential operators and ill-conditioning in the PINN loss and shows the benefits of combining first- and second-order optimization methods. Our work presents valuable insights and more powerful optimization strategies for training PINNs, which could improve the utility of PINNs for solving difficult partial differential equations.

Rathore, Pratik↗

Landscaper v1

Understanding the inner workings of machine learning models through their loss landscapes offers crucial insights into model properties, optimization dynamics, and generalizability. However, accessing these insights has traditionally required specialized mathematical expertise, limiting broader adoption. Landscaper is an open-source Python package designed to bridge this gap. Landscaper seamlessly integrates a suite of multi-dimensional loss landscape analyses with cutting-edge topological data analysis (TDA) methods. This powerful combination makes both fundamental loss landscape analysis and advanced TDA techniques accessible to the broader scientific ML community, without requiring deep pre-existing mathematical knowledge. Landscaper offers three key functionalities: * Construction: Builds detailed loss landscape representations through versatile low and high-dimensional sampling techniques. * Quantification: Applies advanced metrics, including a novel topological data analysis (TDA) based smoothness metric, enabling new perspectives on model behavior. * Visualization: Offers intuitive tools to visualize and interpret loss landscapes, providing actionable insights beyond traditional performance metrics.

Weber, Gunther [Lawrence Berkeley National Laborat↗

A Look at the Truths and Misconceptions of the Variational Quantum Eigensolver and the Implications of Overparameterization

In this work, we investigate loss landscapes of the variational quantum eigensolver (VQE) by quantifying the number of local minima through empirical analyses. We focus on minimal models in chemistry and physics so that we can do a complete analysis using more computationally expensive tools. We employ Hessian eigenvalue calculations and the nudged elastic band algorithm to characterize these landscapes. Our results expand upon the existing literature by highlighting the optimization challenges faced by VQE. We find that, as the number of parameters in our ansatz increases, the number of basins increases while the corresponding loss function values converge toward the global minimum value. This observation implies that overparameterization may lead to an ``effective convexity'' in VQE loss landscapes, a phenomenon supported by theoretical and numerical work in classical machine learning.

quantum computing↗

Theory of overparametrization in quantum neural networks

The prospect of achieving quantum advantage with quantum neural networks (QNNs) is exciting. Understanding how QNN properties (for example, the number of parameters $M$) affect the loss landscape is crucial to designing scalable QNN architectures. Here we rigorously analyze the overparametrization phenomenon in QNNs, defining overparametrization as the regime where the QNN has more than a critical number of parameters $M_c$ allowing it to explore all relevant directions in state space. In this study, our main results show that the dimension of the Lie algebra obtained from the generators of the QNN is an upper bound for $M_c$, and for the maximal rank that the quantum Fisher information and Hessian matrices can reach. Underparametrized QNNs have spurious local minima in the loss landscape that start disappearing when $M$ ≥ $M_c$. Thus, the overparametrization onset corresponds to a computational phase transition where the QNN trainability is greatly improved. We then connect the notion of overparametrization to the QNN capacity, so that when a QNN is overparametrized, its capacity achieves its maximum possible value.

97 MATHEMATICS AND COMPUTING↗

Data efficiency and extrapolation trends in neural network interatomic potentials

Abstract Recently, key architectural advances have been proposed for neural network interatomic potentials (NNIPs), such as incorporating message-passing networks, equivariance, or many-body expansion terms. Although modern NNIP models exhibit small differences in test accuracy, this metric is still considered the main target when developing new NNIP architectures. In this work, we show how architectural and optimization choices influence the generalization of NNIPs, revealing trends in molecular dynamics (MD) stability, data efficiency, and loss landscapes. Using the 3BPA dataset, we uncover trends in NNIP errors and robustness to noise, showing these metrics are insufficient to predict MD stability in the high-accuracy regime. With a large-scale study on NequIP, MACE, and their optimizers, we show that our metric of loss entropy predicts out-of-distribution error and data efficiency despite being computed only on the training set. This work provides a deep learning justification for probing extrapolation and can inform the development of next-generation NNIPs.

36 MATERIALS SCIENCE↗

Divide and conquer: Learning chaotic dynamical systems with multistep penalty neural ordinary differential equations

Forecasting high-dimensional dynamical systems is a fundamental challenge in various fields, such as geosciences and engineering. Neural Ordinary Differential Equations (NODEs), which combine the power of neural networks and numerical solvers, have emerged as a promising algorithm for forecasting complex nonlinear dynamical systems. However, classical techniques used for NODE training are ineffective for learning chaotic dynamical systems. In this work, we propose a novel NODE-training approach that allows for robust learning of chaotic dynamical systems. Here, our method addresses the challenges of non-convexity and exploding gradients associated with underlying chaotic dynamics. Training data trajectories from such systems are split into multiple, non-overlapping time windows. In addition to the deviation from the training data, the optimization loss term further penalizes the discontinuities of the predicted trajectory between the time windows. The window size is selected based on the fastest Lyapunov time scale of the system. Multi-step penalty(MP) method is first demonstrated on Lorenz equation, to illustrate how it improves the loss landscape and thereby accelerates the optimization convergence. MP method can optimize chaotic systems in a manner similar to least-squares shadowing with significantly lower computational costs. Our proposed algorithm, denoted the Multistep Penalty NODE, is applied to chaotic systems such as the Kuramoto-Sivashinsky equation, the two-dimensional Kolmogorov flow, and ERA5 reanalysis data for the atmosphere. It is observed that MP-NODE provide viable performance for such chaotic systems, not only for short-term trajectory predictions but also for invariant statistics that are hallmarks of the chaotic nature of these dynamics.

Chaotic dynamical systems↗

Optimizing the optimizer for physics-informed neural networks and Kolmogorov-Arnold networks

Physics-Informed Neural Networks (PINNs) have revolutionized the computation of PDE solutions by integrating partial differential equations (PDEs) into the neural network’s training process as soft constraints, becoming an important component of the scientific machine learning (SciML) ecosystem. More recently, physics-informed Kolmogorv-Arnold networks (PIKANs) have also shown to be effective and comparable in accuracy with PINNs. In their current implementation, both PINNs and PIKANs are mainly optimized using first-order methods like Adam, as well as quasi-Newton methods such as BFGS and its low-memory variant, L-BFGS. However, these optimizers often struggle with highly nonlinear and non-convex loss landscapes, leading to challenges such as slow convergence, local minima entrapment, and (non)degenerate saddle points. In this study, we investigate the performance of Self- Scaled BFGS (SSBFGS), Self-Scaled Broyden (SSBroyden) methods and other advanced quasi-Newton schemes, including BFGS and L-BFGS with different line search strategies. These methods dynamically rescale updates based on historical gradient information, thus enhancing training efficiency and accuracy. We systematically compare these optimizers – using both PINNs and PIKANs – on key challenging PDEs, including the Burgers, Allen-Cahn, Kuramoto-Sivashinsky, Ginzburg-Landau, and Stokes equations. Additionally, we evaluate the performance of SSBFGS and SSBroyden for Deep Operator Network (DeepONet) architectures, demonstrating their effectiveness for data-driven operator learning. Our findings provide state-of-the-art results with orders-of-magnitude accuracy improvements without the use of adaptive weights or any other enhancements typically employed in PINNs. More broadly, our work reveal insights into the effectiveness of quasi-Newton optimization strategies in significantly improving the convergence and accurate generalization of PINNs and PIKANs.

97 MATHEMATICS AND COMPUTING↗

Multi-module-based CVAE to predict HVCM faults in the SNS accelerator

We present a multi-module framework based on Conditional Variational Autoencoder (CVAE) to detect anomalies in the power signals coming from multiple High Voltage Converter Modulators (HVCMs). We condition the model with the specific modulator type to capture different representations of the $\mathcal{normal}$ waveforms and to improve the sensitivity of the model to identify a specific type of fault when we have limited samples for a given module type. We studied several Artificial Neural Network (ANN) architectures for our CVAE model and evaluated the model performance by looking at their loss landscape for stability and generalization. Our results for the Spallation Neutron Source (SNS) experimental data show that the trained model generalizes well to detecting multiple fault types for several HVCM module types. The results of this study can be used to improve the HVCM reliability and overall SNS uptime.

43 PARTICLE ACCELERATORS↗

Does provable absence of barren plateaus imply classical simulability?

A large amount of effort has recently been put into understanding the barren plateau phenomenon. In this perspective article, we face the increasingly loud elephant in the room and ask a question that has been hinted at by many but not explicitly addressed: Can the structure that allows one to avoid barren plateaus also be leveraged to efficiently simulate the loss classically? We collect evidence-on a case-by-case basis-that many commonly used models whose loss landscapes avoid barren plateaus can also admit classical simulation, provided that one can collect some classical data from quantum devices during an initial data acquisition phase. This follows from the observation that barren plateaus result from a curse of dimensionality, and that current approaches for solving them end up encoding the problem into some small, classically simulable, subspaces. Thus, while stressing that quantum computers can be essential for collecting data, our analysis sheds doubt on the information processing capabilities of many parametrized quantum circuits with provably barren plateau-free landscapes. We end by discussing the (many) caveats in our arguments including the limitations of average case arguments, the role of smart initializations, models that fall outside our assumptions, the potential for provably superpolynomial advantages and the possibility that, once larger devices become available, parametrized quantum circuits could heuristically outperform our analytic expectations.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Triangle Method for Dense ReLU Layers [SWR-25-72]

This software is an implementation of the methods for initializing and training neural networks to be more efficient per parameter, described more fully below and in the related publication: In theory, depth should make a ReLU network EXPONENTIALLY more efficient by enabling it to produce an exponential number of piecewise linear sections in its output. This reasoning is largely based on the work of mathematicians that have hand-constructed networks that make good use of depth. In practice however, even very deep ReLU networks that have been randomly initialized will behave identically to their shallow counterparts - missing an entire exponential dimension of efficiency. The triangle method is a first attempt at realizing the exponential potential of deep networks. Instead of randomly setting weights, we force pairs of neurons in each layer learn to build triangles (i.e. functions from [0,1] -> [0,1] that look like triangles). This is a very efficient pattern for generating lots of linear pieces because composing two triangular functions doubles the number of pieces with each composition. The triangle method is more than just a different initialization, it is a new paradigm of training. Instead of making direct updates to the matrix weights, we do an extra step of backpropagation to collect the derivatives of the loss function with respect to the shapes of the triangles, training them to tilt left or right. This process essentially holds the networks hand throughout the loss landscape and forces it to always use depth effectively by producing triangular shapes internally. This can produce several orders of magnitude of improvement on convex one-dimensional regression problems. Much more theoretical work is needed to realize its full potential beyond this context, but the implementation in this repository will still work in arbitrary numbers of dimensions. The file Triangle_Method.py is a generalized form of the method that will build each neuron its own custom 1-d convex activation function (with exponential efficiency). Example usage on one dimensional problems can be found in Example_Usage.ipynb and an example of using this in a real neural network can be found in Example_VGG16_CIFAR10.ipynb.

Milkert, Max [National Renewable Energy Laboratory↗

Characterizing and mitigating coherent errors in a trapped ion quantum processor using hidden inverses

Quantum computing testbeds exhibit high-fidelity quantum control over small collections of qubits, enabling performance of precise, repeatable operations followed by measurements. Currently, these noisy intermediate-scale devices can support a sufficient number of sequential operations prior to decoherence such that near term algorithms can be performed with proximate accuracy (like chemical accuracy for quantum chemistry problems). While the results of these algorithms are imperfect, these imperfections can help bootstrap quantum computer testbed development. Demonstrations of these algorithms over the past few years, coupled with the idea that imperfect algorithm performance can be caused by several dominant noise sources in the quantum processor, which can be measured and calibrated during algorithm execution or in post-processing, has led to the use of noise mitigation to improve typical computational results. Conversely, benchmark algorithms coupled with noise mitigation can help diagnose the nature of the noise, whether systematic or purely random. Here, we outline the use of coherent noise mitigation techniques as a characterization tool in trapped-ion testbeds. We perform model-fitting of the noisy data to determine the noise source based on realistic physics focused noise models and demonstrate that systematic noise amplification coupled with error mitigation schemes provides useful data for noise model deduction. Further, in order to connect lower level noise model details with application specific performance of near term algorithms, we experimentally construct the loss landscape of a variational algorithm under various injected noise sources coupled with error mitigation techniques. This type of connection enables application-aware hardware codesign, in which the most important noise sources in specific applications, like quantum chemistry, become foci of improvement in subsequent hardware generations.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Estimating Everglades Peat Reserves and Losses on a Landscape-Scale

The Florida Everglades formed from sawgrass and other aquatic plant remains that have accumulated over thousands of years and is a patterned peatland. The region was initially drained for agricultural and urban development in the late 1800s by the lowering of Lake Okeechobee water levels through canal drainage. Over the course of the past century, the Everglades has further been altered by the construction of canals and levees to provide control of Lake Okeechobee and Everglades water levels. The drainage has caused the loss of much of the peat of the Everglades by peat fires and microbial oxidation, much of it occurring between the 1880s and the 1950s. Plans for Everglades restoration are designed to provide additional water to flow through the remaining Everglades. Additional water into the Everglades would prevent further peat loss by reducing oxidation and encouraging the accretion of peat that has been lost. This potentially could result in the sequestration of significant quantities of atmospheric carbon. We used historic records to determine how much peat was there originally and how much has been lost. Over the past century and a half, approximately 50% of the original Everglades has been lost to agricultural and urban development. In addition, much of the peat in the remaining Everglades Protection Area has been lost. We used Geographic Information Systems technologies to determine the volume and mass of peat, the mass of carbon for each of the predrainage landscapes as well as the current Everglades regions. To do this, historical and current surveys within the Everglades were used to create elevation maps. With these maps, along with a bedrock contour map of south Florida, and soils data from the USEPA, we estimated Everglades peat characteristics. These calculations provide rough, quantitative estimates of the historical Everglades, its current condition, and the changes in the peat soil that have taken place since the 19th century. Given the uncertainties in the hindcasting of some of these data sets, this analysis provides rough estimates for the peat volumes and related constituents. Our calculations indicate that the current Everglades roughly contains less than 24% of theoriginal peat volume, 17% of its mass and 19% of its carbon. The work described was fully supported by the South Florida Water Management District, West Palm Beach, FL

Thomas Walter Dreschel↗

Landscape fragmentation overturns classical metapopulation thinking

Habitat loss and isolation caused by landscape fragmentation represent a growing threat to global biodiversity. Existing theory suggests that the process will lead to a decline in metapopulation viability. However, since most metapopulation models are restricted to simple networks of discrete habitat patches, the effects of real landscape fragmentation, particularly in stochastic environments, are not well understood. To close this major gap in ecological theory, we developed a spatially explicit, individual-based model applicable to realistic landscape structures, bridging metapopulation ecology and landscape ecology. This model reproduced classical metapopulation dynamics under conventional model assumptions, but on fragmented landscapes, it uncovered general dynamics that are in stark contradiction to the prevailing views in the ecological and conservation literature. Notably, fragmentation can give rise to a series of dualities: a) positive and negative responses to environmental noise, b) relative slowdown and acceleration in density decline, and c) synchronization and desynchronization of local population dynamics. Furthermore, counter to common intuition, species that interact locally (“residents”) were often more resilient to fragmentation than long-ranging “migrants.” This set of findings signals a need to fundamentally reconsider our approach to ecosystem management in a noisy and fragmented world.

54 ENVIRONMENTAL SCIENCES↗

Ecologically-Informed Precision Conservation: A framework for increasing biodiversity in intensively managed agricultural landscapes with minimal sacrifice in crop production

Conservation actions are urgently needed to tackle biodiversity loss in intensively managed agricultural landscapes. Production lands are usually heterogeneous and contain low-yield areas that can be set aside for biodiversity conservation without serious yield losses. Here, we introduce Ecologically-Informed Precision Conservation, a framework that integrates yield mapping and ecological theory to select the best areas to create new set-asides while ensuring high crop yields at the farm/landscape level. Long-term yield maps can be generated using globally available satellite data and basic information on field/farm crop yield from farmers. Ecological principles are then used to select the subset of areas with the highest potential for biodiversity conservation by prioritising those that increase connectivity, maximise habitat heterogeneity and decrease landscape grain size. Here, the created non-crop habitats can be permanent and thus ensure biodiversity support over time. In addition, agricultural management efficiency can be enhanced by improving field shapes. The framework provides the basis for a practical, user-friendly tool that informs all interested stakeholders on how to rationalise existing agricultural landscapes using already-existing farming systems and available technologies. High cost-effectiveness from an economic and conservation perspective, along with the creation of heterogeneous non-crop habitats, make our framework a promising solution to re-design agricultural landscapes.

Precision agriculture↗

Transmission risk of Oropouche fever across the Americas

Abstract Background Vector-borne diseases (VBDs) are important contributors to the global burden of infectious diseases due to their epidemic potential, which can result in significant population and economic impacts. Oropouche fever, caused by Oropouche virus (OROV), is an understudied zoonotic VBD febrile illness reported in Central and South America. The epidemic potential and areas of likely OROV spread remain unexplored, limiting capacities to improve epidemiological surveillance. Methods To better understand the capacity for spread of OROV, we developed spatial epidemiology models using human outbreaks as OROV transmission-locality data, coupled with high-resolution satellite-derived vegetation phenology. Data were integrated using hypervolume modeling to infer likely areas of OROV transmission and emergence across the Americas. Results Models based on one-support vector machine hypervolumes consistently predicted risk areas for OROV transmission across the tropics of Latin America despite the inclusion of different parameters such as different study areas and environmental predictors. Models estimate that up to 5 million people are at risk of exposure to OROV. Nevertheless, the limited epidemiological data available generates uncertainty in projections. For example, some outbreaks have occurred under climatic conditions outside those where most transmission events occur. The distribution models also revealed that landscape variation, expressed as vegetation loss, is linked to OROV outbreaks. Conclusions Hotspots of OROV transmission risk were detected along the tropics of South America. Vegetation loss might be a driver of Oropouche fever emergence. Modeling based on hypervolumes in spatial epidemiology might be considered an exploratory tool for analyzing data-limited emerging infectious diseases for which little understanding exists on their sylvatic cycles. OROV transmission risk maps can be used to improve surveillance, investigate OROV ecology and epidemiology, and inform early detection.

60 APPLIED LIFE SCIENCES↗