Search NASA⌕ Search

SEARCH · Search NASA

Results for “Backpropagation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Backpropagation-based learning with local derivative approximation and memory replay in biologically plausible neural systems

When learning, the brain modifies individual synaptic connections to reach a desired behavior. Animal and human brains have been shown to be incredibly capable of learning complex and varied functions across a wide variety of tasks. In recent years, artificial neural networks, inspired by human and animal brains, have shown great capabilities in learning a wide variety of difficult tasks. However, artificial neural networks primarily teach themselves through the use of backpropagation, a learning method which has no clear analogue within the brain. Additionally, Artificial Neural Networks primarily use continuous activation functions, which differ significantly from the spiking neuronal behavior present in the brain. In this paper, we discuss and demonstrate a biologically plausible learning method that approximates backpropagation through two techniques on Spiking Neural Networks. First, we show that the local temporal derivatives that are necessary for backpropagation can be approximately recovered through reconstruction using spike timings. Second, we show that through learning during a sleep phase, inspired by neuroscience research into memory replay, the localized parallel feedback path can learn to approximate the derivative through the forward path weight matrix, thus solving the weight transport problem. Lastly, we demonstrate that the combination of these two methods can approach or exceed the accuracy of backpropagation-based methods for a variety of neuromorphic vision tasks while maintaining biological plausibility.

42 ENGINEERING↗

Explainable AI classification for parton density theory

Quantitatively connecting properties of parton distribution functions (PDFs, or parton densities) to the theoretical assumptions made within the QCD analyses which produce them has been a longstanding problem in HEP phenomenology. To confront this challenge, we introduce an ML-based explainability framework, XAI4PDF, to classify PDFs by parton flavor or underlying theoretical model using ResNet-like neural networks (NNs). By leveraging the differentiable nature of ResNet models, this approach deploys guided backpropagation to dissect relevant features of fitted PDFs, identifying x-dependent signatures of PDFs important to the ML model classifications. By applying our framework, we are able to sort PDFs according to the analysis which produced them while constructing quantitative, human-readable maps locating the x regions most affected by the internal theory assumptions going into each analysis. This technique expands the toolkit available to PDF analysis and adjacent particle phenomenology while pointing to promising generalizations.

Artificial Intelligence↗

Topological Optimization with Big Steps

Using persistent homology to guide optimization has emerged as a novel application of topological data analysis. Existing methods treat persistence calculation as a black box and backpropagate gradients only onto the simplices involved in particular pairs. We show how the cycles and chains used in the persistence calculation can be used to prescribe gradients to larger subsets of the domain. In particular, we show that in a special case, which serves as a building block for general losses, the problem can be solved exactly in linear time. This relies on another contribution of this paper, which eliminates the need to examine a factorial number of permutations of simplices with the same value. Here, we present empirical experiments that show the practical benefits of our algorithm: the number of steps required for the optimization is reduced by an order of magnitude.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Entropy-defect synergy for dual luminescence mechanism in spinel: Time-resolved anti-counterfeiting and fingerprint visualization

Multimodal luminescent materials, while promising for anti-counterfeiting, often lack dynamic time-dependent responses and controllable spatial distribution, limiting their encryption capabilities in the spatiotemporal dimension. Here, this work presents a coordinated control strategy based on entropy and defect engineering, and uses a backpropagation (BP) neural network for material screening to successfully prepare spinel Mg 0.8 (Fe 0.04 Co 0.04 Ni 0.04 Cu 0.04 Zn 0.04 )Cr 2 O 4 (MgA 5 CO) phosphors with time-dependent dynamic luminescence behavior. This phosphor simultaneously activated the d-d transition luminescence (∼618 nm) derived from Co 2+ /Cr 3+ and the defect luminescence (∼398 nm) related to zinc vacancies (V Zn ) in a single-phase solid solution. The phosphor exhibits a time-dependent color evolution from pink to purple under fixed-wavelength excitation, due to the different excited-state dynamics and decay lifetimes associated with the d-d transition and defect luminescence. Structural characterization and spectral analysis confirmed the existence of V Zn and its significant role in defect luminescence process. The fluorescent and dynamic luminescent properties of entropy-based spinel oxide enable its use in advanced anti-counterfeiting applications like fingerprint recognition and color-changing dedicated anti-counterfeiting mark, showing promise in high-end and time-dynamic anti-counterfeiting fields. This research not only developed a new type of fluorescent dynamic anti-counterfeiting material, but also provided a new idea for constructing advanced optical functional materials with multiple luminescence mechanisms.

Defect project↗

Regularization via f -Divergence: An Application to Multi-Oxide Spectroscopic Analysis

In this paper, we explore the application of convolutional neural networks (CNNs) for predicting the chemical composition of complex geologic samples in a simulated Martian atmospheric environment. Specifically, we aim to characterize oxide weight percentages (wt.%) of rock samples analyzed by remote Laser-Induced Breakdown Spectroscopy (LIBS), framing the problem as a multi-target regression task . Neural networks trained on LIBS spectra are prone to overfitting due to high spectral complexity, limited labeled data, and measurement noise. While regularization is critical for improving generalization, common methods (e.g., ℓ 2 regularization) impose constraints not directly tied to data distribution properties. We propose a novel regularization method based on a specific ƒ-divergence induced by a graph-based estimator, designed to constrain the distributional discrepancy between predictions and targets. This regularizer serves a dual purpose: (a) mitigating overfitting by enforcing a constraint on the distributional difference between predictions and noisy targets, and (b) acting as an auxiliary loss that penalizes large divergences. To enable backpropagation, we develop a differentiable approximation of this particular ƒ-divergence, making the method feasible for neural networks. Experiments on ChemCam and SuperCam LIBS calibration spectra show that mathematical equation-divergence regularization outperforms or matches standard regularization methods (ℓ 1 , ℓ 2 , dropout) and the classical baseline, partial least squares (PLS). Combining ƒ-divergence regularization with standard regularization yields further performance gains, indicating that distributional regularization is useful in this context giving a promising direction for robust model training in planetary science applications. Source code is publicly available at Klein and Li (2025), https://doi.org/10.11578/dc.20250530.7.

58 GEOSCIENCES↗

Dynamics of episodic supershear in the 2023 M7.8 Kahramanmaraş/Pazarcik earthquake, revealed by near-field records and computational modeling

Abstract The 2023 M7.8 Kahramanmaraş/Pazarcik earthquake was larger and more destructive than what had been expected. Here we analyzed nearfield seismic records and developed a dynamic rupture model that reconciles different currently conflicting inversion results and reveals spatially non-uniform propagation speeds in this earthquake, with predominantly supershear speeds observed along the Narli fault and at the southwest (SW) end of the East Anatolian Fault (EAF). The model highlights the critical role of geometric complexity and heterogeneous frictional conditions in facilitating continued propagation and influencing rupture speed. We also constrained the conditions that allowed for the rupture to jump from the Narli fault to EAF and to generate the delayed backpropagating rupture towards the SW. Our findings have important implications for understanding earthquake hazards and guiding future response efforts and demonstrate the value of physics based dynamic modeling fused with near-field data in enhancing our understanding of earthquake mechanisms and improving risk assessment.

Environmental Sciences & Ecology↗

Coded-mask-based wavefront sensing technique for APS nanofocusing beamline diagnostics

Here, we extend our recently developed coded-mask wavefront sensing technique to enable single-shot measurements of nanofocused x-ray beams. This method accurately reconstructs the focal beam profile by backpropagating the wavefront measured downstream of the beam focus. To validate its performance, we benchmarked it against the conventional fluorescence wire scan method, successfully measuring ∼120 nm focal spots at the 28-ID-B beamline of the Advanced Photon Source using a polymeric compound refractive lens. The results highlight the effectiveness of coded-mask wavefront sensing for high-precision beam profiling and its application as a real-time wavefront monitoring tool.

Shi, Xianbo [Argonne National Laboratory (ANL), Ar↗

Identification of tau leptons using a convolutional neural network with domain adaptation

A tau lepton identification algorithm,DeepTau, based on convolutional neural network techniques, has been developed in the CMS experiment to discriminate reconstructed hadronic decays of tau leptons (τ h ) from quark or gluon jets and electrons and muons that are misreconstructed as τ h candidates. The latest version of this algorithm, v2.5, includes domain adaptation by backpropagation, a technique that reduces discrepancies between collision data and simulation in the region with the highest purity of genuine τh candidates. Additionally, a refined training workflow improves classification performance with respect to the previous version of the algorithm, with a reduction of 30–50% in the probability for quark and gluon jets to be misidentified as τ h candidates for given reconstruction and identification efficiencies. This paper presents the novel improvements introduced in theDeepTau algorithm and evaluates its performance in LHC proton-proton collision data at √(s) = 13 and 13.6 TeV collected in 2018 and 2022 with integrated luminosities of 60 and 35 fb -1 , respectively. Techniques to calibrate the performance of the τ h identification algorithm in simulation with respect to its measured performance in real data are presented, together with a subset of results among those measured for use in CMS physics analyses.

Large detector-systems performance↗

High-resolution source imaging and moment tensor estimation of acoustic emissions during brittle creep of basalt undergoing carbonation

SUMMARY As the high-frequency analogue to field-scale earthquakes, acoustic emissions (AEs) provide a valuable complement to study rock deformation mechanisms. During the load-stepping creep experiments with CO2-saturated water injection into a basaltic sample from Carbfix site in Iceland, 8791 AE events are detected by at least one of the seven piezoelectric sensors. Here, we apply a cross-correlation-based source imaging method, called geometric-mean reverse-time migration (GmRTM) to locate those AE events. Besides the attractive picking-free feature shared with other waveform-based methods (e.g. time-reversal imaging), GmRTM is advantageous in generating high-resolution source images with reduced imaging artefacts, especially for experiments with relatively sparse receivers. In general, the imaged AE locations are found to be scattered across the sample, suggesting a complicated fracture network rather than a well-defined major shear fracture plane, in agreement with X-ray computed tomography imaging results after retrieval of samples from the deformation apparatus. Clustering the events in space and time using the nearest-neighbour approach revealed a group of ‘repeaters’, which are spatially co-located over an elongated period of time and likely indicate crack, or shear band growth. Furthermore, we select 2196 AE events with high signal-to-noise-ratio (SNR) and conduct moment tensor estimation using the adjoint (backpropagated) strain tensor fields at the locations of AE sources. The resulting AE locations and focal mechanisms support our previously assertion that creep of basalt at the experimental conditions is accommodated dominantly by distributed microcracking.

58 GEOSCIENCES↗

Projective Integral Updates for High-Dimensional Variational Inference

Variational inference is an approximation framework for Bayesian inference that seeks to improve quantified uncertainty in predictions by optimizing a simplified distribution over parameters to stand in for the full posterior. Capturing model variations that remain consistent with training data enables more robust predictions by reducing parameter sensitivity. This work introduces a fixed-point optimization for variational inference that is applicable when every feasible log density can be expressed as a linear combination of functions from a given basis. In such cases, the optimizer becomes a fixed-point of projective integral updates. When the basis spans univariate quadratics in each parameter, the feasible distributions are Gaussian mean-fields and the projective integral updates yield quasi-Newton variational Bayes (QNVB). Other bases and updates are also possible. Since these updates require high-dimensional integration, this work begins by proposing an efficient quasirandom sequence of quadratures for mean-field distributions. Each iterate of the sequence contains two evaluation points that combine to correctly integrate all univariate quadratic functions and, if the mean-field factors are symmetric, all univariate cubics. More importantly, averaging results over short subsequences achieves periodic exactness on a much larger space of multivariate polynomials of quadratic total degree. The corresponding variational updates require four loss evaluations with standard (not second-order) backpropagation to eliminate error terms from over half of all multivariate quadratic basis functions. Furthermore, this integration technique is motivated by first proposing stochastic blocked mean-field quadratures, which may be useful in other contexts. A PyTorch implementation of QNVB allows for better control over model uncertainty during training than competing methods. Experiments demonstrate superior generalizability for multiple learning problems and architectures.

Gaussian mean-field↗

Semantic Stealth: Crafting Covert Adversarial Patches for Sentiment Classifiers Using Large Language Models

Deep learning models have been shown to be vulnerable to adversarial attacks, in which perturbations to their inputs cause the model to produce incorrect predictions. As opposed to adversarial attacks in computer vision, where small changes introduced to pixel values can drastically alter a model's output while remaining imperceptible to humans, text-based attacks are difficult to conceal due to the discrete nature of tokens. Consequently, unconstrained gradient-based attacks often produce adversarial examples that lack semantic meaning, rendering them detectable through visual inspection or perplexity filters. In contrast to methods that rely on gradient-based optimization in the embedding space, we propose an approach that leverages a Large Language Model's ability to generate grammatically correct and semantically meaningful text to craft adversarial patches that seamlessly blend in with the original input text. These patches can be used to alter the behavior of a target model, such as a text classifier. Since our approach does not rely on gradient backpropagation, it only requires access to the target model's confidence scores, making it a grey-box attack. We demonstrate the feasibility of our approach using open-source LLMs, including Intel's Neural Chat, Llama2, and Mistral-Instruct, to generate adversarial patches capable of altering the predictions of a distilBERT model fine-tuned on the IMDB reviews dataset for sentiment classification.

Roa Carvajal, Maria↗

Triangle Method for Dense ReLU Layers [SWR-25-72]

This software is an implementation of the methods for initializing and training neural networks to be more efficient per parameter, described more fully below and in the related publication: In theory, depth should make a ReLU network EXPONENTIALLY more efficient by enabling it to produce an exponential number of piecewise linear sections in its output. This reasoning is largely based on the work of mathematicians that have hand-constructed networks that make good use of depth. In practice however, even very deep ReLU networks that have been randomly initialized will behave identically to their shallow counterparts - missing an entire exponential dimension of efficiency. The triangle method is a first attempt at realizing the exponential potential of deep networks. Instead of randomly setting weights, we force pairs of neurons in each layer learn to build triangles (i.e. functions from [0,1] -> [0,1] that look like triangles). This is a very efficient pattern for generating lots of linear pieces because composing two triangular functions doubles the number of pieces with each composition. The triangle method is more than just a different initialization, it is a new paradigm of training. Instead of making direct updates to the matrix weights, we do an extra step of backpropagation to collect the derivatives of the loss function with respect to the shapes of the triangles, training them to tilt left or right. This process essentially holds the networks hand throughout the loss landscape and forces it to always use depth effectively by producing triangular shapes internally. This can produce several orders of magnitude of improvement on convex one-dimensional regression problems. Much more theoretical work is needed to realize its full potential beyond this context, but the implementation in this repository will still work in arbitrary numbers of dimensions. The file Triangle_Method.py is a generalized form of the method that will build each neuron its own custom 1-d convex activation function (with exponential efficiency). Example usage on one dimensional problems can be found in Example_Usage.ipynb and an example of using this in a real neural network can be found in Example_VGG16_CIFAR10.ipynb.

Milkert, Max [National Renewable Energy Laboratory↗

Graph-based Reversible Evaluation and Tangents Library

GRETL is a C++ library for evaluation, re-evaluation and algorithmic differentiation of functional operations on an arbitrary computational graph with limited memory usage. Similar to popular machine learning frameworks in Python, like PyTorch and JAX, it tracks and stores both operations and output data as functions are evaluated. Once this composition of functions is built up, the entire chain of operations can be back propagated to compute sensitivities of the final result with respect to any number of inputs. In contrast to most machine learning applications, memory usage becomes the bottleneck for back propagation in many physics applications, especially for time-dependent PDEs. Dynamic check pointing becomes essential. An important distinguishing feature of GRETL is its ability to limit the maximum memory usage by automatically dynamic checkpointing the data output for each graph operation (see Wang, Moin, Iaccarino, 2009). During backpropagation, parts of the graph that are no longer in memory are automatically re-evaluated from upstream checkpointed states as needed for derivative sensitivity calculations (or more precisely, for vector-Jacobian products). GRETL is particularly beneficial for applications, such as coupled multi-physics, where deriving adjoint-based sensitivities and managing checkpoint memory across modules becomes onerous. Cases which can be readily handled by the GRETL library include: different time-integration algorithms per physics (e.g., coupled predictor-corrector algorithms, IMEX, etc.), sub-cycling, asynchronous integrators, state dependent timestep sizes, iterative solvers and coupling algorithms, controller algorithms, and more.

Tupek, MichaelR [Lawrence Livermore National Labor↗

xesn: Echo state networks powered by Xarray and Dask

Xesn is a Python package that allows scientists to easily design Echo State Networks (ESNs) for forecasting problems. ESNs are a Recurrent Neural Network architecture introduced by Jaeger (2001) that are part of a class of techniques termed Reservoir Computing. One defining characteristic of these techniques is that all internal weights are determined by a handful of global, scalar parameters, thereby avoiding problems during backpropagation and reducing training time significantly. Because this architecture is conceptually simple, many scientists implement ESNs from scratch, leading to questions about computational performance. Xesn offers a straightforward, standard implementation of ESNs that operates efficiently on CPU and GPU hardware. The package leverages optimization tools to automate the parameter selection process, so that scientists can reduce the time finding a good architecture and focus on using ESNs for their domain application. Importantly, the package flexibly handles forecasting tasks for out-of-core, multi-dimensional datasets, eliminating the need to write parallel programming code. Xesn was initially developed to handle the problem of forecasting weather dynamics, and so it integrates naturally with Python packages that have become familiar to weather and climate scientists such as Xarray (Hoyer & Hamman, 2017). However, the software is ultimately general enough to be utilized in other domains where ESNs have been useful, such as in signal processing (Jaeger & Haas, 2004).

97 MATHEMATICS AND COMPUTING↗

Technical Report on Subsurface Monitoring of the Brady Hot Spring Geothermal Site, Nevada, based upon Full Waveform Inversion

Abilities to accurately characterize the subsurface in a geothermal setting is key to assess and support production. An important element of geothermal reservoir monitoring is also the ability to investigate fluid transport within fracture network. This report focuses on improving subsurface imaging and monitoring in geothermal settings using full waveform inversion based on the adjoint method and time-lapse imaging. To assess our method, we rely on a dense seismic dataset collected in 2016 at the Brady Hot Springs geothermal site in Nevada for the DOE-funded project Poroelastic Tomography by Adjoint Inverse Modeling of Data from Seismology, Geodesy, and Hydrology. This dataset captures subsurface changes across four stages of geothermal power plant operations, which involve varying rates of fluid injection and extraction. Two velocity models were previously derived from this dataset using different methods: one based on travel times and another on sweep interferometry. Our first step is to refine these models using adjoint tomography, which has been applied successfully at global and regional-scales but is less common at the reservoir-scale. Two approaches are then explored for time-lapse analysis: directly comparing refined tomographic models from different stages or backpropagating waveform differences relative to a baseline tomographic model. The main take away is that both approaches highlight similar reservoir behaviors, but the latter approach is more computationally effective in capturing small-scale changes in subsurface properties. For this work, we leverage the use of Salvus (www.mondaic.com), an end-to-end seismic imaging solution, relying on the spectral element method to compute forward and adjoint simulations, and developed by Mondaic Ltd. It includes integrated workflow management that handles waveform and metadata, launches simulations, computes waveform misfits and adjoint sources, and iterates for model updates by nonlinear optimization.

15 GEOTHERMAL ENERGY↗

CRCNS22 Learning Rules in the Hippocampus and their Mapping to Neuromorphic Systems (Final Technical Report)

Large scale biologically-realistic computational models are key to investigating the interplay between structure and function in nervous systems, thus paving the way to new clinical methods and neuro-inspired computing solutions. This project focuses on the hippocampus, in particular the CA3-CA1 regions, due to their role in associative learning and memory, pattern separation and completion, and spatial navigation. Investigations into the neuronal organization and learning rule(s) of this circuit can shed light into how declarative memories are formed, stored, recalled and forgotten and inform computational, experimental and clinical neuroscience work. Our project aims at developing a novel data-driven methodology supported by a broad heterogeneous base of neuroscience experimental knowledge and inspired from advances in computer science and engineering. Specifically, this work will benchmark existing and new learning rules within a full-scale spiking neural network simulation of the CA3-CA1 region. The model will be based on an open-source repository, called the Hippocampome, which contains neuronal morphologies, firing patterns, synapse probabilities, and most other required parameters for all known neuron types in the rodent hippocampal formation. The model will be first trained in a supervised fashion for associative memory tasks using backpropagation through time traditionally used in computer science, enhanced with a new technique called the surrogate gradient method. This optimization method will be used to obtain a global loss minimization, but it is not biologically inspired as it assumes the use of data not locally available to the synapses. However, we propose its use as a benchmarking tool, to compare the training performance of local biologically plausible and hardware-mappable learning rules at scale. New rules or combinations will be proposed and tested as needed, based on the obtained results. Progress in this area will also drive the development of novel hardware-mappable algorithms for continual lifelong learning and categorization of new events from few presented examples. This project goes beyond the existing state-of-the-art by looking at large scale realistic neuronal circuits as networks trainable via global optimization methods such as surrogate gradient descent. The objective function of the brain that supports learning is largely unknown, but it is likely that it operates through local learning rules. Studying network trajectories around local minima as proposed in this work represents a useful strategy for understanding whether a network is training by using a specific (set of) learning rule(s). Starting from a completely untrained network is a challenging test since it is difficult to determine how the learning rule affects the trajectory of the network. This interdisciplinary project will help understand what rule governs learning in these regions or if multiple learning rules are involved. The work will develop a robust methodology to measure if the network is converging to the target solution, oscillating around it, or diverging away.

59 BASIC BIOLOGICAL SCIENCES↗