Search NASASearch

SEARCH · Search NASA

Results for “generalization error”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Generalization error guaranteed auto-encoder-based nonlinear model reduction for operator learning

Many physical processes in science and engineering are naturally represented by operators between infinite-dimensional function spaces. The problem of operator learning, in this context, seeks to extract these physical processes from empirical data, which is challenging due to the infinite or high dimensionality of data. An integral component in addressing this challenge is model reduction, which reduces both the data dimensionality and problem size. In this paper, we utilize low-dimensional nonlinear structures in model reduction by investigating Auto-Encoder-based Neural Network (AENet). AENet first learns the latent variables of the input data and then learns the transformation from these latent variables to corresponding output data. Our numerical experiments validate the ability of AENet to accurately learn the solution operator of nonlinear partial differential equations. Furthermore, we establish a mathematical and statistical estimation theory that analyzes the generalization error of AENet. Finally, our theoretical framework shows that the sample complexity of training AENet is intricately tied to the intrinsic dimension of the modeled process, while also demonstrating the robustness of AENet to noise.

Auto-encoder

Parameter uncertainties for imperfect surrogate models in the low-noise regime

Abstract Bayesian regression determines model parameters by minimizing the expected loss, an upper bound to the true generalization error. However, this loss ignores model form error, or misspecification, meaning parameter uncertainties are significantly underestimated and vanish in the large data limit. As misspecification is the main source of uncertainty for surrogate models of low-noise calculations, such as those arising in atomistic simulation, predictive uncertainties are systematically underestimated. We analyze the true generalization error of misspecified, near-deterministic surrogate models, a regime of broad relevance in science and engineering. We show that posterior parameter distributions must cover every training point to avoid a divergence in the generalization error and design a compatible ansatz which incurs minimal overhead for linear models. The approach is demonstrated on model problems before application to thousand-dimensional datasets in atomistic machine learning. Our efficient misspecification-aware scheme gives accurate prediction and bounding of test errors in terms of parameter uncertainties, allowing this important source of uncertainty to be incorporated in multi-scale computational workflows.

Swinburne, Thomas D. (ORCID:0000000232554257)

Neural Scaling Laws of Deep ReLU and Deep Operator Network: A Theoretical Study

Neural scaling laws play a pivotal role in the performance of deep neural networks and have been observed in a wide range of tasks. However, a complete theoretical framework for understanding these scaling laws remains underdeveloped. In this paper, we explore the neural scaling laws for deep operator networks, which involve learning mappings between function spaces, with a focus on the Chen and Chen style architecture. These approaches, which include the popular Deep Operator Network (DeepONet), approximate the output functions using a linear combination of learnable basis functions and coefficients that depend on the input functions. We establish a theoretical framework to quantify the neural scaling laws by analyzing its approximation and generalization errors. We articulate the relationship between the approximation and generalization errors of deep operator networks and key factors such as network model size and training data size. Moreover, we address cases where input functions exhibit low-dimensional structures, allowing us to derive tighter error bounds. These results also hold for deep ReLU networks and other similar structures. Our results offer a partial explanation of the neural scaling laws in operator learning and provide a theoretical foundation for their applications.

97 MATHEMATICS AND COMPUTING

Deep nonparametric estimation of operators between infinite dimensional spaces

Learning operators between infinitely dimensional spaces is an important learning task arising in machine learning, imaging science, mathematical modeling and simulations, etc. This paper studies the nonparametric estimation of Lipschitz operators using deep neural networks. Non-asymptotic upper bounds are derived for the generalization error of the empirical risk minimizer over a properly chosen network class. Under the assumption that the target operator exhibits a low dimensional structure, our error bounds decay as the training sample size increases, with an attractive fast rate depending on the intrinsic dimension in our estimation. Our assumptions cover most scenarios in real applications and our results give rise to fast rates by exploiting low dimensional structures of data in operator estimation. We also investigate the influence of network structures (e.g., network width, depth, and sparsity) on the generalization error of the neural network estimator and propose a general suggestion on the choice of network structures to maximize the learning efficiency quantitatively.

97 MATHEMATICS AND COMPUTING

Scalable Hybrid Learning Techniques for Scientific Data Compression

Data compression is becoming critical for storing scientific data because many scientific applications need to store large amounts of data and post process this data for scientific discovery. Unlike image and video compression algorithms that limit errors to primary data (PD), scientists require compression techniques that accurately preserve derived quantities of interest (QoIs). Here, this article presents a physics-informed compression technique implemented as an end-to-end, scalable, GPU-based pipeline for data compression that addresses this requirement. Our hybrid compression technique combines machine learning techniques and standard compression methods. Specifically, we combine an autoencoder, an error-bounded lossy compressor to provide guarantees on raw data error, and a constraint satisfaction post-processing step to preserve the QoIs within a minimal error (generally less than floating point error). The effectiveness of the data compression pipeline is demonstrated by compressing nuclear fusion simulation data generated by a large-scale fusion code, XGC, which produces hundreds of terabytes of data in a single day. Our approach works within the ADIOS framework and results in compression by a factor of more than 150 while requiring only a few percent of the computational resources necessary for generating the data, making the overall approach highly effective for practical scenarios.

ITER

Coefficient-to-Basis Network: a fine-tunable operator learning framework for inverse problems with adaptive discretizations and theoretical guarantees

We propose a Coefficient-to-Basis Network (C2BNet), a novel framework for solving inverse problems within the operator learning paradigm. C2BNet efficiently adapts to different discretizations through fine-tuning, using a pre-trained model to significantly reduce computational cost while maintaining high accuracy. Unlike traditional approaches that require retraining from scratch for new discretizations, our method enables seamless adaptation without sacrificing predictive performance. Furthermore, we establish theoretical approximation and generalization error bounds for C2BNet by exploiting low-dimensional structures in the underlying datasets. Our analysis demonstrates that C2BNet adapts to low-dimensional structures without relying on explicit encoding mechanisms, highlighting its robustness and efficiency. To validate our theoretical findings, we conducted extensive numerical experiments that showcase the superior performance of C2BNet on several inverse problems. The results confirm that C2BNet effectively balances computational efficiency and accuracy, making it a promising tool to solve inverse problems in scientific computing and engineering applications.

97 MATHEMATICS AND COMPUTING

Theory and numerics of subspace approximation of eigenvalue problems

Large-scale eigenvalue problems arise in various fields of science and engineering and demand computationally efficient solutions. In this study, we investigate the subspace approximation for parametric linear eigenvalue problems, aiming to mitigate the computational burden associated with high-fidelity systems. Furthermore, we provide general error estimates under non-simple eigenvalue conditions, establishing some theoretical foundations for understanding the convergence behavior of subspace approximations. Numerical examples, including problems with one-dimensional to three-dimensional spatial domain and one-dimensional to two-dimensional parameter domain, are presented to demonstrate the efficacy of reduced basis method in handling parametric variations in boundary conditions and coefficient fields to achieve significant computational savings while maintaining high accuracy, making them promising tools for practical applications in large-scale eigenvalue computations.

Eigenvalue problems

Modeling and Experimental Validation of a Direct-Contact Counter-Flow Fluidized Bed Heat Exchanger for Thermal Energy Storage (TES) Applications

Particle-based thermal energy storage (TES) systems are an emerging energy storage technology. The technological advances have reduced costs, making TES more competitive and reliable in the marketplace but an efficient and reliable operation is heavily dependent on coherent heat transfer between air to particles or vice versa. The particle-based TES technologies provide an intermediate system that can store energy for short (0-10 h), long (10-200 h) and seasonal (> 200 h) timescales. The TES systems store energy by converting electricity to thermal energy; electricity can be directly sourced intermittent generation technologies and/or the grid, helping manage peak loads and other mismatches in supply and demand. The overall efficiency of the TES system depends on the performance of system components (particle storage silos and particle transfer mechanism etc.). The particle heat exchanger is one of the key system components that affects the system efficiency. The pressurized fluidized bed heat exchanger (PFB HX) performance is challenging to predict due to the chaotic behavior of particle and fluid interaction. This research presents a computational study of a novel direct-contact, counter-flow and air-to-particles PFB HX, that contributes in advancing the particle-based long-duration TES technologies. For the current analysis an unsteady Eulerian-Eulerian CFD model was developed and validated against experiments performed at the National Laboratory of the Rockies for two particle sizes (600 ..mu..m and 825 ..mu..m ). Following validation, parametric simulations were conducted to evaluate the effects of interphase drag models (Syamlal-O'Brien and Gidaspow), particle size, bed height and the influence of a frictional-viscosity term on hydrodynamics and heat transfer between the air & particles. The key findings from the analysis are: (1) for the studied operating window Syamlal-O'Brien provides superior agreement with measured gas temperatures (errors generally < 10%) while Gidaspow shows large deviations for the coarse particle case; (2) model predictions are most sensitive in the lower 0.2 m above the air distributor where bubble initiation and local mixing dominate interphase heat transfer; (3) representation of the distributor (number of inlet ports) materially affects predicted local mixing and temperature stratification; and (4) the Eulerian-Eulerian framework reproduces bulk thermal trends but shows regime dependent limitations for coarse particles, motivating mesoscale informed closures for scale-up analysis for future studies. These results provide validated guidance for drag selection and distributor design in particle-based thermal energy storage applications. Collectively, the validated model and parametric results quantify key drivers of PHB-HX performance and provide practical guidance for design and optimization. The results provide confidence in the model predictability and provide a step forward to improve on heat exchange performance. The demonstrated performance and modeling approach support the deployment and further development of this novel PHB-HX concept for robust, particle-based long-duration thermal energy storage systems.

25 ENERGY STORAGE

A tutorial review of machine learning-based model predictive control methods

Abstract This tutorial review provides a comprehensive overview of machine learning (ML)-based model predictive control (MPC) methods, covering both theoretical and practical aspects. It provides a theoretical analysis of closed-loop stability based on the generalization error of ML models and addresses practical challenges such as data scarcity, data quality, the curse of dimensionality, model uncertainty, computational efficiency, and safety from both modeling and control perspectives. The application of these methods is demonstrated using a nonlinear chemical process example, with open-source code available on GitHub. The paper concludes with a discussion on future research directions in ML-based MPC.

Wu, Zhe [Department of Chemical and Biomolecular E

Deep Neural Networks are Adaptive to Function Regularity and Data Distribution in Approximation and Estimation

Deep learning has exhibited remarkable results across diverse areas. To understand its success, substantial research has been directed towards its theoretical foundations. Nev- ertheless, the majority of these studies examine how well deep neural networks can model functions with uniform regularities. In this paper, we explore a different angle: how deep neural networks can adapt to varying degrees of smoothness in functions and nonuni- form data distributions across different locations and scales. More precisely, we focus on a broad class of functions defined by nonlinear tree-based approximation methods. This class encompasses a range of function types, such as functions with uniform regularities and discontinuous functions. We develop nonparametric approximation and estimation theories for this class using deep ReLU networks. Our results show that deep neural networks are adaptive to the nonuniform smoothness of functions and nonuniform data distributions at different locations and scales. We apply our results to several function classes, and derive the corresponding approximation and generalization errors. The validity of our results is demonstrated through numerical experiments.

97 MATHEMATICS AND COMPUTING

A General Framework for Error-controlled Unstructured Scientific Data Compression

Data compression plays a key role in reducing storage and I/O costs. Traditional lossy methods primarily target data on rectilinear grids and cannot leverage the spatial coherence in unstructured mesh data, leading to suboptimal compression ratios. We present a multi-component, error-bounded compression framework designed to enhance the compression of floating-point unstructured mesh data, which is common in scientific applications. Our approach involves interpolating mesh data onto a rectilinear grid and then separately compressing the grid interpolation and the interpolation residuals. This method is general, independent of mesh types and typologies, and can be seamlessly integrated with existing lossy compressors for improved performance. We evaluated our framework across twelve variables from two synthetic datasets and two real-world simulation datasets. The results indicate that the multi-component framework consistently outperforms state-of-the-art lossy compressors on unstructured data, achieving, on average, a 2.3 − 3.5× improvement in compression ratios, with error bounds ranging from 1 × 10 the −6 to 1×10−2. We further investigate impact of hyperparameters, such as grid spacing and error allocation, to deliver optimal compression ratios in diverse datasets.

Gong, Qian

On the discretization error of the discrete generalized quantum master equation

The transfer tensor method (TTM) [Cerrillo and Cao, Phys. Rev. Lett. 112 , 110401 (2014)] can be considered a discrete-time formulation of the Nakajima–Zwanzig quantum master equation (NZ-QME) for modeling non-Markovian quantum dynamics. A recent paper [Makri, J. Chem. Theory Comput. 21 , 5037 (2025)] raised concerns regarding the consistency of the TTM discretization, particularly a spurious term at the initial time t = 0. Here, this work presents a detailed analysis of the discretization structure of the TTM, clarifying the origin of the initial-time correction and establishing a consistent relationship between the TTM discrete-time memory kernel K N and the continuous-time NZ-QME kernel $\mathscr{K}$( N Δ t ). This relationship is validated numerically using the spin-boson model, demonstrating convergence of reconstructed memory kernels and accurate dynamical evolution as Δ t → 0. While the TTM provides a consistent discretization, we note that alternative schemes are also viable, such as the midpoint derivative/midpoint integral scheme proposed in Makri’s work. The relative performance of various schemes for either computing accurate $\mathscr{K}$( N Δ t ) from exact dynamics or obtaining accurate dynamics from exact $\mathscr{K}$( N Δ t ) warrants further investigation.

Density-matrix

Leveraging Qubit Loss Detection in Fault-Tolerant Quantum Algorithms

Qubit loss errors constitute a dominant source of noise in many quantum hardware systems, particularly in neutral-atom quantum computers. We develop a theoretical framework to effectively detect and correct loss errors in logical algorithms and leverage such loss information in decoding. Considering general quantum error correction codes and logical circuits, we introduce a delayed-erasure decoder for experimentally motivated error models which leverages information from delayed loss detection to accurately correct loss errors, even when the precise moment of the error is unknown. Using this decoder, we identify strategies for detecting and correcting loss errors based on the logical circuit structure. For deep circuits prior to logical measurement, we explore methods to integrate loss detection into syndrome extraction with minimal overhead, identifying optimal strategies depending on the qubit loss fraction in the noise and hardware capabilities. In contrast, we find that many key algorithmic subroutines involve frequent gate teleportation, shortening the circuit depth before logical measurement and naturally replacing qubits with no additional experimental overhead. We simulate this setting using a toy model algorithm for small-angle synthesis and find a significant performance improvement as the loss fraction increases. These results provide a path forward for advancing large-scale fault-tolerant quantum computation in systems with loss error detection.

atoms

Early calendar life and health prediction of silicon batteries via machine learning with uncertainty quantification

Lithium-ion batteries with silicon anodes promise high energy density but are limited by calendar lifetime. Reducing the long iteration time to obtain experimental results requires predicting calendar lifetime early in a cell's life. In this study, we demonstrate that lightweight machine learning models with feature engineering can provide calendar lifetime estimates from early electrochemical signals. After 1 month of electrochemical aging, the best models achieve 10% error in calendar-life prediction and can separate "bad" from "good" lifetime cells with a mean F1 score of 0.857. As battery systems exhibit inherent variability, four methods for uncertainty quantification are compared, and confidence intervals are demonstrated with an uncertainty of +-3.6 months in lifetime prediction. A feature importance analysis indicates that early patterns in voltage decay are the strongest indicators of calendar lifetime. Finally, this modeling approach has high error when generalizing to new electrode chemistries or testing conditions but with appropriately low confidence.

25 ENERGY STORAGE

The Atacama Cosmology Telescope: DR6 constraints on extended cosmological models

We use new cosmic microwave background (CMB) primary temperature and polarization anisotropy measurements from the Atacama Cosmology Telescope (ACT) Data Release 6 (DR6) to test foundational assumptions of the standard cosmological model, ΛCDM, and set constraints on extensions to it. We derive constraints from the ACT DR6 power spectra alone, as well as in combination with legacy data from the Planck mission. To break geometric degeneracies, we include ACT and Planck CMB lensing data and baryon acoustic oscillation data from DESI Year-1. To test the dependence of our results on non-ACT data, we also explore combinations replacing Planck with WMAP and DESI with BOSS, and further add supernovae measurements from Pantheon+ for models that affect the late-time expansion history. We verify the near-scale-invariance (running of the spectral index dn s /d ln k = 0.0062 ± 0.0052) and adiabaticity of the primordial perturbations. Neutrino properties are consistent with Standard Model predictions: we find no evidence for new light, relativistic species that are free-streaming (N eff = 2.86 ± 0.13, which combined with astrophysical measurements of primordial helium and deuterium abundances becomes N eff = 2.89 ± 0.11), for non-zero neutrino masses (∑m ν < 0.089 eV at 95% CL), or for neutrino self-interactions. We also find no evidence for self-interacting dark radiation (N idr < 0.134), or for early-universe variation of fundamental constants, including the fine-structure constant (α EM /α EM,0 = 1.0043 ± 0.0017) and the electron mass (m e /m e,0 = 1.0063 ± 0.0056). Our data are consistent with standard big bang nucleosynthesis (we find Y p = 0.2312 ± 0.0092), the COBE/FIRAS-inferred CMB temperature (we find T CMB = 2.698 ± 0.016 K), a dark matter component that is collisionless and with only a small fraction allowed as axion-like particles, a cosmological constant (w = -0.986 ± 0.025), and the late-time growth rate predicted by general relativity (γ = 0.663 ± 0.052). We find no statistically significant preference for a departure from the baseline ΛCDM model. In fits to models invoking early dark energy, primordial magnetic fields, or an arbitrary modified recombination history, we find H 0 = 69.9 +0.8 -1.5 , 69.1 ± 0.5, or 69.6 ± 1.0 km/s/Mpc, respectively; using BOSS instead of DESI BAO data reduces the central values of these constraints by 1–1.5 km/s/Mpc while only slightly increasing the error bars. In general, models introduced to increase the Hubble constant or to decrease the amplitude of density fluctuations inferred from the primary CMB are not favored over ΛCDM by our data.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Geometric Interpretation of the Cluster Location Problem Part I: Theory

We present a new framing of the seismic location problem using principles drawn from differential geometry. Our interpretation relies upon the common assumption that travel times observed across a network are continuous, differentiable functions of source location. In consequence, travel‐time functions constitute a differentiable map between the source region and a Riemannian manifold. The manifold is said to be the image of the source region embedded in a generally high‐dimension travel‐time vector space. A cluster of events in the source region has an image of discrete points on the manifold, that, except in the simplest cases, cannot be viewed directly. However, it is possible to project the image of a cluster into a tangent space of the manifold for direct visualization. The projection operator can be computed directly from the data without a velocity model, but produces a distorted rendering of the cluster geometry. With a model we can predict the distortions and correct them to estimate cluster geometry. We develop these points with the simplest possible example, one for which direct visualization of the manifold is possible, using the example as an introduction to the relevant concepts from differential geometry in a familiar setting. The tangent space, a local linearization of the manifold, plays a key role. We develop a metric to estimate the limits of linearization, that is, to determine when the curvature of the manifold invalidates the linear assumption. We also examine the interplay of model error, inadequate network geometry, and pick error. We then generalize our results from the simple case to the general case of 3D source regions observed by general networks. Although we do suggest a new “project and correct” method for location, we do not develop it into a practical algorithm. In conclusion, our intention rather is to highlight new analytical methods grounded in differential geometry.

East Pacific Ocean Islands