Search NASASearch

SEARCH · Search NASA

Results for “generalization error”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Generalization error guaranteed auto-encoder-based nonlinear model reduction for operator learning

Many physical processes in science and engineering are naturally represented by operators between infinite-dimensional function spaces. The problem of operator learning, in this context, seeks to extract these physical processes from empirical data, which is challenging due to the infinite or high dimensionality of data. An integral component in addressing this challenge is model reduction, which reduces both the data dimensionality and problem size. In this paper, we utilize low-dimensional nonlinear structures in model reduction by investigating Auto-Encoder-based Neural Network (AENet). AENet first learns the latent variables of the input data and then learns the transformation from these latent variables to corresponding output data. Our numerical experiments validate the ability of AENet to accurately learn the solution operator of nonlinear partial differential equations. Furthermore, we establish a mathematical and statistical estimation theory that analyzes the generalization error of AENet. Finally, our theoretical framework shows that the sample complexity of training AENet is intricately tied to the intrinsic dimension of the modeled process, while also demonstrating the robustness of AENet to noise.

Auto-encoder

AveBoost2: Boosting for Noisy Data

AdaBoost is a well-known ensemble learning algorithm that constructs its constituent or base models in sequence. A key step in AdaBoost is constructing a distribution over the training examples to create each base model. This distribution, represented as a vector, is constructed to be orthogonal to the vector of mistakes made by the pre- vious base model in the sequence. The idea is to make the next base model's errors uncorrelated with those of the previous model. In previous work, we developed an algorithm, AveBoost, that constructed distributions orthogonal to the mistake vectors of all the previous models, and then averaged them to create the next base model s distribution. Our experiments demonstrated the superior accuracy of our approach. In this paper, we slightly revise our algorithm to allow us to obtain non-trivial theoretical results: bounds on the training error and generalization error (difference between training and test error). Our averaging process has a regularizing effect which, as expected, leads us to a worse training error bound for our algorithm than for AdaBoost but a superior generalization error bound. For this paper, we experimented with the data that we used in both as originally supplied and with added label noise-a small fraction of the data has its original label changed. Noisy data are notoriously difficult for AdaBoost to learn. Our algorithm's performance improvement over AdaBoost is even greater on the noisy data than the original data.

Oza, Nikunj C.

Performance analysis of a generalized concurrent error detection procedure

A general procedure for error detection in complex systems, called the data block capture and analysis monitoring process, is described and analyzed. It is assumed that, in addition to being exposed to potential external fault sources, a complex system will in general always contain embedded hardware and software fault mechanisms which can cause the system to perform incorrect computations and/or produce incorrect output. Thus, in operation, the system continuously moves back and forth between error and no-error states. These external fault sources or internal fault mechanisms are extremely difficult to detect. The data block capture and analysis monitoring process is concerned with detecting deviations from the normal performance of the system, known as errors, which are symptomatic of fault conditions. The process consists of repeatedly recording a fixed amount of data from a set of predetermined observation lines of the system being monitored (i.e., capturing a block of data) and then analyzing the captured block in an attempt to determine whether the system is functioning correctly. The performances of linear, quadratic, and logarithmic data analysis algorithms are rigorously characterized in terms of the probability of correctly detecting an error, the expectation and variance of the number of false alarms per error, and the expectation and variance of the latency in detection of errors. Insight into the nature of the general problem of error detection is obtained.

Blough, Douglas M.

An extremum principle for computation of the zone of tooth contact and generalized transmission error of spiral bevel gears

For a given set of forces transmitted by the gears, each of the three components of the generalized transmission error of spiral bevel gears is shown to be stationary with respect to small independent variations in the positions of the endpoints of the lines of tooth contact about their true values. The tangential generalized transmission error component is shown to take on a minimum value at the true endpoint positions. A computational procedure based on the method of steepest descent is described for computing the true line of contact endpoint positions and the three components of the generalized transmission error. A method for computing the Fourier series coefficients of the tooth meshing harmonics of the three generalized transmission error components also is provided.

Mark, W. D.

Use of the generalized transmission error in the equations of motion of gear systems

The vibratory excitation arising from a gear pair is widely recognized to be a consequence of the nonuniform transmission of motion by the gear pair. In the preceding paper in this issue, it is shown that a three-component transmission error is required to describe the nonuniform transmission of motion by bevel gears for vibration excitation characterization purposes. The expression for the three-component transmission error derived in that paper is combined in the present paper with an analysis of the mesh forces and mesh elasticity to yield an equation of constraint involving the six degree-of-freedom unknown vibratory displacements of the gear shaft centerliners, the three unknown components of the generalized force transmitted by the mesh, and the geometric deviations of the tooth running surfaces from perfect involute surfaces which are assumed known. This matrix equation can be combined with the equations of motion of a gear system to predict the vibratory response of the system to the generalized transmission error excitation arising from meshing gear pairs within the system.

Mark, W. D.

Parameter uncertainties for imperfect surrogate models in the low-noise regime

Abstract Bayesian regression determines model parameters by minimizing the expected loss, an upper bound to the true generalization error. However, this loss ignores model form error, or misspecification, meaning parameter uncertainties are significantly underestimated and vanish in the large data limit. As misspecification is the main source of uncertainty for surrogate models of low-noise calculations, such as those arising in atomistic simulation, predictive uncertainties are systematically underestimated. We analyze the true generalization error of misspecified, near-deterministic surrogate models, a regime of broad relevance in science and engineering. We show that posterior parameter distributions must cover every training point to avoid a divergence in the generalization error and design a compatible ansatz which incurs minimal overhead for linear models. The approach is demonstrated on model problems before application to thousand-dimensional datasets in atomistic machine learning. Our efficient misspecification-aware scheme gives accurate prediction and bounding of test errors in terms of parameter uncertainties, allowing this important source of uncertainty to be incorporated in multi-scale computational workflows.

Swinburne, Thomas D. (ORCID:0000000232554257)

Jet stream velocity errors in general circulation models

The longitude and time dependence of excessive wind speed errors above subtropical jets in current GCM forecasts is studied for 14 five-day winter forecasts using the NASA Goddard Laboratory for Atmospheres fourth-order GCM. Several distinct phenomena, which may be divided into four categories, are found to contribute to the excess winds. These categories are: (1) error growth above the jet near the Himalayas; (2) error growth above the jet initiated elsewhere followed by advection; (3) tropical moisture bursts appearing in equatorial regions and migrating northeastward to merge with and distort the jet; and (4) undulatory growth of waves in the meridional component of the jet stream velocity. The results of additional sensitivity studies of the Himalayan region error are also reported.

Tenenbaum, J.

Neural Scaling Laws of Deep ReLU and Deep Operator Network: A Theoretical Study

Neural scaling laws play a pivotal role in the performance of deep neural networks and have been observed in a wide range of tasks. However, a complete theoretical framework for understanding these scaling laws remains underdeveloped. In this paper, we explore the neural scaling laws for deep operator networks, which involve learning mappings between function spaces, with a focus on the Chen and Chen style architecture. These approaches, which include the popular Deep Operator Network (DeepONet), approximate the output functions using a linear combination of learnable basis functions and coefficients that depend on the input functions. We establish a theoretical framework to quantify the neural scaling laws by analyzing its approximation and generalization errors. We articulate the relationship between the approximation and generalization errors of deep operator networks and key factors such as network model size and training data size. Moreover, we address cases where input functions exhibit low-dimensional structures, allowing us to derive tighter error bounds. These results also hold for deep ReLU networks and other similar structures. Our results offer a partial explanation of the neural scaling laws in operator learning and provide a theoretical foundation for their applications.

97 MATHEMATICS AND COMPUTING

Attitude determination error analysis - General model and specific application

This paper presents a comprehensive approach to filter and dynamics modeling for attitude determination error analysis. The discussion includes models for both batch least-squares and sequential estimators, a specific dynamic model for attitude determination error analysis of a three-axis stabilized spacecraft equipped with strapdown gyros, and the incorporation of general attitude sensor observations. An analyst using this approach to perform an error analysis chooses a subset of the spacecraft parameters to be 'solve-for' parameters, which are to be estimated, and another subset to be 'consider' parameters, which are assumed to have errors but not to be estimated. The result of the error analysis is an indication of overall uncertainties in the 'solve-for' parameters, as well as the contributions of the various error sources to these uncertainties, including those of errors in the a priori 'solve-for' estimates, of measurement noise, of dynamic noise (also known as process noise or plant noise), and of 'consider' parameter uncertainties. The analysis of attitude, star tracker alignment, and gyro bias uncertainties for the Gamma Ray Observatory spacecraft provide a specific example of the use of a general-purpose software package incorporating these models.

Markley, F. Landis

The generalized transmission error of spiral bevel gears

The traditional definition of the transmission error of parallel-axis gear pairs is reviewed and shown to be unsuitable for characterizing the deviation from conjugate action of bevel gear pairs for vibration excitation characterization purposes. This situation is rectified by generalizing the concept of the transmission error of parallel-axis gears to a three-component transmission error for spiral bevel gears of nominal spherical involute design. A general relationship is derived which expresses the contributions to the three-component transmission error from each gear of a meshing spiral bevel pair as a linear transformation of the six coordinates that describe the deviation of the shaft centerline position of each gear of the pair from the position of its rigid perfect involute counterpart.

Mark, W. D.

Deep nonparametric estimation of operators between infinite dimensional spaces

Learning operators between infinitely dimensional spaces is an important learning task arising in machine learning, imaging science, mathematical modeling and simulations, etc. This paper studies the nonparametric estimation of Lipschitz operators using deep neural networks. Non-asymptotic upper bounds are derived for the generalization error of the empirical risk minimizer over a properly chosen network class. Under the assumption that the target operator exhibits a low dimensional structure, our error bounds decay as the training sample size increases, with an attractive fast rate depending on the intrinsic dimension in our estimation. Our assumptions cover most scenarios in real applications and our results give rise to fast rates by exploiting low dimensional structures of data in operator estimation. We also investigate the influence of network structures (e.g., network width, depth, and sparsity) on the generalization error of the neural network estimator and propose a general suggestion on the choice of network structures to maximize the learning efficiency quantitatively.

97 MATHEMATICS AND COMPUTING

Scalable Hybrid Learning Techniques for Scientific Data Compression

Data compression is becoming critical for storing scientific data because many scientific applications need to store large amounts of data and post process this data for scientific discovery. Unlike image and video compression algorithms that limit errors to primary data (PD), scientists require compression techniques that accurately preserve derived quantities of interest (QoIs). Here, this article presents a physics-informed compression technique implemented as an end-to-end, scalable, GPU-based pipeline for data compression that addresses this requirement. Our hybrid compression technique combines machine learning techniques and standard compression methods. Specifically, we combine an autoencoder, an error-bounded lossy compressor to provide guarantees on raw data error, and a constraint satisfaction post-processing step to preserve the QoIs within a minimal error (generally less than floating point error). The effectiveness of the data compression pipeline is demonstrated by compressing nuclear fusion simulation data generated by a large-scale fusion code, XGC, which produces hundreds of terabytes of data in a single day. Our approach works within the ADIOS framework and results in compression by a factor of more than 150 while requiring only a few percent of the computational resources necessary for generating the data, making the overall approach highly effective for practical scenarios.

ITER

TRMM On-Orbit Performance Re-Accessed After Control Change

The Tropical Rainfall Measuring Mission (TRMM) spacecraft, a joint mission between the U.S. and Japan, launched onboard an HI1 rocket on November 27,1997 and transitioned in August, 2001 from an average operating altitude of 350 kilometers to 402.5 kilometers. Due to problems using the Earth Sensor Assembly (ESA) at the higher altitude, TRMM switched to a backup attitude control mode. Prior to the orbit boost TRMM controlled pitch and roll to the local vertical using ESA measurements while using gyro data to propagate yaw attitude between yaw updates from the Sun sensors. After the orbit boost, a Kalman filter used 3-axis gyro data with Sun sensor and magnetometers to estimate onboard attitude. While originally intended to meet a degraded attitude accuracy of 0.7 degrees, the new control mode met the original 0.2 degree attitude accuracy requirement after improving onboard ephemeris prediction and adjusting the magnetometer calibration onboard. Independent roll attitude checks using a science instrument, the Precipitation Radar (PR) which was built in Japan, provided a novel insight into the pointing performance. The PR data helped identify the pointing errors after the orbit boost, track the performance improvements, and show subtle effects from ephemeris errors and gyro bias errors. It also helped identify average bias trends throughout the mission. Roll errors tracked by the PR from sample orbits pre-boost and post-boost are shown in Figure 1. Prior to the orbit boost the largest attitude errors were due to occasional interference in the ESA. These errors were sometime larger than 0.2 degrees in pitch and roll, but usually less, as estimated from a comprehensive review of the attitude excursions using gyro data. Sudden jumps in the onboard roll show up as spikes in the reported attitude since the control responds within tens of seconds to null the pointing error. The PR estimated roll tracks well with an estimate of the roll history propagated using gyro data. After the orbit boost, the attitude errors shown by the PR roll have a smooth sine-wave type signal because of the way that attitude errors propagate with the use of gyro data. Yaw errors couple at orbit period to roll with '/4 orbit lag. By tracking the amplitude, phase, and bias of the sinusoidal PR roll error signal, it was shown that the average pitch rotation axis tends to be offset from orbit normal in a direction perpendicular to the Sun direction, as shown in Figure 2 for a 200 day period following the orbit boost. This is a result of the higher accuracy and stability of the Sun sensor measurements relative to the magnetometer measurements used in the Kalman filter. In November, 2001 a magnetometer calibration adjustment was uploaded which improved the pointing performance, keeping the roll and yaw amplitudes within about 0.1 degrees. After the boost, onboard ephemeris errors had a direct effect on the pitch pointing, being used to compute the Earth pointing reference frame. Improvements after the orbit boost have kept the the onboard ephemeris errors generally below 20 kilometers. Ephemeris errors have secondary effects on roll and yaw, especially during high beta angle when pitch effects can couple into roll and yaw. This is illustrated in figure 3. The onboard roll bias trends as measured by PR data show correlations with the Kalman filter's gyro bias error. This particularly shows up after yaw turns (every 2 to 4 weeks) as shown in Figure 3, when a slight roll bias is observed while the onboard computed gyro biases settle to new values. As for longer term trends, the PR data shows that the roll bias was influenced by Earth horizon radiance effects prior to the boost, changing values at yaw turns, and indicated a long term drift as shown in Figure 4. After the boost, the bias variations were smaller and showed some possible correlation with solar beta angle, probably due to sun sensor misalignment effects.

Bilanow, Steve

Error analysis of multi-conic techniques

A general error analysis of three recently developed multi-conic methods of three-body trajectory integration has been carried out. Single-step error functions for position and velocity have been derived as Taylor series in powers of the time step and also in integral form. These error functions are used to investigate the relative accuracy of the three methods in various regions of the earth-moon space and to provide a method of variable step size control for the trajectory integration procedure. Numerical results are used to compare the multi-step performance of the methods for both large and small step sizes.

D'Amario, L. A.

A general analysis of anti-jam communication systems

A general error bound is derived for a general anti-jam communication system which will serve as the basis for evaluating the performance of all such complex communication systems. The two most common spread spectrum techniques, coherent DS/BPSK and noncoherent FH/MFSK, are analyzed. Pulse jamming represents the worst type of jammer for DS/BPSK systems, and several receiver structures against such a jammer are examined. It is found that for low values of chip energy-to-noise ratios of O dB or less there is little difference between having or not having jammer state knowledge with a hard decision receiver. Soft decision receivers are shown to be useless against very narrow pulses without jammer state knowledge. Partial band jammers are close to the worst case jammer for FH/MFSK systems. The conclusions found for these systems are similar to those for the DS/BPSK systems.

Omura, J. K.

Coefficient-to-Basis Network: a fine-tunable operator learning framework for inverse problems with adaptive discretizations and theoretical guarantees

We propose a Coefficient-to-Basis Network (C2BNet), a novel framework for solving inverse problems within the operator learning paradigm. C2BNet efficiently adapts to different discretizations through fine-tuning, using a pre-trained model to significantly reduce computational cost while maintaining high accuracy. Unlike traditional approaches that require retraining from scratch for new discretizations, our method enables seamless adaptation without sacrificing predictive performance. Furthermore, we establish theoretical approximation and generalization error bounds for C2BNet by exploiting low-dimensional structures in the underlying datasets. Our analysis demonstrates that C2BNet adapts to low-dimensional structures without relying on explicit encoding mechanisms, highlighting its robustness and efficiency. To validate our theoretical findings, we conducted extensive numerical experiments that showcase the superior performance of C2BNet on several inverse problems. The results confirm that C2BNet effectively balances computational efficiency and accuracy, making it a promising tool to solve inverse problems in scientific computing and engineering applications.

97 MATHEMATICS AND COMPUTING