Slightly non-linear estimation with noisy data.
Nonlinear estimation with noisy data, considering least squares fit, sequential /Kalman/ estimate and iterated sequential scheme
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Nonlinear estimation with noisy data, considering least squares fit, sequential /Kalman/ estimate and iterated sequential scheme
AdaBoost is a well-known ensemble learning algorithm that constructs its constituent or base models in sequence. A key step in AdaBoost is constructing a distribution over the training examples to create each base model. This distribution, represented as a vector, is constructed to be orthogonal to the vector of mistakes made by the pre- vious base model in the sequence. The idea is to make the next base model's errors uncorrelated with those of the previous model. In previous work, we developed an algorithm, AveBoost, that constructed distributions orthogonal to the mistake vectors of all the previous models, and then averaged them to create the next base model s distribution. Our experiments demonstrated the superior accuracy of our approach. In this paper, we slightly revise our algorithm to allow us to obtain non-trivial theoretical results: bounds on the training error and generalization error (difference between training and test error). Our averaging process has a regularizing effect which, as expected, leads us to a worse training error bound for our algorithm than for AdaBoost but a superior generalization error bound. For this paper, we experimented with the data that we used in both as originally supplied and with added label noise-a small fraction of the data has its original label changed. Noisy data are notoriously difficult for AdaBoost to learn. Our algorithm's performance improvement over AdaBoost is even greater on the noisy data than the original data.
Multi-query applications such as parameter estimation, uncertainty quantification and design optimization for parameterized partial differential equation (PDE) systems are expensive. While reduced/latent state dynamics approaches for parameterized PDEs offer a viable alternative, these approaches rely on high-quality data and struggle with highly sparse spatiotemporal noisy measurements typically obtained from experiments. Furthermore, there is no guarantee that these models satisfy governing physical conservation laws. In this article, we propose a reduced state dynamics approach, referred to as ECLEIRS, that embeds exact conservation in the solution and flux representation by utilizing a space-time divergence-free neural network formulation. We compare ECLEIRS with other reduced state dynamics approaches, those that do not enforce any physical constraints and those with physics-informed loss functions, for three shock-propagation problems: 1-D advection, 1-D Burgers and 2-D Euler equations. In conclusion, the numerical experiments conducted in this study demonstrate that ECLEIRS provides the most accurate prediction of dynamics for unseen parameters even in the presence of highly sparse and noisy data.
Non-negative matrix factorization (NMF) is a dimensionality reduction technique that has shown promise for analyzing noisy data, especially astronomical data. For these datasets, the observed data may contain negative values due to noise even when the true underlying physical signal is strictly positive. Prior NMF work has not treated negative data in a statistically consistent manner, which becomes problematic for low signal-to-noise data with many negative values. In this paper we present two algorithms, Shift-NMF and Nearly-NMF, that can handle both the noisiness of the input data and also any introduced negativity. Both of these algorithms use the negative data space without clipping or masking and recover non-negative signals without any introduced positive offset that occurs when clipping or masking negative data. We demonstrate this numerically on both simple and more realistic examples, and prove that both algorithms have monotonically decreasing update rules.
Step input linear system modeling from nonlinear system sampled noisy data based on method of perturbation or quasilinearization of automatic control system data
Abstract We propose a new approach to generate a reliable reduced model for a parametric elliptic problem, in the presence of noisy data. The reference model reduction procedure is the directional HiPOD method, which combines Hierarchical Model reduction with a standard Proper Orthogonal Decomposition, according to an offline/online paradigm. In this paper we show that directional HiPOD looses in terms of accuracy when problem data are affected by noise. This is due to the interpolation driving the online phase, since it replicates, by definition, the noise trend. To overcome this limit, we replace interpolation with Machine Learning fitting models which better discriminate relevant physical features in the data from irrelevant unstructured noise. The numerical assessment, although preliminary, confirms the potentialities of the new approach.
There is an increasing aspiration to utilize machine learning (ML) for various tasks of relevance to national security. ML models have thus far been mostly applied to tasks and domains that, while impactful, have sufficient volume of data. For predictive tasks of national security relevance, ML models of great capacity (ability to approximate nonlinear trends in input-output maps) are often needed to capture the complex underlying physics. However, scientific problems of relevance to national security are often accompanied by various sources of sparse and/or incomplete data, including experiments and simulations, across different regimes of operation, of varying degrees of fidelity, and include noise with different characteristics and/or intensity. State-of-the-art ML models, despite exhibiting superior performance on the task and domain they were trained on, may suffer detrimental loss in performance in such sparse data environments. This report summarizes the results of the Laboratory Directed Research and Development project entitled Trust-Enhancing Probabilistic Transfer Learning for Sparse and Noisy Data Environments. The objective of the project was to develop a new transfer learning (TL) framework that aims to adaptively blend the data across different sources in tackling one task of interest, resulting in enhanced trustworthiness of ML models for mission- and safety-critical systems. The proposed framework determines when it is worth applying TL and how much knowledge is to be transferred, despite uncontrollable uncertainties. The framework accomplishes this by leveraging concepts and techniques from the fields of Bayesian inverse modeling and uncertainty quantification, relying on strong mathematical foundations of probability and measure theories to devise new uncertainty-aware TL workflows.
The goal of our Early Career Research Project (ECRP) is to establish a modern mathematical foundation that will enable next-generation computational methods for polynomial approximation of high-dimensional systems, having a certain set of constraints, from a limited amount of noisy data. Such a foundation is critical to realizing the future potential of the DOE user facilities, and will ultimately empower scientists to address a fundamental question, namely, “how many realizations of a nonlinear manifold are required to recover the entire high-dimensional solution map, with optimal approximation guarantees and minimal computational cost?” The central theme of this effort aims to conquer this challenge by pioneering the development of extraordinarily innovative theoretical analysis and transformational non-intrusive computational methodologies. Such approaches will enable the reconstruction of the entire high-dimensional solution map, with accuracy comparable to the best approximation, while utilizing an optimal number of samples. During this reporting period we have made significant progress on four thrusts.
Two reports present numerical study of performance of feedforward neural network trained by back-propagation algorithm in learning continuous-valued mappings from data corrupted by noise. Two types of noise considered: plant noise which affects dynamics of controlled process and data-processing noise, which occurs during analog processing and digital sampling of signals. Study performed with view toward use of neural networks as neurocontrollers to substitute for, or enhance, performances of human experts in controlling mechanical devices in presence of sensor and actuator noise and to enhance performances of more-conventional digital feedback electronic process controllers in noisy environments.
Differential equations and numerical methods are extensively used to model various real-world phenomena in science and engineering. With modern developments, we aim to find the underlying differential equation from a single observation of time-dependent data. If we assume that the differential equation is a linear combination of various linear and nonlinear differential terms, then the identification problem can be formulated as solving a linear system. The goal then reduces to finding the optimal coefficient vector that best represents the time derivative of the given data. We review some recent works on the identification of differential equations. We find some common themes for the improved accuracy: (i) The formulation of linear system with proper denoising is important, (ii) how to utilize sparsity and model selection to find the correct coefficient support needs careful attention, and (iii) there are ways to improve the coefficient recovery. We present an overview and analysis of recent developments on the topic.
Recent advances in the field of data-driven dynamics allow for the discovery of ODE systems using state measurements. One approach, known as Sparse Identification of Nonlinear Dynamics (SINDy), assumes the dynamics are sparse within a predetermined basis in the states and finds the expansion coefficients through linear regression with sparsity constraints. This approach requires an accurate estimation of the state time derivatives, which is not necessarily possible in the high-noise regime without additional constraints. We present an approach called Derivative-based SINDy (DSINDy) that combines two novel methods to improve ODE recovery at high-noise levels. First, we denoise the state variables by applying a projection operator that leverages the assumed basis for the system dynamics. Second, we use a second order cone program (SOCP) to find the derivative and governing equations simultaneously. We derive theoretical results for the projection-based denoising step, which allow us to estimate the values of hyperparameters used in the SOCP formulation. This underlying theory helps limit the number of required user-specified parameters. Finally, we present results demonstrating that our approach leads to improved system recovery for the Van der Pol oscillator, the Duffing oscillator, the Rössler attractor, and the Lorenz 96 model.
Not Available
In this work, we present a machine learning approach to calculating electronic specific heat capacities for a variety of benchmark molecular systems. Our models are based on data from density matrix quantum Monte Carlo, which is a stochastic method that can calculate the electronic energy at finite temperature. As these energies typically have noise, numerical derivatives of the energy can be challenging to find reliably. In order to circumvent this problem, we use Gaussian process regression to model the energy and use analytical derivatives to produce the specific heat capacity. From there, we also calculate the entropy by numerical integration. We compare our results to cubic splines and finite differences in a variety of molecules in which Hamiltonians can be diagonalized exactly with full configuration interaction. We finally apply this method to look at larger molecules where exact diagonalization is not possible and make comparisons with more approximate ways to calculate the specific heat capacity and entropy.
Understanding the equation of state (EOS) of pure neutron matter is necessary for interpreting multimessenger observations of neutron stars. Reliable data analyses of these observations require well-quantified uncertainties for the EOS input, ideally propagating uncertainties from nuclear interactions directly to the EOS. This, however, requires calculations of the EOS for a prohibitively larger number of nuclear Hamiltonians, solving the nuclear many-body problem for each one. Quantum Monte Carlo methods, such as auxiliary-field diffusion Monte Carlo (AFDMC), provide precise and accurate results for the neutron matter EOS, but they are very computationally expensive, making them unsuitable for the fast evaluations necessary for uncertainty propagation. Here, we employ parametric matrix models to develop fast emulators for AFDMC calculations of neutron matter and use them to directly propagate uncertainties of coupling constants in the Hamiltonian to the EOS. As these uncertainties include estimates of the effective field theory truncation uncertainty, this approach provides robust uncertainty estimates for use in astrophysical data analyses. In conclusion, this Letter will enable novel applications such as using astrophysical observations to put constraints on coupling constants for nuclear interactions.
Not provided.
Explore the source record for details and available documents.
Vector smoothing splines on the sphere are defined. Theoretical properties are briefly alluded to. The appropriate Hilbert space norms used in a specific meteorological application are described and justified via a duality theorem. Numerical procedures for computing the splines as well as the cross validation estimate of two smoothing parameters are given. A Monte Carlo study is described which suggests the accuracy with which upper air vorticity and divergence can be estimated using measured wind vectors from the North American radiosonde network.