Search NASASearch

SEARCH · Search NASA

Results for “robust regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

A data-driven approach to real-time vertical position estimation for NSTX-U vertical stability control

In this paper, a database of 77 996 plasma equilibrium reconstructions from 727 discharges during the initial operation of the NSTX-U spherical tokamak is analyzed to develop a statistically robust model of the plasma vertical position for real-time control. A variety of regression models are developed and tested, ranging in complexity from linear models to deep neural networks, and including input signals ranging from the four pairs of flux loops used historically on NSTX-U up to the full set of 389 real-time signals available to the plasma control system. A linear model based on 140 real-time magnetics signals is found to offer excellent accuracy, with a coefficient of determination R 2 = 0.906. The robustness of this model to limited training data, new operating scenarios, and signal errors is tested, and a procedure is demonstrated to tune the model parameters to optimize its robustness. A time-dependent plasma equilibrium solver, TokaMaker, is used to simulate vertical stability control in NSTX-U, demonstrating that it should be possible to iteratively tune the parameters of a linear vertical position model to stabilize both positive and negative triangularity plasmas in future experiments.

magnetic diagnostics

Chemical studies of H chondrites. 4: New data and comparison of Antarctic suites

We report data for the trace elements Au, Co, Sb, Ga, Rb, Ag, Se, Cs, Te, Zn, Cd, Bi, Ti, and In (ordered by putative volatility during nebular condensation and accretion) determined by neutron activation analysis in 13 H5 chondrites from Victoria Land and 20 H4-6 chondrites from Queen Maud Land, Antarctica. These and earlier results provide Antarctic sample suites of 34 chondrites from Victoria Land and 25 from Queen Maud Land. Treatment of data for the most volatile 10 elements (Rb to In) in these studies by multivariate statistical techniques more robust, as well as more conservative, than conventional linear discriminant analysis and logistic regression demonstrates that compositions differ at marginally significant levels. This difference cannot be explained by trivial (terrestrial) causes and becomes more significant, despite the smaller size of the database, when comparisons are limited to data from a single analyst and when all upper limits are eliminated from consideration. The Victoria Land and Queen Maud Land suites have different mean terrestrial ages (approximately 300 kyr and approximately 100 kyr, respectively) and age distributions, suggesting that a time-dependent variation of chondritic sources with different thermal histories is responsible. As a result, these two Antarctic suites are, on average, chemically distinguishable from each other. Since H chondrites serve as a paradigm for other meteorite classes, these results indicate that the near-Earth populations of planetary materials varied with time on the 10(exp 5)-year timescale.

Wolf, Stephen F.

Ames Hybrid Combustion Facility

The report summarizes the design, fabrication, safety features, environmental impact, and operation of the Ames Hybrid-Fuel Combustion Facility (HCF). The facility is used in conducting research into the scalability and combustion processes of advanced paraffin-based hybrid fuels for the purpose of assessing their applicability to practical rocket systems. The facility was designed to deliver gaseous oxygen at rates between 0.5 and 16.0 kg/sec to a combustion chamber operating at pressures ranging from 300 to 900. The required run times were of the order of 10 to 20 sec. The facility proved to be robust and reliable and has been used to generate a database of regression-rate measurements of paraffin at oxygen mass flux levels comparable to those of moderate-sized hybrid rocket motors.

Zilliac, Greg

A diagnostic analysis of the VVP single-doppler retrieval technique

A diagnostic analysis of the VVP (volume velocity processing) retrieval method is presented, with emphasis on understanding the technique as a linear, multivariate regression. Similarities and differences to the velocity-azimuth display and extended velocity-azimuth display retrieval techniques are discussed, using this framework. Conventional regression diagnostics are then employed to quantitatively determine situations in which the VVP technique is likely to fail. An algorithm for preparation and analysis of a robust VVP retrieval is developed and applied to synthetic and actual datasets with high temporal and spatial resolution. A fundamental (but quantifiable) limitation to some forms of VVP analysis is inadequate sampling dispersion in the n space of the multivariate regression, manifest as a collinearity between the basis functions of some fitted parameters. Such collinearity may be present either in the definition of these basis functions or in their realization in a given sampling configuration. This nonorthogonality may cause numerical instability, variance inflation (decrease in robustness), and increased sensitivity to bias from neglected wind components. It is shown that these effects prevent the application of VVP to small azimuthal sectors of data. The behavior of the VVP regression is further diagnosed over a wide range of sampling constraints, and reasonable sector limits are established.

Boccippio, Dennis J.

Developing a robust strength model using physically-informed genetic programming

The strength of materials is influenced by a range of external conditions, such as temperature and deformation rate. Consequently, materials that demonstrate substantial variations in their mechanical behavior due to fluctuations in temperature and strain rate require complex strength models to accurately predict material performance in real-world applications. To predict such complex behavior, a robust and flexible strength model is necessary. In this work, we utilize genetic programming-based symbolic regression (GPSR) to develop data-driven strength models that accurately represent the measured stress–strain responses of tin across a wide range of strain, strain rate and temperature regimes. The GPSR models are constrained by physically-informed conditions, which leads to significant improvement in extrapolation. The best model is integrated into a multi-physics code to perform Taylor impact simulations, validating the model’s accuracy and robustness. In conclusion, the model predictions showed excellent agreement with experimental results, particularly when compared to predictions using traditional strength models.

Genetic programming

Understanding and Verifying Neural Networks

Deep Neural Networks (DNNs) have gained immense popularity in recent times and have widespread use in applications such as image classification, sentiment analysis, speech recognition and also in safety-critical applications such as autonomous driving. However, they suffer limitations such as lack of explainability and robustness which raise safety and security concerns in their usage. Further, the complex structure and large input spaces of DNNs act as an impediment to thorough verification and testing. The SafeDNN project at the Robust Software Engineering (RSE) group at NASA aims at exploring techniques to ensure that systems that use deep neural networks are safe, robust and interpretable. In this talk, I will be presenting our technique Prophecy that automatically infers formal properties of deep neural network models. The tool extracts patterns based on neuron activations as preconditions that imply certain desirable output properties of the model. I would be highlighting case studies that use Prophecy in obtaining explanations for network decisions, understanding correct and incorrect behavior, providing formal guarantees wrt safety and robustness, and debugging neural network models. We have applied the tool on image classification networks, neural network controllers providing turn advisories in unmanned aircrafts, regression models used for autonomous center-line tracking in aircrafts and neural network object detectors

Deep Neural Networks

A Markov chain Monte Carlo (MCMC) Bayesian inference approach to analyze apparent activation barriers and reaction orders from microreactor data

Statistical analysis of steady-state catalytic kinetic data is often limited by data sparsity due to the slow pace at which the data is collected. Data sparsity and limitations in statistical analysis make it difficult to differentiate between mechanistic models and catalytic sites. A Bayesian inference tool is reported for catalysis researchers to estimate error in the determination of reaction orders from steady state microreactor data. The benefits of a Bayesian inference approach are discussed, as an alternative to the more common frequentist approach. The approach incorporates prior knowledge of the system and the data collected to form an error estimate on reaction orders. We investigated the effects of three distinct data treatments—individual fitting of trials, pooled analysis, and constrained regression methods—on the precision and uncertainty of reaction order determinations. To assess the robustness of our findings, we conducted sensitivity analyses to evaluate the influence of Bayesian parameters on uncertainty estimation. Additionally, we utilized synthetic data to illustrate how data quality impacts the precision of uncertainty assessments. We show Bayesian analysis can obtain a more precise estimation of error with a sparse data set than a frequentist analysis. Finally, this work provides strong evidence that the adoption of Bayesian analysis of kinetic data may help researchers make more precise arguments as to the strength of their evidence for a particular mechanistic hypothesis, or in comparing across different catalysts.

42 ENGINEERING

Analysis of Present Day and Future OH and Methane Lifetime in the ACCMIP Simulations

Results from simulations performed for the Atmospheric Chemistry and Climate Modeling Intercomparison Project (ACCMIP) are analysed to examine how OH and methane lifetime may change from present day to the future, under different climate and emissions scenarios. Present day (2000) mean tropospheric chemical lifetime derived from the ACCMIP multi-model mean is 9.8+/-1.6 yr (9.3+/-0.9 yr when only including selected models), lower than a recent observationally-based estimate, but with a similar range to previous multi-model estimates. Future model projections are based on the four Representative Concentration Pathways (RCPs), and the results also exhibit a large range. Decreases in global methane lifetime of 4.5 +/- 9.1% are simulated for the scenario with lowest radiative forcing by 2100 (RCP 2.6), while increases of 8.5+/-10.4% are simulated for the scenario with highest radiative forcing (RCP 8.5). In this scenario, the key driver of the evolution of OH and methane lifetime is methane itself, since its concentration more than doubles by 2100 and it consumes much of the OH that exists in the troposphere. Stratospheric ozone recovery, which drives tropospheric OH decreases through photolysis modifications, also plays a partial role. In the other scenarios, where methane changes are less drastic, the interplay between various competing drivers leads to smaller and more diverse OH and methane lifetime responses, which are difficult to attribute. For all scenarios, regional OH changes are even more variable, with the most robust feature being the large decreases over the remote oceans in RCP8.5. Through a regression analysis, we suggest that differences in emissions of non-methane volatile organic compounds and in the simulation of photolysis rates may be the main factors causing the differences in simulated present day OH and methane lifetime. Diversity in predicted changes between present day and future OH was found to be associated more strongly with differences in modelled temperature and stratospheric ozone changes. Finally, through perturbation experiments we calculated an OH feedback factor (F) of 1.24 from present day conditions (1.50 from 2100 RCP8.5 conditions) and a climate feedback on methane lifetime of 0.33+-0.13 yr/K, on average. Models that did not include interactive stratospheric ozone effects on photolysis showed a stronger sensitivity to climate, as they did not account for negative effects of climate-driven stratospheric ozone recovery on tropospheric OH, which would have partly offset the overall OH/methane lifetime response to climate change.

atmospheric composition

Refining T c Prediction in Hydrides via Symbolic‐Regression‐Enhanced Electron‐Localization‐Function‐Based Descriptors

Hydrogen‐based materials are able to possess extremely high superconducting critical temperatures, T c s , due to hydrogen's low atomic mass and strong electron–phonon interaction. Recently, a descriptor based on the Electron Localization Function (ELF) has enabled the rapid estimation of the T c of hydrogen‐containing compounds from electronic networking properties, but its applicability has been limited by the small size and homogeneity of the training dataset used. Herein, the model is re‐examined, compiling a publicly available combined dataset of 244 binary and ternary hydride superconductors. The analysis shows that though ELF‐based networking remains a valuable descriptor, its predictive power declines with increasing compositional complexity. However, by introducing the molecularity index, defined as the highest value of the ELF at which two hydrogen atoms connect, and applying symbolic regression, the accuracy of the predictions can be substantially enhanced. These results establish a more robust framework for assessing superconductivity in hydride materials, facilitating accelerated screening of novel candidates through integration with crystal structure prediction methods or high‐throughput searches.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Evidential Deep Learning: Enhancing Predictive Uncertainty Estimation for Earth System Science Applications

Abstract Robust quantification of predictive uncertainty is a critical addition needed for machine learning applied to weather and climate problems to improve the understanding of what is driving prediction sensitivity. Ensembles of machine learning models provide predictive uncertainty estimates in a conceptually simple way but require multiple models for training and prediction, increasing computational cost and latency. Parametric deep learning can estimate uncertainty with one model by predicting the parameters of a probability distribution but does not account for epistemic uncertainty. Evidential deep learning, a technique that extends parametric deep learning to higher-order distributions, can account for both aleatoric and epistemic uncertainties with one model. This study compares the uncertainty derived from evidential neural networks to that obtained from ensembles. Through applications of the classification of winter precipitation type and regression of surface-layer fluxes, we show evidential deep learning models attaining predictive accuracy rivaling standard methods while robustly quantifying both sources of uncertainty. We evaluate the uncertainty in terms of how well the predictions are calibrated and how well the uncertainty correlates with prediction error. Analyses of uncertainty in the context of the inputs reveal sensitivities to underlying meteorological processes, facilitating interpretation of the models. The conceptual simplicity, interpretability, and computational efficiency of evidential neural networks make them highly extensible, offering a promising approach for reliable and practical uncertainty quantification in Earth system science modeling. To encourage broader adoption of evidential deep learning, we have developed a new Python package, Machine Integration and Learning for Earth Systems (MILES) group Generalized Uncertainty for Earth System Science (GUESS) (MILES-GUESS) ( https://github.com/ai2es/miles-guess ), that enables users to train and evaluate both evidential and ensemble deep learning. Significance Statement This study demonstrates a new technique, evidential deep learning, for robust and computationally efficient uncertainty quantification in modeling the Earth system. The method integrates probabilistic principles into deep neural networks, enabling the estimation of both aleatoric uncertainty from noisy data and epistemic uncertainty from model limitations using a single model. Our analyses reveal how decomposing these uncertainties provides valuable insights into reliability, accuracy, and model shortcomings. We show that the approach can rival standard methods in classification and regression tasks within atmospheric science while offering practical advantages such as computational efficiency. With further advances, evidential networks have the potential to enhance risk assessment and decision-making across meteorology by improving uncertainty quantification, a longstanding challenge. This work establishes a strong foundation and motivation for the broader adoption of evidential learning, where properly quantifying uncertainties is critical yet lacking.

Schreck, John S.

Constraining Galaxy-Halo connection using machine learning

We investigate the potential of machine learning (ML) methods to model small-scale galaxy clustering for constraining Halo Occupation Distribution (HOD) parameters. Our analysis reveals that while many ML algorithms report good statistical fits, they often yield likelihood contours that are significantly biased in both mean values and variances relative to the true model parameters. This highlights the importance of careful data processing and algorithm selection in ML applications for galaxy clustering, as even seemingly robust methods can lead to biased results if not applied correctly. ML tools offer a promising approach to exploring the HOD parameter space with significantly reduced computational costs compared to traditional brute-force methods if their robustness is established. Using our ANN-based pipeline, we successfully recreate some standard results from recent literature. Properly restricting the HOD parameter space, transforming the training data, and carefully selecting ML algorithms are essential for achieving unbiased and robust predictions. Among the methods tested, artificial neural networks (ANNs) outperform random forests (RF) and ridge regression in predicting clustering statistics, when the HOD prior space is appropriately restricted. We demonstrate these findings using the projected two-point correlation function (w p (r p )), angular multipoles of the correlation function (ξ ℓ (r)), and the void probability function (VPF) of Luminous Red Galaxies from Dark Energy Spectroscopic Instrument mocks. Our results show that while combining w p (r p ) and VPF improves parameter constraints, adding the multipoles ξ 0 , ξ 2 , and ξ 4 to w p (r p ) does not significantly improve the constraints.

cosmology

Regression Analysis with the Directed Infusion of Data

Integrating artificial intelligence and machine learning tools into industry necessitates large-scale collaborative efforts that ensure the robust and accurate execution of downstream analytics such as time series prediction, uncertainty quantification, grid optimization, and condition monitoring. However, concerns related to data privacy pervade the nuclear industry due to the proprietary nature of its data and the possibility of data leakage. Legacy techniques such as encryption often require the explicit transmission of data to trustworthy parties, thereby inviting data leakage concerns. The ideal collaboration scenario avoids the explicit dissemination of data/code while maintaining experimental fidelity, which is currently accomplished using various techniques such as trusted execution environments, homomorphic encryption, differential privacy, and multimatrix masking. These techniques, however, often necessitate a trade-off between trust, efficiency, and utility. This article extends a previously proposed technique called the directed infusion of data (DIOD) that ensures data privacy, allows for scalable obfuscation, and combats the risk of data leakage without compromising utility. The experiments discussed in this article examine a regression-type scenario using DIOD with the goal of preserving the inferential link between two variables. Using the point-kinetics equations, regression experiments compare the performance of a model trained using the original data to that of a model trained using the obfuscated data, which produced identical results. Our claim is further strengthened by an information theoretic proof and experiment, which showed that the inferential content between variables remains the same after obfuscation, thereby avoiding the required communication of the proprietary data.

47 - OTHER INSTRUMENTATION

HighDimMixedModels.jl: Robust high-dimensional mixed-effects models across omics data

High-dimensional mixed-effects models are an increasingly important form of regression in which the number of covariates rivals or exceeds the number of samples, which are collected in groups or clusters. The penalized likelihood approach to fitting these models relies on a coordinate descent algorithm that lacks guarantees of convergence to a global optimum. Here, we empirically study the behavior of this algorithm on simulated and real examples of three types of data that are common in modern biology: transcriptome, genome-wide association, and microbiome data. Our simulations provide new insights into the algorithm’s behavior in these settings, and, comparing the performance of two popular penalties, we demonstrate that the smoothly clipped absolute deviation (SCAD) penalty consistently outperforms the least absolute shrinkage and selection operator (LASSO) penalty in terms of both variable selection and estimation accuracy across omics data. To empower researchers in biology and other fields to fit models with the SCAD penalty, we implement the algorithm in a Julia package, HighDimMixedModels.jl .

Gorstein, Evan

Uncertainty quantification for misspecified machine learned interatomic potentials

The use of high-dimensional regression techniques from machine learning has significantly improved the quantitative accuracy of interatomic potentials. Atomic simulations can now plausibly target quantitative predictions in a variety of settings, which has brought renewed interest in robust means to quantify uncertainties. In many practical settings where model complexity is constrained (e.g., due to performance considerations), misspecification — the inability of any one choice of model parameters to exactly match all training data — is a key contributor to errors that is often disregarded. Here, we employ a recent misspecification-aware regression technique to quantify parameter uncertainties, which is then propagated to a broad range of phase and defect properties in tungsten. The propagation is performed through both brute-force resampling and implicit Taylor expansion. The propagated misspecification uncertainties robustly quantify and bound errors on a broad range of material properties. We demonstrate application to recent foundational machine learning interatomic potentials, accurately predicting and bounding errors in MACE-MPA-0 energy predictions across the diverse materials project database.

36 MATERIALS SCIENCE

Computationally efficient and error aware surrogate construction for numerical solutions of subsurface flow through porous media

Limiting the injection rate to restrict the pressure below a threshold at a critical location can be an important goal of simulations that model the subsurface pressure between injection and extraction wells. The pressure is approximated by the solution of Darcy’s partial differential equation for a given permeability field. The subsurface permeability is modeled as a random field since it is known only up to statistical properties. This induces uncertainty in the computed pressure. Solving the partial differential equation for an ensemble of random permeability simulations enables estimating a probability distribution for the pressure at the critical location. These simulations are computationally expensive, and practitioners often need rapid online guidance for real-time pressure management. An ensemble of numerical partial differential equation solutions is used to construct a Gaussian process regression model that can quickly predict the pressure at the critical location as a function of the extraction rate and permeability realization. The Gaussian process surrogate analyzes the ensemble of numerical pressure solutions at the critical location as noisy observations of the true pressure solution, enabling robust inference using the conditional Gaussian process distribution. Our first novel contribution is to identify a sampling methodology for the random environment and matching kernel technology for which fitting the Gaussian process regression model scales as O ( n log n ) instead of the typical O ( n 3 ) rate in the number of samples n used to fit the surrogate. The surrogate model allows almost instantaneous predictions for the pressure at the critical location as a function of the extraction rate and permeability realization. Our second contribution is a novel algorithm to calibrate the uncertainty in the surrogate model to the discrepancy between the true pressure solution of Darcy’s equation and the numerical solution. Finally, although our method is derived for building a surrogate for the solution of Darcy’s equation with a random permeability field, the framework broadly applies to solutions of other partial differential equations with random coefficients.

54 ENVIRONMENTAL SCIENCES

End–to–End Metasurface Design for Temperature Imaging via Broadband Planck‐Radiation Regression

A theoretical framework is presented for temperature imaging from long-wavelength infrared (LWIR) thermal radiation (e.g., 8–12 µm) through the end-to-end design of a metasurface-optics frontend and a computational-reconstruction backend. A new nonlinear reconstruction algorithm, “Planck regression”, is introduced to reconstruct the temperature map from a gray scale sensor image, even in the presence of severe chromatic aberration, by exploiting black body and optical physics particular to thermal imaging. This algorithm is combined with an end-to-end approach that optimizes manufacturable, single-layer metasurfaces to yield the most accurate reconstruction. The designs demonstrate high-quality, noise-robust reconstructions of arbitrary temperature maps (including completely random images) in simulations of an ultra-compact thermal-imaging device. Here, it is also shown that Planck regression is much more generalizable to arbitrary images than a straightforward neural-network reconstruction, which requires a large training set of domain-specific images.

36 MATERIALS SCIENCE

Quantifying PM2.5-Meteorology Sensitivities in a Global Climate Model

Climate change can influence fine particulate matter concentrations (PM2.5) through changes in air pollution meteorology. Knowledge of the extent to which climate change can exacerbate or alleviate air pollution in the future is needed for robust climate and air pollution policy decision-making. To examine the influence of climate on PM2.5, we use the Geophysical Fluid Dynamics Laboratory Coupled Model version 3 (GFDL CM3), a fully-coupled chemistry-climate model, combined with future emissions and concentrations provided by the four Representative Concentration Pathways (RCPs). For each of the RCPs, we conduct future simulations in which emissions of aerosols and their precursors are held at 2005 levels while other climate forcing agents evolve in time, such that only climate (and thus meteorology) can influence PM2.5 surface concentrations. We find a small increase in global, annual mean PM2.5 of about 0.21 micro-g/cu m3 (5%) for RCP8.5, a scenario with maximum warming. Changes in global mean PM2.5 are at a maximum in the fall and are mainly controlled by sulfate followed by organic aerosol with minimal influence of black carbon. RCP2.6 is the only scenario that projects a decrease in global PM2.5 with future climate changes, albeit only by -0.06 micro-g/cu m (1.5%) by the end of the 21st century. Regional and local changes in PM2.5 are larger, reaching upwards of 2 micro-g/cu m for polluted (eastern China) and dusty (western Africa) locations on an annually averaged basis in RCP8.5. Using multiple linear regression, we find that future PM2.5 concentrations are most sensitive to local temperature, followed by surface wind and precipitation. PM2.5 concentrations are robustly positively associated with temperature, while negatively related with precipitation and wind speed. Present-day (2006-2015) modeled sensitivities of PM2.5 to meteorological variables are evaluated against observations and found to agree reasonably well with observed sensitivities (within 10e50% over the eastern United States for several variables), although the modeled PM2.5 is less sensitive to precipitation than in the observations due to weaker convective scavenging. We conclude that the hypothesized "climate penalty" of future increases in PM2.5 is relatively minor on a global scale compared to the influence of emissions on PM2.5 concentrations.

PM2.5

Robust Design Under Uncertainty in Quantum Error Mitigation

Error mitigation techniques are crucial to achieving near-term quantum advantage. Classical postprocessing of quantum computation outcomes is a popular approach for error mitigation, which includes methods, such as zero noise extrapolation, virtual distillation, and learning-based error mitigation. However, these techniques have limitations due to the propagation of uncertainty resulting from the finite shot number of a quantum measurement. In this work, we introduce general and unbiased methods for quantifying the uncertainty and error of error-mitigated observables based on the strategic sampling of error mitigation outcomes. We then extend our approach to demonstrate the optimization of performance and robustness of error mitigation under uncertainty. To illustrate our methods, we apply them to zero noise extrapolation and Clifford date regression in the ground state of the XY model simulated using depolarizing and International Business Machines Corporation (IBM) Toronto noise models, respectively. In particular, we optimize the choice of noise levels and the allocation of shots for zero noise extrapolation and the distribution of the training circuits for Clifford data regression. While our methods are readily applicable to any postprocessing-based error mitigation approach, in practice they must not be prohibitively expensive—even though they perform optimizations of the error mitigation hyperparameters requiring sampling of a statistical distribution of error mitigation outcomes. By leveraging surrogate-based optimization, we show that our methods can efficiently perform optimal design for a zero noise extrapolation implementation. We then further demonstrate the transferability of learned zero noise extrapolation hyperparameters to other similar circuits.

97 MATHEMATICS AND COMPUTING