Search NASASearch

SEARCH · Search NASA

Results for “robust regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

How Universal is the Relationship Between Remotely Sensed Vegetation Indices and Crop Leaf Area Index? A Global Assessment

Leaf Area Index (LAI) is a key variable that bridges remote sensing observations to the quantification of agroecosystem processes. In this study, we assessed the universality of the relationships between crop LAI and remotely sensed Vegetation Indices (VIs). We first compiled a global dataset of 1459 in situ quality-controlled crop LAI measurements and collected Landsat satellite images to derive five different VIs including Simple Ratio (SR), Normalized Difference Vegetation Index (NDVI), two versions of the Enhanced Vegetation Index (EVI and EVI2), and Green Chlorophyll Index (CI(sub Green)). Based on this dataset, we developed global LAI-VI relationships for each crop type and VI using symbolic regression and Theil-Sen (TS) robust estimator. Results suggest that the global LAI-VI relationships are statistically significant, crop-specific, and mostly non-linear. These relationships explain more than half of the total variance in ground LAI observations (R2 greater than 0.5), and provide LAI estimates with RMSE below 1.2 m2/m2. Among the five VIs, EVI/EVI2 are the most effective, and the crop-specific LAI-EVI and LAI-EVI2 relationships constructed by TS, are robust when tested by three independent validation datasets of varied spatial scales. While the heterogeneity of agricultural landscapes leads to a diverse set of local LAI-VI relationships, the relationships provided here represent global universality on an average basis, allowing the generation of large-scale spatial-explicit LAI maps. This study contributes to the operationalization of large-area crop modeling and, by extension, has relevance to both fundamental and applied agroecosystem research.

Vegetation Index

A robust approach to Gaussian process implementation

Abstract. Gaussian process (GP) regression is a flexible modeling technique used to predict outputs and to capture uncertainty in the predictions. However, the GP regression process becomes computationally intensive when the training spatial dataset has a large number of observations. To address this challenge, we introduce a scalable GP algorithm, termed MuyGPs, which incorporates nearest-neighbor and leave-one-out cross-validation during training. This approach enables the evaluation of large spatial datasets with state-of-the-art accuracy and speed in certain spatial problems. Despite these advantages, conventional quadratic loss functions used in the MuyGPs optimization, such as root mean squared error (RMSE), are highly influenced by outliers. We explore the behavior of MuyGPs in cases involving outlying observations and, subsequently, develop a robust approach to handle and mitigate their impact. Specifically, we introduce a novel leave-one-out loss function based on the pseudo-Huber function (LOOPH) that effectively accounts for outliers in large spatial datasets within the MuyGPs framework. Our simulation study shows that the LOOPH loss method maintains accuracy despite outlying observations, establishing MuyGPs as a powerful tool for mitigating unusual observation impacts in the large data regime. In the analysis of US ozone data, MuyGPs provides accurate predictions and uncertainty quantification, demonstrating its utility in managing data anomalies. Through these efforts, we advance the understanding of GP regression in spatial contexts.

Mukangango, Juliette

Real-time capable modeling of ICRF heating on NSTX and WEST via machine learning approaches

Abstract A real-time capable core Ion Cyclotron Range of Frequencies (ICRF) heating model on NSTX and WEST is developed. The model is based on two nonlinear regression algorithms, the random forest ensemble of decision trees and the multilayer perceptron neural network. The algorithms are trained on TORIC ICRF spectrum solver simulations of the expected flat-top operation scenarios in NSTX and WEST assuming Maxwellian plasmas. The surrogate models are shown to successfully capture the multi-species core ICRF power absorption predicted by the original model for the high harmonic fast wave and the ion cyclotron minority heating schemes while reducing the computational time by six orders of magnitude. Although these models can be expanded, the achieved regression scoring, computational efficiency and increased model robustness suggest these strategies can be implemented into integrated modeling frameworks for real-time control applications.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

A robust estimator of rainfall rate using differential reflectivity

Conventional estimator of rainfall rate using reflectivity factor and differential reflectivity Z(sub DR) becomes unstable when the measured values of Z(sub DR) are small due to measurement errors. An alternate estimator of rainfall rate using reflectivity factor and Z(sub DR) is derived, so that this estimator is fairly robust over the full dynamic range of reflectivity factor and Z(sub DR). Simulations are used to study the error structure of this robust estimator in comparison with the conventional estimator of rainfall rate. It is shown that the alternate estimator performs better than the conventional estimator of rainfall rate at all rainfall values. In particular the largest improvement of this estimator is proved to be in light rain. The robust estimator is obtained as a direct regression of rainfall rate against reflectivity factor and Z(sub DR) instead of solving for the drop size distribution.

Gorgucci, Eugenio

Enhancing Solar Power Forecasting with Regularized Constrained Quantile Regression Averaging and Bootstrapping Techniques

Probabilistic solar power forecasting (SPF) plays an essential role in optimizing power-grid operations by quantifying the forecast uncertainty. To improve the accuracy and robustness of probabilistic SPF, this paper introduces the regularized constrained quantile regression averaging (rCQRA) method to combine outputs from multiple PSPF models. In addition, a bootstrapping method was used to quantify model uncertainty, providing insights into the reliability and significance of each ensemble component. To evaluate its efficacy, the proposed rCQRA method is used to integrate four PSPF methods. The resulting SPF models are trained and validated using a real-world six-year dataset from a rooftop solar plant in the USA. The performance of the proposed rCQRA method is evaluated and compared with two benchmark methods under three categories of weather conditions. It is shown that the rCQRA method has superior performance in its forecast reliability, sharpness, and accuracy.

Ensemble learning, probabilistic solar power forec

Simultaneous Development and Robust Optimization of a Microstructure Dependent Material

Recent microstructure characterization techniques combined with Symbolic Regression(SR)analysis has been proven to generate white box plasticity models well suited for incorporation into FEA software.The current work builds upon those efforts and demonstrates the applicability of Sequential Monte-Carlo (SMC) methods within SR analysis to condense model development and robust optimization into a single, co-dependent process. In this project, SMC methods provide a mechanism through which the observed microstructure features and associated variability can be incorporated into the discovery phase of model development and simultaneously recover approximate parameter distributions through SR analysis. The demonstration utilized a data set consisting of tensile test results from a limited number of sample specimens with corresponding EBSD data from which microstructure features were characterized.The maximum threshold stress model in the Visco-Plastic Self-Consistent (VPSC) code developed by Los Alamos National Laboratories was calibrated using mechanical test data.Synthetic volume elements with statistically equivalent microstructure were generated with DREAM3Dbased on the observed EBSD data. VPSC was used to simulate the corresponding tensile test response for each of the synthetic volume elements. The simulated microstructure and tensile test data was used astraining datafor SMC-SR algorithm and the resulting model was validated with data from the original empirical data set.

Karl Garbrecht

An Alternative Flight Software Trigger Paradigm: Applying Multivariate Logistic Regression to Sense Trigger Conditions using Inaccurate or Scarce Information

In late 2014, NASA will fly the Orion capsule on a Delta IV-Heavy rocket for the Exploration Flight Test-1 (EFT-1) mission. For EFT-1, the Orion capsule will be flying with a new GPS receiver and new navigation software. Given the experimental nature of the flight, the flight software must be robust to the loss of GPS measurements. Once the high-speed entry is complete, the drogue parachutes must be deployed within the proper conditions to stabilize the vehicle prior to deploying the main parachutes. When GPS is available in nominal operations, the vehicle will deploy the drogue parachutes based on an altitude trigger. However, when GPS is unavailable, the navigated altitude errors become excessively large, driving the need for a backup barometric altimeter. In order to increase overall robustness, the vehicle also has an alternate method of triggering the drogue parachute deployment based on planet-relative velocity if both the GPS and the barometric altimeter fail. However, this velocity-based trigger results in large altitude errors relative to the targeted altitude. Motivated by this challenge, this paper demonstrates how logistic regression may be employed to automatically generate robust triggers based on statistical analysis. Logistic regression is used as a ground processor pre-flight to develop a classifier. The classifier would then be implemented in flight software and executed in real-time. This technique offers excellent performance even in the face of highly inaccurate measurements. Although the logistic regression-based trigger approach will not be implemented within EFT-1 flight software, the methodology can be carried forward for future missions and vehicles.

Smith, Kelly M.

Remote sensing of Pu in uranyl nitrate crystals using reflectance spectroscopy and chemometrics

Remote quantification of Pu(VI) (0–5 mol%) co-crystallized with U in uranyl nitrate hexahydrate (UNH) crystals was achieved in a glove box using reflectance spectroscopy coupled with chemometric modeling. Reflectance spectra were also acquired for Pu(IV) and Np(VI) (0–5 mol%) crystallized with UNH; revealing spectral features consistent with their solution-phase analogs. Principal component analysis revealed Pu(IV/VI) and Np(VI) concentrations as the primary source of variation in the data, informing the development of a supervised partial least squares regression model for Pu(VI). The resulting calibration demonstrated robust performance, with replicate root mean square errors near 10% and quantifiable limits near 0.2 mol% Pu(VI) relative to U. The Pu(VI) remained stable in the crystalline UNH matrix for at least one week with minimal reduction to Pu(IV). Notably, Pu(VI) and Np(VI) incorporation in UNH quenched U(VI) fluorescence while Pu(IV) did not. This study presents a noninvasive, spectroscopic approach for solid-state Pu quantification, with direct implications for material accountability and nuclear nonproliferation monitoring.

Sadergaski, Luke R. [Oak Ridge National Laboratory

D–MOPH–25: diverse MOF–molecule pairs for Henry’s constants prediction

Computational methods like grand-canonical Monte Carlo simulations and machine learning (ML) have accelerated metal–organic frameworks (MOF) exploration but are typically limited to a narrow range of adsorbates due to data availability and force field constraints. In this study, we introduce a dataset of diverse MOF–molecule pairs for Henry’s constant prediction, D–MOPH–25, which systematically explores a diverse chemical space by combining 113 molecular adsorbates with over 5000 MOF structures through an active learning process. D–MOPH–25 constitutes the most diverse adsorbate dataset used in any ML study of molecular adsorption in MOFs to date. Our workflow builds a benchmark for predicting Henry’s constants at 300 K, leveraging conformal prediction for uncertainty quantification. Assessment through Shannon entropy and uniform manifold approximation and projection confirms the comprehensiveness of D–MOPH–25 while highlighting the importance of robust classification to filter out unphysical data points in regression tasks. Although future enhancements in model architecture and sampling criteria could improve predictive performance, our dataset already spans the target space using only 2.31% of total possibilities. This comprehensive dataset facilitates assessment of model generalizability across adsorbate species and can establish a foundation for high-throughput MOF screening and ML-driven separation processes.

active learning

An Alternative Flight Software Trigger Paradigm: Applying Multivariate Logistic Regression to Sense Trigger Conditions Using Inaccurate or Scarce Information

In late 2014, NASA will fly the Orion capsule on a Delta IV-Heavy rocket for the Exploration Flight Test-1 (EFT-1) mission. For EFT-1, the Orion capsule will be flying with a new GPS receiver and new navigation software. Given the experimental nature of the flight, the flight software must be robust to the loss of GPS measurements. Once the high-speed entry is complete, the drogue parachutes must be deployed within the proper conditions to stabilize the vehicle prior to deploying the main parachutes. When GPS is available in nominal operations, the vehicle will deploy the drogue parachutes based on an altitude trigger. However, when GPS is unavailable, the navigated altitude errors become excessively large, driving the need for a backup barometric altimeter to improve altitude knowledge. In order to increase overall robustness, the vehicle also has an alternate method of triggering the parachute deployment sequence based on planet-relative velocity if both the GPS and the barometric altimeter fail. However, this backup trigger results in large altitude errors relative to the targeted altitude. Motivated by this challenge, this paper demonstrates how logistic regression may be employed to semi-automatically generate robust triggers based on statistical analysis. Logistic regression is used as a ground processor pre-flight to develop a statistical classifier. The classifier would then be implemented in flight software and executed in real-time. This technique offers improved performance even in the face of highly inaccurate measurements. Although the logistic regression-based trigger approach will not be implemented within EFT-1 flight software, the methodology can be carried forward for future missions and vehicles.

Smith, Kelly M.

An Alternative Flight Software Paradigm: Applying Multivariate Logistic Regression to Sense Trigger Conditions using Inaccurate or Scarce Information

In late 2014, NASA will fly the Orion capsule on a Delta IV-Heavy rocket for the Exploration Flight Test-1 (EFT-1) mission. For EFT-1, the Orion capsule will be flying with a new GPS receiver and new navigation software. Given the experimental nature of the flight, the flight software must be robust to the loss of GPS measurements. Once the high-speed entry is complete, the drogue parachutes must be deployed within the proper conditions to stabilize the vehicle prior to deploying the main parachutes. When GPS is available in nominal operations, the vehicle will deploy the drogue parachutes based on an altitude trigger. However, when GPS is unavailable, the navigated altitude errors become excessively large, driving the need for a backup barometric altimeter to improve altitude knowledge. In order to increase overall robustness, the vehicle also has an alternate method of triggering the parachute deployment sequence based on planet-relative velocity if both the GPS and the barometric altimeter fail. However, this backup trigger results in large altitude errors relative to the targeted altitude. Motivated by this challenge, this paper demonstrates how logistic regression may be employed to semi-automatically generate robust triggers based on statistical analysis. Logistic regression is used as a ground processor pre-flight to develop a statistical classifier. The classifier would then be implemented in flight software and executed in real-time. This technique offers improved performance even in the face of highly inaccurate measurements. Although the logistic regression-based trigger approach will not be implemented within EFT-1 flight software, the methodology can be carried forward for future missions and vehicles

Smith, Kelly

Areal Distribution of the Oxygen-Isotope Ratio in Greenland

Mean values of the oxygen-isotope ratio relative to standard mean ocean water reported for 46 sites on the Greenland ice sheet are compiled together with data on mean annual surface temperature, latitude, 6180 elevation, and mean annual shortest distance to the open ocean denoted by the 10% sea-ice concentration boundary. Stepwise regression analyses, with 6180 as the dependent variable, define two robust models. In the forward mode at the 99.9% confidence level, only temperature enters the model. In the backward mode at the 95% confidence level, only temperature, latitude, and distance to the open ocean remain in the model. Inversions of the models on the basis of 160 gridpoint locations 100 km apart in the area delimited by the surface equilibrium line produce four contoured distributions of 6"0. Two distributions are based on the bivariate model and two on the multivariate model. The second distribution for each model is obtained substituting mean annual surface-temperature values obtained from the Nimbus-7 Temperature Humidity Infrared Radiometer (THIR) database. All four distributions are considered valid, and differences between them are evaluated using contoured anomaly maps. It is suggested that the inversion of the multivariate model using THIR data provides the more reliable pattern for studies of atmospheric advection or for the derivation of ice-flow adjustments for 6180 series obtained from deep-core or ablation-zone sites.

Zwally, H. Jay

Standardising the “Gregory method” for calculating equilibrium climate sensitivity

The equilibrium climate sensitivity (ECS) – the equilibrium global mean temperature response to a doubling of atmospheric CO 2 – is a high-profile metric for quantifying the Earth system's response to human-induced climate change. A widely applied approach to estimating the ECS is the “Gregory method” (Gregory et al., 2004), which uses an ordinary least squares (OLS) regression between the net radiative flux, N, and surface air temperature anomalies, ΔT, from a 150 year experiment in which atmospheric CO 2 concentrations are quadrupled. The ECS is determined by extrapolating the linear fit to N=0, i.e. the ΔT-intercept, indicating the point at which the system is back in equilibrium. This method has been used to compare ECS estimates across the CMIP5 and CMIP6 ensembles and will likely be a key diagnostic for CMIP7. Despite its widespread application, there is little consistency or transparency between studies in how the climate model data is processed prior to the regression, leading to potential discrepancies in ECS estimates. We identify 32 alternative data processing pathways, varying by differences in global mean weighting, net radiative flux variable, anomaly calculation method, and linear regression fit. Using 44 CMIP6 models, we systematically assess the impact of these choices on ECS estimates and calculate uncertainty ranges using two bootstrap approaches. While the inter-model ECS range is insensitive to the data processing pathway, individual outlier models exhibit notable differences. Approximating a model's native grid cell area (if irregular) with cosine of the latitude can decrease the ECS by 11 %, the choice of N-variable can change the ECS by 6 %, and some anomaly calculation methods can introduce spurious temporal correlations in the processed data. Beyond data processing choices, we also evaluate an alternative linear regression method – total least squares (TLS) – which has a more statistically robust basis than OLS. However, for consistency with previous literature, and given TLS may reduce the ECS compared to OLS (by up to 24 %), thereby making a known bias in the Gregory method worse, we do not feel there is sufficient clarity to recommend a transition to TLS in all cases. To improve reproducibility and comparability in future studies, we recommend a standardised Gregory method: weighting the global mean by cell area, using the top of the atmosphere (as opposed to the top of model) N-variable, and calculating anomalies by first applying a rolling average to the preindustrial control timeseries then subtracting from the raw CO 2 quadrupling experiment. This approach accounts for model drift while reducing noise in the data to best meet the pre-conditions of the linear regression. While CMIP6 results of the multi-model mean ECS appear insensitive to these processing choices, similar assumptions may not hold for CMIP7, underscoring the need for standardised data preparation in future climate sensitivity assessments.

Geosciences

Multilevel Probit Regression for 3-Alternative Forced Choice Audibility Testing

A recent psychoacoustic test at NASA Langley generated a dataset of 3-alternative forced-choice responses for 40 subjects that measured the audibility of a tone complex in a shaped broadband masker. The task was completed by 4 subjects at a time in a small theatre-like environment using predetermined stimuli levels. These data were subject to 4 forms of probit regression: a “complete pooling” analysis in which all data from the test was fit with one curve, two forms of “no pooling” analyses in which subjects’ data were treated individually (using both packaged and custom software), and a “partial pooling” analysis in which multilevel-regression software fit both individual curves as well as population-level parameters at the same time. The results of the analyses are compared in terms of both individual- and population-level parameters. Partial pooling appears to give the most consistent results at both levels, as well as provide the most robustness among the packaged approaches (albeit with added complexity over single-level regression). These results are also contained in a recent NASA technical memorandum entitled “Comparisons of Analysis Methods Applied to Alternative Forced Choice Audibility Data.”

Psychoacoustics

Extreme Temperature Cryptography Based On Nitrogen-Incorporated Ultrananocrystalline Diamond

Physical entropy sources that remain stable under extreme temperatures are essential for cryptography in emerging technological frontiers in deep space exploration, geothermal energy harvesting, and nuclear energy. However, conventional semiconductor platforms fail to generate stable and reliable cryptographic keys above 200 degrees C due to performance degradation. Here, we report a diamond-based cryptographic primitive that exploits the defect-rich sp 2 -bonded grain boundary network in nitrogen-incorporated ultrananocrystalline diamond (n-UNCD) film as a robust entropy source to generate cryptographic keys that remain operationally stable even after enduring extreme temperatures of 700 degrees C for 54 h while also surviving thermal cycling between room temperature and 700 degrees C for 48 h. The strength of the generated keys is assessed through several cryptographic metrics such as bit uniformity, entropy, hamming distances, and correlation coefficients, all of which are found to be near their respective ideal values. Moreover, the generated keys pass the NIST SP 800 and SP 800-90B tests and are also resilient to supply bias variations and a regression-based machine learning attack model based on the Fourier series. The robustness of the keys is attributed to the better thermal stability and chemical inertness of the n-UNCD film. This is supported by high-resolution energy-dispersive X-ray spectroscopy (EDS), which shows no significant lateral diffusion of metal atoms into the n-UNCD layer, and by Raman spectroscopy, which reveals no significant changes in the bonding configuration of the n-UNCD structure. Our findings highlight the remarkable potential of n-UNCD film for extreme environment cryptography by expanding the operational limits of conventional hardware security platforms.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Discovering nuclear models from symbolic machine learning

Numerous phenomenological nuclear models have been proposed to describe specific observables within different regions of the nuclear chart. However, developing a unified model that describes the complex behavior of all nuclei remains an open challenge. Here, we explore whether symbolic Machine Learning (ML) can rediscover traditional nuclear physics models or identify alternatives with improved simplicity, fidelity, and predictive power. To address this challenge, we developed a Multi-objective Iterated Symbolic Regression approach that handles symbolic regressions over multiple target observables, accounts for experimental uncertainties and is robust against high-dimensional problems. As a proof of principle, we applied this method to describe the nuclear binding energies and charge radii of light and medium mass nuclei. Our approach identified simple analytical relationships based on the number of protons and neutrons, providing interpretable models with precision comparable to state-of-the-art nuclear models. Additionally, we integrated this ML-discovered model with an existing complementary model to estimate the limits of nuclear stability. These results highlight the potential of symbolic ML to develop accurate nuclear models and guide our description of complex many-body problems.

Nuclear structure

Data-Driven Mean-Corrected Recursive Estimation-Based Optimal DER Dispatch for Distribution System Voltage Control

Recent advances in smart inverters offer opportunities to mitigate adverse grid impacts caused by high penetrations of distributed photovoltaics (PV) in distribution grids, such as voltage violations. Here, this paper proposes a novel measurement-driven optimal power flow (OPF)-based distributed energy resource management system (DERMS) voltage regulation via recursive sensitivity estimation informed coordinated control of distributed PV inverters. The proposed approach leverages available grid and controllable DER measurements, eliminating reliance on system model information while being adaptive and robust to volatile operating conditions. A mean-corrected recursive ridge regression (MCRRR) algorithm is proposed for sensitivity estimation, continuously refining the sensitivity model through a closed-form solution. It effectively manages varying grid operating conditions, such as changes in power injections and topology reconfiguration, to facilitate a time-varying update of the Load Sensitivity Factors (LSF). The proposed approach is formulated as a linear programming (LP) problem and is thus scalable to larger-scale distribution systems. Its effectiveness and efficiency are demonstrated on a realistic distribution feeder with high PV penetrations in Southern California, USA.

14 SOLAR ENERGY

Joint Modeling of Quasar Variability and Accretion Disk Reprocessing Using Latent Stochastic Differential Equations

Quasars are bright active galactic nuclei powered by the accretion of matter around supermassive black holes at the center of galaxies. Their stochastic brightness variability depends on the physical properties of the accretion disk and black hole. The upcoming Rubin Observatory Legacy Survey of Space and Time (LSST) is expected to observe tens of millions of quasars, so there is a need for efficient techniques like machine learning that can handle the large volume of data. Quasar variability is believed to be driven by an X-ray corona, which is reprocessed by the accretion disk and emitted as UV/optical variability. We are the first to introduce an auto-differentiable simulation of the accretion disk and reprocessing. We use the simulation as a direct component of our neural network to jointly model the driving variability and reprocessing, trained with supervised learning on simulated LSST-like 10 yr quasar light curves. We encode the light curves using a transformer encoder, and the driving variability is reconstructed using latent stochastic differential equations, a physically motivated generative deep learning method that can model continuous-time stochastic dynamics. By embedding the physical processes of the driving signal and reprocessing into our network, we achieve a model that is more robust and interpretable. We demonstrate that our model outperforms a Gaussian process regression baseline and can infer accretion disk parameters and time delays between wave bands, even for out-of-distribution driving signals. Our approach provides a powerful framework that can be adapted to solve other inverse problems in multivariate time series.

Fagin, Joshua [City Univ. of New York (CUNY), NY (