Search NASA⌕ Search

SEARCH · Search NASA

Results for “Regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Surrogate-Based Autotuning for Randomized Sketching Algorithms in Regression Problems

Algorithms from Randomized Numerical Linear Algebra (RandNLA) are known to be effective in handling high-dimensional computational problems, providing high-quality empirical performance as well as strong probabilistic guarantees. However, their practical application is complicated by the fact that the user needs to set various algorithm-specific tuning parameters which are different from those used in traditional NLA. This paper demonstrates how a surrogate-based autotuning approach can be used to address fundamental problems of parameter selection in RandNLA algorithms. In particular, we provide a detailed investigation of surrogate-based autotuning for sketch-and-precondition (SAP)-based randomized least squares methods, which have been one of the great success stories in modern RandNLA. Empirical results show that our surrogate-based autotuning approach can achieve near-optimal performance with much less tuning cost than a random search (up to about 7.6x fewer trials of different parameter configurations). Moreover, while our experiments focus on least squares, our results demonstrate a general-purpose autotuning pipeline applicable to any kind of RandNLA algorithm.

Cho, Younghyun↗

Open Call LDRD: Physically Informed Autoencoders for Galactic Redshift Regression

Physical constraints have been suggested to make neural network models more generalizable, act scientifically plausible, and be more data-efficient over unconstrained baselines. In this report, we present preliminary work on evaluating the effects of adding soft physical constraints to computer vision neural networks trained to estimate the conditional density of redshift on input galaxy images for the Sloan Digital Sky Survey. We introduce physically motivated soft constraint terms that are not implemented with differential or integral operators. We frame this work as a simple ablation study where the effect of including soft physical constraints is compared to an unconstrained baseline. We compare networks using standard point estimate metrics for photometric redshift estimation, as well as metrics to evaluate how faithful our conditional density estimate represents the probability over the ensemble of our test dataset. We find no evidence that the implemented soft physical constraints are more effective regularizers than augmentation.

97 MATHEMATICS AND COMPUTING↗

Meteoric 10Be Flux Calibration Data for the East River Watershed, Colorado, USA

This data package contains tabular and geospatial data used to quantify and model meteoric beryllium-10 fluxes in the East River watershed, Colorado, USA. The tabular component includes calibration-site data from five glacial moraine sites and includes environmental variables used to evaluate spatial controls on meteoric 10Be delivery, including elevation, mean annual precipitation (MAP), mean snow depth, and mean snow water equivalent (SWE). These site-level data were used to compare observed fluxes with environmental gradients across the watershed and to evaluate the effects of erosion correction on flux estimates. The package also includes supporting slope and curvature values used to assess topographic inputs to the erosion analysis. A second component of the data package contains updated manuscript tables and regression outputs used to summarize the relationships between meteoric 10Be flux and environmental predictors. These tables include meteoric 10Be sample information and AMS results, site-level environmental values, site-level meteoric 10Be inventory and flux values, watershed-averaged predicted fluxes, soil bulk density measurements, fine-fraction values, soil pH measurements, and regression statistics including slope, intercept, coefficient of determination, and p-value. The regression products include both standard linear regressions and regressions in which the intercept is constrained to pass through zero, and they support the analyses presented in the companion manuscript. Together, these tabular files provide the numerical basis for the manuscript tables and the regression-based interpretation of meteoric 10Be flux variability in a snow-dominated mountain watershed. The geospatial component of the package consists of GeoTIFF raster files used to generate the map products presented in Figures 2 and 6 of the companion manuscript. These rasters represent watershed-scale spatial layers for environmental variables and regression-based predictions of meteoric 10Be flux. This dataset contains comma-separated values files (.csv), Microsoft Excel files (.xlsx), GeoTIFF raster files (.tif), and upporting metadata files, including CSV data dictionaries and readme text files (.csv, .txt). The tabular files can be opened with standard spreadsheet software, and the raster files can be viewed and analyzed in GIS software such as ArcGIS Pro or QGIS. Together, these files document the numerical and spatial datasets used to calibrate and predict meteoric 10Be delivery in the East River watershed.

East River↗

Investigating the ecological fallacy through sampling distributions constructed from finite populations

Correlation coefficients and linear regression values computed from group averages can differ from correlation coefficients and linear regression values computed using individual scores. This observation known as the ecological fallacy often assumes that all the individual scores are available from a population. In many situations, one must use a sample from the larger population. In such cases, the computed correlation coefficient and linear regression values will depend on the sample that is chosen and the underlying sampling distribution. The sampling distribution of correlation coefficients and linear regression values for group averages will be identical to the sampling distribution for individuals for normally distributed variables for random samples drawn from infinitely large continuous distributions. However, data that is acquired in practice is often acquired when sampling without replacement from a finite population. Our objective is to demonstrate through Monte Carlo simulations that the sampling distributions for correlation and linear regression will also be similar for individuals and group averages when sampling without replacement from normally distributed variables. These simulations suggest that when a random sample from a population is selected, the correlation coefficients and linear regression values computed from individual scores will not be more accurate in estimating the entire population values compared to samples when group averages are used as long as the sample size is the same.

97 MATHEMATICS AND COMPUTING↗

Machine Learning–Augmented Laser-Induced Breakdown Spectroscopy for Spectral Discrimination of Iron Oxalates

Enhanced characterization and phase identification of post-PUREX Pu Oxalates (PuOXA) are pivotal for nonproliferation and pre-detonation nuclear forensics. Despite significant advances in the characterization of PuO 2 samples, little is known about the impact of both the chemical structure and oxidation states of PuOXA (i.e., Pu(III) and Pu(IV)) have on optical emission signatures. Here, we demonstrate the analytical capabilities of laser-induced breakdown spectroscopy (LIBS) applied to Fe(II) and Fe(III) oxalate samples as surrogates for PuOXA, highlighting the discriminating features in the LIBS emission spectra arising from differences in the oxidation states within mixed FeOXA samples. We report the enhancement of spectral feature selection using Principal Component Analysis (PCA), which enables the analytical superiority of machine learning algorithms such as Linear Discriminant Analysis (LDA), Quadratic Discriminant Analysis (QDA), Partial Least Squares Regression (PLSR), Support Vector Regression (SVR), and Random Forest Regression (RFR) over conventional univariate techniques for phase discrimination and chemometric analysis. Cluster analysis revealed how both matrix effects and laser ablation influence cluster separability by introducing spectral artifacts that misdirect the maximization of variance. PCA-selected emission lines were used in the regression models, demonstrating that both univariate and multivariate linear regression models (i.e., PLSR and SVR) can achieve acceptable performance, with machine learning models outperforming conventional calibration regressions. Furthermore, the application of non-linearly activated PCA-selected emission lines illustrates how simplifying the data while retaining captured variance enables the use of less complex and more computationally efficient models. Furthermore, this is particularly evident in the underperformance of RFR, which suffers from increased computational costs and overfitting owing to its high complexity.

Oxalates↗

Evaluating the potential of short-term instrument deployment to improve distributed wind resource assessment

Distributed wind projects, which are connected at the distribution level of an electricity system or in off-grid applications to serve specific or local energy needs, often rely solely on wind resource models to establish wind speed and energy generation expectations. Historically, anemometer loan programs have provided an affordable avenue for more accurate onsite wind resource assessment, and the lowering cost of lidar systems has shown similar advantages for more recent assessments. While a full 12 months of onsite wind measurement is the standard for correcting model-based long-term wind speed estimates for utility-scale wind farms, the time and capital investment involved in gathering onsite measurements must be reconciled with the energy needs and funding opportunities that drive expedient deployment of distributed wind projects. Much literature exists to quantify the performance of correcting long-term wind speed estimates with 1 or more years of observational data, but few studies explore the impacts of correcting with months-long observational periods. This study aims to answer the question of how short you can go in terms of the observational time period needed to make impactful improvements to model-based long-term wind speed estimates. Three algorithms, multivariable linear regression, adaptive regression splines, and regression trees, are evaluated for their skill at correcting long-term wind resource estimates from the European Centre for Medium-Range Weather Forecasts Reanalysis version 5 (ERA5) using months-long periods of observational data from 66 locations across the US. On average, correction with even 1 month of observations provides significant improvement over the baseline ERA5 wind speed estimates and produces median bias magnitudes and relative errors within 0.22 m s −1 and 4 percentage points of the median bias magnitudes and relative errors achieved using the standard 12 months of data for correction. However, in cases when the shortest observational periods (1 to 2 months) used for correction are not well correlated with the overlapping ERA5 reference, the resultant long-term wind speed errors are worse than those produced using ERA5 without correction. Summer months, which are characterized by weaker relative wind speeds and standard deviations for most of the evaluation sites, tend to produce the worst results for long-term correction using months-long observations. The three tested algorithms perform similarly for long-term wind speed bias; however, regression trees perform notably worse than multivariable linear regression and adaptive regression splines in terms of correlation when using 6 months or less of observational data for correction. Translating the analysis to wind energy, median relative errors in the capacity factor are on average within 10 % using 1 month of training. If the observation period used for correction is not well correlated with the reference data, however, misrepresentation of the observed capacity factor can be substantial. The risk associated with poor correlation between the observed and reference datasets decreases with increasing training period length. In the worst-correlation scenarios, the median capacity factor relative errors from using 1, 3, and 6 months are within 47 %, 26 %, and 16 %, respectively.

17 WIND ENERGY↗

Generalized Bayesian MARS: Tools for Stochastic Computer Model Emulation

The multivariate adaptive regression spline (MARS) approach of Friedman and its Bayesian counterpart are effective approaches for the emulation of computer models. The traditional assumption of Gaussian errors limits the usefulness of MARS, and many popular alternatives, when dealing with stochastic computer models. Here, we propose a generalized Bayesian MARS (GBMARS) framework which admits the broad class of generalized hyperbolic distributions as the induced likelihood function. This allows us to develop tools for the emulation of stochastic simulators which are parsimonious, scalable, and interpretable and require minimal tuning, while providing powerful predictive and uncertainty quantification capabilities. GBMARS is capable of robust regression with t distributions, quantile regression with asymmetric Laplace distributions, and a general form of “Normal-Wald” regression in which the shape of the error distribution and the structure of the mean function are learned simultaneously. We demonstrate the effectiveness of GBMARS on various stochastic computer models, and we show that it compares favorably to several popular alternatives.

97 MATHEMATICS AND COMPUTING↗

Machine Learning-Based Process Control for Injection Molding of Recycled Polypropylene

The increased interest in artificial intelligence in manufacturing has driven the adoption of machine learning to optimize processes and improve efficiency. A key challenge in injection molding is the variability of recycled materials, which affects part quality and processing stability. This study presents a novel closed-loop process control approach for injection molding, leveraging machine learning to adaptively predict processing inputs and quality outcomes. The methodology was tested on five blends of recycled polypropylene (rPP), using artificial neural networks (ANNs), linear regression, and polynomial regression to model the relationships between material properties and process parameters. The dataset was split 80/20 into training and testing sets. The ANN model was implemented using TensorFlow and Keras, with six hidden layers of 32 neurons per layer, ReLU activation, and an Adam optimizer. Empirical tuning and early stopping were used to optimize performance and prevent overfitting. Predictions were evaluated based on mean absolute error (MAE), mean squared error (MSE), and percentage error. The results showed that yield stress, ultimate elongation, and part weight were accurately predicted within a 5% error for linear and polynomial regression models and within a 10% error for the ANN. However, modulus predictions were less reliable, with errors of ~11% for ANN and linear regression and ~40% for polynomial regression, reflecting the inherent variability of this property in rPP blends. Predictions of processing inputs had errors ranging from 3% to 25%, depending on the model and response variable. No single modeling approach was consistently superior across all responses, highlighting the complexity of the relationship between material properties, process parameters, and quality metrics. Overall, the work demonstrates that closed-loop process control, powered by machine learning, can effectively predict key quality parameters in injection molding of recycled materials. The proposed approach can improve process stability and material utilization, facilitating increased adoption of sustainable materials.

Krantz, Joshua↗

Standardising the “Gregory method” for calculating equilibrium climate sensitivity

The equilibrium climate sensitivity (ECS) – the equilibrium global mean temperature response to a doubling of atmospheric CO 2 – is a high-profile metric for quantifying the Earth system's response to human-induced climate change. A widely applied approach to estimating the ECS is the “Gregory method” (Gregory et al., 2004), which uses an ordinary least squares (OLS) regression between the net radiative flux, N, and surface air temperature anomalies, ΔT, from a 150 year experiment in which atmospheric CO 2 concentrations are quadrupled. The ECS is determined by extrapolating the linear fit to N=0, i.e. the ΔT-intercept, indicating the point at which the system is back in equilibrium. This method has been used to compare ECS estimates across the CMIP5 and CMIP6 ensembles and will likely be a key diagnostic for CMIP7. Despite its widespread application, there is little consistency or transparency between studies in how the climate model data is processed prior to the regression, leading to potential discrepancies in ECS estimates. We identify 32 alternative data processing pathways, varying by differences in global mean weighting, net radiative flux variable, anomaly calculation method, and linear regression fit. Using 44 CMIP6 models, we systematically assess the impact of these choices on ECS estimates and calculate uncertainty ranges using two bootstrap approaches. While the inter-model ECS range is insensitive to the data processing pathway, individual outlier models exhibit notable differences. Approximating a model's native grid cell area (if irregular) with cosine of the latitude can decrease the ECS by 11 %, the choice of N-variable can change the ECS by 6 %, and some anomaly calculation methods can introduce spurious temporal correlations in the processed data. Beyond data processing choices, we also evaluate an alternative linear regression method – total least squares (TLS) – which has a more statistically robust basis than OLS. However, for consistency with previous literature, and given TLS may reduce the ECS compared to OLS (by up to 24 %), thereby making a known bias in the Gregory method worse, we do not feel there is sufficient clarity to recommend a transition to TLS in all cases. To improve reproducibility and comparability in future studies, we recommend a standardised Gregory method: weighting the global mean by cell area, using the top of the atmosphere (as opposed to the top of model) N-variable, and calculating anomalies by first applying a rolling average to the preindustrial control timeseries then subtracting from the raw CO 2 quadrupling experiment. This approach accounts for model drift while reducing noise in the data to best meet the pre-conditions of the linear regression. While CMIP6 results of the multi-model mean ECS appear insensitive to these processing choices, similar assumptions may not hold for CMIP7, underscoring the need for standardised data preparation in future climate sensitivity assessments.

Geosciences↗

Multi deep learning-based stochastic microstructure reconstruction and high-fidelity micromechanics simulation of time-dependent ceramic matrix composite response

A multi deep learning-based framework is developed for efficient, automated microstructure reconstruction and generation of stochastic representative volume elements (SRVEs) with periodic boundary conditions (PBCs) for accurate modeling of ceramic matrix composite (CMC) response. The methodology comprises a convolutional neural network coupled with regression layers to act as a vanilla regression network for semantic segmentation of the microstructure, allowing accurate characterization of the phases and their distributions at the microscale. Scanning electron microscope and confocal microscope are used to obtain C/SiNC and SiC/SiNC CMCs micrographs for vanilla regression testing. Microstructure variability in terms of fiber volume fraction and porosity are quantified through the output regression layer, ensuring accurate representation of material variability in SRVE construction. Generative adversarial network (GAN) and its variants are designed to produce high-fidelity SRVE, spanning CMCs microstructure variability space. A circular padding algorithm is developed to generate SRVEs with PBCs during training of GANs. The accuracy of the generated SRVEs is established through micromechanics simulations, where an efficient formulation of the high-fidelity generalized methods of cells (HFGMC) approach is used to compute the effective mechanical properties. Furthermore, an iterative algorithm is implemented in the HFGMC solver to simulate time-dependent deformation of SiC/SiNC subjected to creep loading conditions.

36 MATERIALS SCIENCE↗