Search NASA⌕ Search

SEARCH · Search NASA

Results for “regularized regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

An ℓ 0 ℓ 2 -norm regularized regression model for construction of robust cluster expansions in multicomponent systems

In this work we introduce ℓ 0 ℓ 2 -norm regularization and hierarchy constraints into linear regression for the construction of cluster expansions to describe configurational disorder in materials. The approach is implemented through mixed integer quadratic programming (MIQP). The ℓ 2 -norm regularization is used to suppress intrinsic data noise, while the ℓ 0 -norm is used to penalize the number of nonzero elements in the solution. The hierarchy relation between clusters imposes relevant physics and is naturally included by the MIQP paradigm. As such, sparseness and cluster hierarchy can be well optimized to obtain a robust, converged set of effective cluster interactions with improved physical meaning. We demonstrate the effectiveness of ℓ 0 ℓ 2 -norm regularization in two high-component disordered rocksalt cathode material systems, where we compare the cross-validation, convergence speed, and the reproduction of phase diagrams, voltage profiles, and Li-occupancy energies with those of the conventional ℓ 1 -norm regularized cluster expansion models.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Improving Prediction of Peroxide Value of Edible Oils Using Regularized Regression Models

We present four unique prediction techniques, combined with multiple data pre-processing methods, utilizing a wide range of both oil types and oil peroxide values (PV) as well as incorporating natural aging for peroxide creation. Samples were PV assayed using a standard starch titration method, AOCS Method Cd 8-53, and used as a verified reference method for PV determination. Near-infrared (NIR) spectra were collected from each sample in two unique optical pathlengths (OPLs), 2 and 24 mm, then fused into a third distinct set. All three sets were used in partial least squares (PLS) regression, ridge regression, LASSO regression, and elastic net regression model calculation. While no individual regression model was established as the best, global models for each regression type and pre-processing method show good agreement between all regression types when performed in their optimal scenarios. Furthermore, small spectral window size boxcar averaging shows prediction accuracy improvements for edible oil PVs. Best-performing models for each regression type are: PLS regression, 25 point boxcar window fused OPL spectral information RMSEP = 2.50; ridge regression, 5 point boxcar window, 24 mm OPL, RMSEP = 2.20; LASSO raw spectral information, 24 mm OPL, RMSEP = 1.80; and elastic net, 10 point boxcar window, 24 mm OPL, RMSEP = 1.91. The results show promising advancements in the development of a full global model for PV determination of edible oils.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

MINLP for regularized symbolic regression with applications to data-driven modeling of critical minerals processes

The poster summarizes recent advances in symbolic regression developed as part of the PrOMMiS project over the past year. In particular, it describes the comparison of surrogates for critical minerals (CM) & rare earth element (REE) recovery flowsheets obtained via symbolic regression and ALAMO. It also compares the predictive ability and solvability of optimization models that incorporate these surrogates.

36 MATERIALS SCIENCE↗

Adaptive spectra-to-exposure conversion using ridge regularized polynomial response models

Real-time gamma spectra-to-exposure conversion in aerial and ground monitoring commonly relies on calibration-derived, detector- or system-specific conversion coefficients that are assumed to generalize across operational environments. In practice, deployment specific differences in spectral composition and transport conditions can introduce systematic bias relative to reference instruments, motivating methods that adapt coefficients using minimal field supervision while explicitly limiting overfitting. In this work, we present a conservative coefficient adaptation framework that updates a baseline polynomial energy-weighting function using ridge-regularized regression, with leave-one-out cross-validation (LOOCV) used to select the regularization strength. The findings support ridge-constrained minimal-supervision adaptation as a practical mechanism to suppress site-specific bias without destabilizing a calibration-derived baseline.

61 RADIATION PROTECTION AND DOSIMETRY↗

Numerical characterization of support recovery in sparse regression with correlated design

Sparse regression is employed in diverse scientific settings as a feature selection method. A pervasive aspect of scientific data is the presence of correlations between predictive features. These correlations hamper both feature selection and estimation and jeopardize conclusions drawn from estimated models. On the other hand, theoretical results on sparsity-inducing regularized regression have largely addressed conditions for selection consistency via asymptotics, and disregard the problem of model selection, whereby regularization parameters are chosen. In this numerical study, we address these issues through exhaustive characterization of the performance of several regression estimators, coupled with a range of model selection strategies. These estimators and selection criteria were examined across correlated regression problems with varying degrees of signal to noise, distributions of non-zero model coefficients, and model sparsity. Our results reveal a fundamental tradeoff between false positive and false negative control in all regression estimators and model selection criteria examined. Additionally, we numerically explore a transition point modulated by the signal-to-noise ratio and spectral properties of the design covariance matrix at which the selection accuracy of all considered algorithms degrades. Overall, we find that SCAD coupled with BIC or empirical Bayes model selection performs the best feature selection across the regression problems considered.

97 MATHEMATICS AND COMPUTING↗

Code Coverage Status of ARC Code-DIF3D

The Argonne Reactor Code (ARC) software system supports users in their fast reactor design goals by providing neutronic, thermal-hydraulic, and structural analysis capabilities. DIF3D plays a pivotal role in the ARC system as the primary homogenized assembly neutronic calculation methodology for fast reactor problems. Over its 40 years history, ARC software usage with DIF3D has been applied to numerous fast and thermal spectrum reactor analysis projects with good to excellent comparison against experiments. With continued improvement of computation resources, many of the geometry modeling capabilities in DIF3D that were primarily used in low order schemes are not really needed anymore. Today, the diffusion and transport capabilities of DIF3D-VARIANT are primarily used in the reactor design process with some scattered usage of DIF3D-FD and DIF3D-Nodal. In recent work, the DIF3D software verification was completed for DIF3D-FD and DIF3D-VARIANT on the geometry options used in the Versatile Test Reactor project. While we can be confident that these capabilities of DIF3D are well used and thus trusted, it does not demonstrate that all possible input options of DIF3D are actually working, but just those that were tested as part of VTR are and that they are correct. Thus, the purpose of the present work is to identify a set of test problems for DIF3D and assess the code coverage of DIF3D for those test problems. The goal is to document what parts of the existing DIF3D code are touched by the set of test problems and which are not. Because the verification work done on DIF3D-VARIANT and DIF3D-FD was focused on the most common uses of DIF3D for fast reactor analysis, the code coverage assessment of those capabilities is the highest priority. This will ensure that nothing is being missed by the existing verification test problems that DIF3D relies upon. The DIF3D-Nodal capability will also be inspected for code coverage as part of this work to further ensure that regular regression testing of DIF3D will trap any likely errors the end user might experience with the DIF3D software. The code coverage analysis of DIF3D was performed with the Code Coverage Tool of the Intel Fortran compiler which requires modifications to the compilation of DIF3D. The detailed coverage tables are given for each submodule of DIF3D separately, and for the submodules which are primarily developed for DIF3D, most of the source files could be at least partially touched. Most of the uncovered parts/files could be easily ignored, because they are either for error message and debugging output or obviously not needed by DIF3D. Out of the entire source codes of DIF3D, only a few uncovered modules deserve further investigation.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Code Coverage Status of the ARC Code DIF3D

The Argonne Reactor Code (ARC) software system supports users in their fast reactor design goals by providing neutronic, thermal-hydraulic, and structural analysis capabilities. DIF3D plays a pivotal role in the ARC system as the primary homogenized assembly neutronic calculation methodology for fast reactor problems. Over its 40 years history, ARC software usage with DIF3D has been applied to numerous fast and thermal spectrum reactor analysis projects with good to excellent comparison against experiments. With continued improvement of computation resources, many of the geometry modeling capabilities in DIF3D that were primarily used in low order schemes are not really needed anymore. Today, the diffusion and transport capabilities of DIF3D-VARIANT are primarily used in the reactor design process with some scattered usage of DIF3D-FD and DIF3D-Nodal. In recent work, the DIF3D software verification was completed for DIF3D-FD and DIF3D-VARIANT on the geometry options used in the Versatile Test Reactor project. While we can be confident that these capabilities of DIF3D are well used and thus trusted, it does not demonstrate that all possible input options of DIF3D are actually working, but just those that were tested as part of VTR are and that they are correct. Thus, the purpose of the present work is to identify a set of test problems for DIF3D and assess the code coverage of DIF3D for those test problems. The goal is to document what parts of the existing DIF3D code are touched by the set of test problems and which are not. Because the verification work done on DIF3D-VARIANT and DIF3D-FD was focused on the most common uses of DIF3D for fast reactor analysis, the code coverage assessment of those capabilities is the highest priority. This will ensure that nothing is being missed by the existing verification test problems that DIF3D relies upon. The DIF3D-Nodal capability will also be inspected for code coverage as part of this work to further ensure that regular regression testing of DIF3D will trap any likely errors the end user might experience with the DIF3D software. The code coverage analysis of DIF3D was performed with the Code Coverage Tool of the Intel Fortran compiler which requires modifications to the compilation of DIF3D. The detailed coverage tables are given for each submodule of DIF3D separately, and for the submodules which are primarily developed for DIF3D, most of the source files could be at least partially touched. Most of the uncovered parts/files could be easily ignored, because they are either for error message and debugging output or obviously not needed by DIF3D. Out of the entire source codes of DIF3D, only a few uncovered modules deserve further investigation.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Enhancing Solar Power Forecasting with Regularized Constrained Quantile Regression Averaging and Bootstrapping Techniques

Probabilistic solar power forecasting (SPF) plays an essential role in optimizing power-grid operations by quantifying the forecast uncertainty. To improve the accuracy and robustness of probabilistic SPF, this paper introduces the regularized constrained quantile regression averaging (rCQRA) method to combine outputs from multiple PSPF models. In addition, a bootstrapping method was used to quantify model uncertainty, providing insights into the reliability and significance of each ensemble component. To evaluate its efficacy, the proposed rCQRA method is used to integrate four PSPF methods. The resulting SPF models are trained and validated using a real-world six-year dataset from a rooftop solar plant in the USA. The performance of the proposed rCQRA method is evaluated and compared with two benchmark methods under three categories of weather conditions. It is shown that the rCQRA method has superior performance in its forecast reliability, sharpness, and accuracy.

Ensemble learning, probabilistic solar power forec↗

Next-Level Energy Management in Manufacturing: Facility-Level Energy Digital Twin Framework Based on Machine Learning and Automated Data Collection

This research introduces an energy prediction framework at the facility level supported by automated data collection and machine learning models. It investigates whether reducing the prediction time scale allows for applying more complex machine learning techniques and if those techniques improve the prediction accuracy. The primary advantages of this framework lie in its automation of the energy prediction process and its provision of real-time energy data suitable for use in energy dashboards or digital twins. A sitewide dataset was created by combining 15 min energy and daily production data of five shops—assembly, battery, body (electric), body (gas), and paint—from a globally recognized electric vehicle manufacturer. Various machine learning models were evaluated on daily, weekly, and monthly datasets, including, in increasingly complex order: naïve, simple linear regression, net regularized generalized linear regression, principal component regression, k-nearest neighbor, random forest, and Bayesian regularized neural network. Compared to the current state-of-the-art energy consumption prediction for the industrial facility level, this research investigates more complex models and smaller time intervals for higher accuracy. The findings revealed that the more complex monthly models require a minimum of a year and a half of data to operate, while weekly models demand a year of data to achieve improved accuracy. Daily models can operate with only six months of data but exhibit poor performance due to reduced prediction accuracy of production. Key challenges identified include access to reliable, high-quality energy and production data and the initial demand for human labor.

digital twin↗

Inference for Nonparanormal Partial Correlation via Regularized Rank-Based Nodewise Regression

Abstract Partial correlation is a common tool in studying conditional dependence for Gaussian distributed data. However, partial correlation being zero may not be equivalent to conditional independence under non-Gaussian distributions. In this paper, we propose a statistical inference procedure for partial correlations under the high-dimensional nonparanormal (NPN) model where the observed data are normally distributed after certain monotone transformations. The NPN partial correlation is the partial correlation of the normal transformed data under the NPN model, which is a more general measure of conditional dependence. We estimate the NPN partial correlations by regularized nodewise regression based on the empirical ranks of the original data. A multiple testing procedure is proposed to identify the nonzero NPN partial correlations. The proposed method can be carried out by a simple coordinate descent algorithm for lasso optimization. It is easy-to-implement and computationally more efficient compared to the existing methods for estimating NPN graphical models. Theoretical results are developed to show the asymptotic normality of the proposed estimator and to justify the proposed multiple testing procedure. Numerical simulations and a case study on brain imaging data demonstrate the utility of the proposed procedure and evaluate its performance compared to the existing methods. Data used in preparation of this article were obtained from the Alzheimer's Disease Neuroimaging Initiative (ADNI) database.

97 MATHEMATICS AND COMPUTING↗

Improved Bayesian regularization of inverse problems in vibrations and acoustics using noise-only measurements

Here, this paper studies Tikhonov regularization (ridge regression) parameter selection for problems in vibrations and acoustics. The selection method is based on a popular Bayesian method, but it incorporates measurements of sensor noise. The regularization parameter is closely related to the ratio of system input energy to noise energy, so noise measurements inform the inference procedure and improve parameter identification. In cases where standard Bayesian regularization identifies zero as the optimal regularization parameter, noise measurements guarantee a unique nonzero optimum. Sufficient theoretical criteria are developed for this guarantee. The method is verified in even-determined and under-determined configurations in an acoustic source localization simulation and a vibration load identification experiment. It is shown to yield significant improvements over existing empirical Bayesian regularization. Improvements are larger in the even-determined case and smaller in the under-determined case, wherein the inverse solution is less sensitive to the regularization parameter.

42 ENGINEERING↗

Convolutional Neural Networks Trained on Internal Variability Predict Forced Response of TOA Radiation by Learning the Pattern Effect

Abstract Predicting forced, long‐term radiative feedbacks from internal climate variability has been a decades‐long quest in climate science. We train a convolutional neural network (CNN) to predict annual‐ and global‐mean top of the atmosphere radiation anomalies from time‐varying maps of near‐surface temperature in climate models. Trained on internal variability alone, the nonlinear CNN can predict radiation under strong climate change, outperforms a regularized linear regression approach, and works within and across different climate models. We show with explainable artificial intelligence methods that the CNN draws predictive skill from physically meaningful regions but at much smaller spatial scales than currently assumed.

Rugenstein, Maria [Colorado State University Fort ↗

Ultra-fast interpretable machine-learning potentials

Abstract All-atom dynamics simulations are an indispensable quantitative tool in physics, chemistry, and materials science, but large systems and long simulation times remain challenging due to the trade-off between computational efficiency and predictive accuracy. To address this challenge, we combine effective two- and three-body potentials in a cubic B-spline basis with regularized linear regression to obtain machine-learning potentials that are physically interpretable, sufficiently accurate for applications, as fast as the fastest traditional empirical potentials, and two to four orders of magnitude faster than state-of-the-art machine-learning potentials. For data from empirical potentials, we demonstrate the exact retrieval of the potential. For data from density functional theory, the predicted energies, forces, and derived properties, including phonon spectra, elastic constants, and melting points, closely match those of the reference method. The introduced potentials might contribute towards accurate all-atom dynamics simulations of large atomistic systems over long-time scales.

36 MATERIALS SCIENCE↗

Applying Infrared Thermography as a Method for Online Monitoring of Turbine Blade Coolant Flow

As gas turbine engine manufacturers strive to implement condition-based operation and maintenance, there is a need for blade monitoring strategies capable of early fault detection and root-cause determination. Given the importance of blade cooling flows to turbine blade health and longevity, there is a distinct lack of methodologies for coolant flowrate monitoring. The present study addresses this identified opportunity by applying an infrared thermography system on an engine-representative research turbine to generate data-driven models for prediction of blade coolant flowrate. Thermal images were used as inputs to a linear regression and regularization algorithm to relate blade surface temperature distribution with blade coolant flowrate. Additionally, this study investigates how coolant flowrate prediction accuracy is influenced by the number and breadth of diagnostic measurements. Here, the results of this study indicate that a source of high-fidelity training data can be used to predict blade coolant flowrate within about six percent error. Furthermore, identification of prioritized sensor placement supports application of this technique across multiple sensor technologies capable of measuring blade surface temperature in operating gas turbine engines, including spatially resolved and point-based measurement techniques.

42 ENGINEERING↗