Search NASASearch

SEARCH · Search NASA

Results for “principal component regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Next-Level Energy Management in Manufacturing: Facility-Level Energy Digital Twin Framework Based on Machine Learning and Automated Data Collection

This research introduces an energy prediction framework at the facility level supported by automated data collection and machine learning models. It investigates whether reducing the prediction time scale allows for applying more complex machine learning techniques and if those techniques improve the prediction accuracy. The primary advantages of this framework lie in its automation of the energy prediction process and its provision of real-time energy data suitable for use in energy dashboards or digital twins. A sitewide dataset was created by combining 15 min energy and daily production data of five shops—assembly, battery, body (electric), body (gas), and paint—from a globally recognized electric vehicle manufacturer. Various machine learning models were evaluated on daily, weekly, and monthly datasets, including, in increasingly complex order: naïve, simple linear regression, net regularized generalized linear regression, principal component regression, k-nearest neighbor, random forest, and Bayesian regularized neural network. Compared to the current state-of-the-art energy consumption prediction for the industrial facility level, this research investigates more complex models and smaller time intervals for higher accuracy. The findings revealed that the more complex monthly models require a minimum of a year and a half of data to operate, while weekly models demand a year of data to achieve improved accuracy. Daily models can operate with only six months of data but exhibit poor performance due to reduced prediction accuracy of production. Key challenges identified include access to reliable, high-quality energy and production data and the initial demand for human labor.

digital twin

Quantifying market volume sensitivity to material property modifications in polyhydroxybutyrate: A parametric analysis approach

Polyhydroxybutyrate (PHB), a biodegradable biopolymer, represents a promising alternative to petroleum-based thermoplastics. However, despite consistent market growth, PHB faces persistent commercialization challenges that limit widespread adoption. Existing research has focused predominantly on optimizing PHB production processes, leaving a critical gap in understanding which material property modifications would most effectively enhance market competitiveness. This study addresses this gap by systematically analyzing the relationship between polymer material properties and market performance using U.S. market data from 2008 to 2021 for 21 thermoplastic polymers across 19 material properties. We employed principal component regression to identify property modifications that could maximize market volume while reducing CO 2 emissions. Our parametric analysis revealed that two specific material properties – Hardness Shore A and Sheet Extrusion Temperature – significantly influence PHB marketability across different price points. Market simulations demonstrated that a 10% increase in Hardness Shore A could increase PHB market volume by 431.5 million kg while reducing emissions by 188.7 kg CO 2 . A similar 10% increase to Sheet Extrusion Temperature could yield a 297.5 million kg volume increase and a 99.2 kg CO 2 reduction in emissions. Critically, this approach is agnostic to the specific methods required to achieve these property changes, instead providing material scientists with quantitative, data-driven targets for R&D prioritization. Here, this framework offers a novel methodology for evaluating biopolymer competitiveness and supporting strategic decisions to accelerate PHB market adoption and contribute to decarbonization of the plastics industry.

09 BIOMASS FUELS

Silver diamine fluoride differentially affects dentin and hypomineralized enamel permeabilities

OBJECTIVES: To investigate the physicochemical effect of silver diamine fluoride (SDF) by correlating permeability with mineral density and elemental composition of hypomineralized enamel and carious dentin. METHODS: Enamel and dentin from human carious primary teeth with and without SDF treatment in-vivo, and hypomineralized enamel from permanent molars with and without SDF treatment in-vitro were scanned using micro X-ray computed tomography. Spatial maps of biometals (calcium, zinc), phosphorus, and silver were generated using X-ray fluorescence microprobe. Permeabilities were computed using Porous Microstructure Analysis software. RESULTS: The intrinsic permeability of SDF-treated carious dentin was 14.3 % lower than untreated sound dentin (6.39e-15 ± 3.01e-15 m² vs 7.46e-15 ± 1.82e-15 m²; P < 0.0001), while untreated carious dentin was 98.4 % higher (1.48e-14 ± 7.11e-15 m²; P < 0.0001). SDF-treated and untreated transparent dentin showed similar reduced permeabilities (75.6 % and 78.4 % lower than untreated sound dentin, respectively; P = 0.93). Severely hypomineralized enamel showed permeability reaching 108.1 % of adjacent sound dentin (5.71e-15 ± 2.04e-15 m² vs 5.28e-15 ± 1.30e-15 m²; P = 0.1409) and was significantly higher than mildly hypomineralized enamel (1.39e-15 ± 1.04e-15 m²; P < 0.0001). SDF treatment did not significantly impact the permeability of severely hypomineralized enamel (12.4 % reduction; P = 0.07). Principal component regression identified Zn level as a significant effector of tissue permeabilities in carious primary teeth (P < 0.0001). SIGNIFICANCE: This study introduces a computational method to measure dental tissue permeability, and demonstrates that SDF significantly reduces permeability in carious dentin but not intact hypomineralized enamel. The study reveals biometal Zn localization can alter dentin and enamel permeabilities, providing new insights into pathobiological mechanisms underlying caries and hypomineralization.

Chou, Conrad

GP-BayesOpInf

SAND2025-01851O GP-BayesOpInf is a software tool that uses algorithms to combine Gaussian process regression, principal component analysis, and linear Bayesian inference to produce a probabilistic reduced-order model for time-dependent systems. Numerical examples include the compressible Euler equations for an ideal gas, a heat diffusion process with a nonlinear reaction term, and a set of ordinary differential equations describing a compartmental model in epidemiology. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

SciDAC

Transient uncertainty quantification and Global Sensitivity Analysis of the open-source Molten Chloride Reactor Experiment (MCRE) using GP-PCA surrogate models

Uncertainties in the thermophysical properties of molten salts impact both the steady-state and transient behavior of Molten Salt Reactors (MSRs). In this work, we aim to quantify the influence of such uncertainties on the transient operation of the Molten Chloride Reactor Experiment (MCRE), utilizing the open-source specifications provided for this reactor. Seven representative transient scenarios are considered. For each scenario, we evaluate the impact of thermophysical property uncertainties on four key multiphysics model output variables of interest (VoIs): maximum power density, maximum fuel temperature, maximum reflector temperature, and average fuel velocity magnitude. In addition, we perform a Global Sensitivity Analysis (GSA) by computing Sobol’ indices for the uncertain input parameters to determine their contribution to the variability of each VoI. Conducting GSA is computationally intensive due to the large number of required evaluations of the high-fidelity multiphysics model. To mitigate this cost, we develop a surrogate modeling framework that combines Gaussian Process (GP) regression with Principal Component Analysis (PCA), enabling efficient sample generation for the GSA. Our results show that for energy-related VoIs, thermal conductivity is the dominant contributor to uncertainty. In contrast, for flow-related VoIs, density and dynamic viscosity are the primary sources of uncertainty. The specific heat of the fuel salt was found to play a secondary role in the transient analyses.

42 - ENGINEERING

Multi-Parameter Optical Fiber for Distributed Sensing of Humidity, CH4, CO2, and Corrosion

This work describes the use of the optical fiber sensor (OFS) for successful monitoring of humidity, CH4, and CO2 based on the strain produced along the single-mode fiber (SMF) sensor. This is enabled by absorption of H2O/gases onto the commercially available polyacrylate coated jacketed portion of the fiber resulting in a change in strain. Under equilibrium, a differential microstrain was observed along the jacketed portion of the SMF with N2, CH4, and CO2 at different relative humidity (RH) conditions. The strain response of the SMF under various mixed gas composition and different RH conditions were also measured and made calibration curves accordingly. Linear regression and principal component analysis of the strain datasets provided deconvolution of the impact of strain from H2O, N2, CH4, and CO2. Additionally, modified OFS comprised of the Fe coated fiber section was employed to monitor corrosion based on the increase in backscattered intensity amplitude of the light being passed once corrosion of Fe occurs. Also, corrosion of Fe was studied under soil by installing the fiber in soil with protective measures which would prevent the sensor being mechanically disturbed or broken during installation. The corrosion rates were studied by monitoring the rate at which the intensity of backscattered light amplitude attains a steady state value when complete corrosion of Fe with a specific coating thickness occurs.

Mainali, Badri

Uncertainty Quantification for Smooth Functional Data with Application to Material Properties

This document outlines a method for processing functional output (i.e., curves) for the ultimate purpose of sampling curves under specified input conditions for use in modeling and simulation uncertainty quantification (UQ) studies. A set of benchmark curves sufficiently representative of the relevant scenario(s) being simulated are provided to the process and formatted as described in Section 1. Principal Component Analysis (PCA) is utilized to discover the components of uncertainty in the benchmark curves and is outlined in Section 2. Section 3 describes the application of uncertainty quantification to the PCA results for the purpose of sampling curves to be used in UQ analysis. Section 4 applies these techniques to an example benchmark dataset. Concluding remarks are provided in the final section.

36 MATERIALS SCIENCE

Probabilistic projections of the Amery Ice Shelf catchment, Antarctica, under conditions of high ice-shelf basal melt

Abstract. Antarctica's Lambert Glacier drains about one-sixth of the ice from the East Antarctic Ice Sheet and is considered stable due to the strong buttressing provided by the Amery Ice Shelf. While previous projections of the sea-level contribution from this sector of the ice sheet have predicted significant mass loss only with near-complete removal of the ice shelf, the ocean warming necessary for this was deemed unlikely. Recent climate projections through 2300 indicate that sufficient ocean warming is a distinct possibility after 2100. This work explores the impact of parametric uncertainty on projections of the response of the Lambert–Amery system (hereafter “the Amery sector”) to abrupt ocean warming through Bayesian calibration of a perturbed-parameter ice-sheet model ensemble. We address the computational cost of uncertainty quantification for ice-sheet model projections via statistical emulation, which employs surrogate models for fast and inexpensive parameter space exploration while retaining critical features of the high-fidelity simulations. To this end, we build Gaussian process (GP) emulators from simulations of the Amery sector at a medium resolution (4–20 km mesh) using the Model for Prediction Across Scales (MPAS)-Albany Land Ice (MALI) model. We consider six input parameters that control basal friction, ice stiffness, calving, and ice-shelf basal melting. From these, we generate 200 perturbed input parameter initializations using space filling Sobol sampling. For our end-to-end probabilistic modeling workflow, we first train emulators on the simulation ensemble and then calibrate the input parameters using observations of the mass balance, grounding line movement, and calving front movement with priors assigned via expert knowledge. Next, we use MALI to project a subset of simulations to 2300 using ocean and atmosphere forcings from a climate model for both low- and high-greenhouse-gas-emission scenarios. From these simulation outputs, we build multivariate emulators by combining GP regression with principal component dimension reduction to emulate multivariate sea-level contribution time series data from the MALI simulations. We then use these emulators to propagate uncertainty from model input parameters to predictions of glacier mass loss through 2300, demonstrating that the calibrated posterior distributions have both greater mass loss and reduced variance compared to the uncalibrated prior distributions. Parametric uncertainty is large enough through about 2130 that the two projections under different emission scenarios are indistinguishable from one another. However, after rapid ocean warming in the first half of the 22nd century, the projections become statistically distinct within decades. Overall, this study demonstrates an efficient Bayesian calibration and uncertainty propagation workflow for ice-sheet model projections and identifies the potential for large sea-level rise contributions from the Amery sector of the Antarctic Ice Sheet after 2100 under high-greenhouse-gas-emission scenarios.

54 ENVIRONMENTAL SCIENCES

Exploring biofiber properties and their influence on biocomposite tensile properties

Biofibers serve as effective reinforcements for neat polylactic acid (PLA) in biocomposites, offering an attractive opportunity to decarbonize the manufacturing sector of the United States by displacing fossil-based reinforcement fibers such as carbon fibers. Also, biofiber production can stimulate economic growth in rural economies, fueling sustainable development. PLA resins are commonly compounded with biofibers to create biocomposites suitable for additive manufacturing. PLA-biofiber composites often exhibit better overall material properties than neat (pure) PLA, but the associations between biofiber properties and the material properties of their biocomposites remain largely unexplored. Hence, this research delves into a comprehensive exploration of diverse biofibers, scrutinizing their physical and chemical attributes, including size, shape, ash content and biochemical composition. The study meticulously analyzes the flow properties of each biofiber and elucidates the ultimate tensile strengths and Young's modulus of corresponding biocomposite samples. Noteworthy correlations between biofiber and biocomposite tensile properties are uncovered, shedding light on critical interrelationships. The study introduces an approach employing regression models to predict the ultimate tensile strength and Young's modulus of biocomposites. These models, validated with a cross-validation technique, exhibit remarkable predictive accuracy, particularly in estimating ultimate tensile strength. © 2024 Oak Ridge National Laboratory managed by UT-Battelle, LLC and The Author(s). Polymer International published by John Wiley & Sons Ltd on behalf of Society of Chemical Industry.

36 MATERIALS SCIENCE

Unraveling Fundamental Activity–Stability Relationships in Rutile Oxides

The oxygen evolution reaction (OER) is a key anodic half-cell reaction that accompanies several critical electrochemical reduction reactions of interest to a variety of applications. Despite steady advances in understanding and qualitatively predicting OER activity and selectivity trends, a comprehensive description or prediction of material aqueous (in)stability and degradation mechanisms remains elusive, even though these processes critically influence device lifetime and economic feasibility. In this work, we investigate the interplay, or lack thereof, between OER activity and material aqueous stability across rutile oxides, with a particular focus on iridium oxide (IrO 2 ). By applying a Born–Haber cycle, we calculate the thermodynamic driving force for metal dissolution as a function of the applied bias and electrolyte conditions. We apply interpretable machine learning techniques, including principal component analysis and symbolic regression, to analyze trends across rutile oxides and find that key thermodynamic descriptors for OER activity and surface stability are only very weakly correlated. Instead, the local atomic environment─especially electronic structure signatures for interactions between the active site and its neighbors─plays a more important role in predicting material stability. Leveraging these insights, we investigate the impact of doping IrO 2 with a range of transition metals and show that the stability of Ir active sites can be tuned largely independently of its predicted OER activity. These insights lay the foundation for material design to improve stability with respect to corrosion, with the ultimate aim to enhance long-term stability without sacrificing catalytic performance in the OER.

evolution reactions

Projection-based multifidelity linear regression for data-scarce applications

Surrogate modeling for systems with high-dimensional quantities of interest remains challenging, particularly when training data are costly to acquire. This work develops multifidelity methods for multiple-input multiple-output linear regression targeting data-limited applications with high-dimensional outputs. Multifidelity methods integrate many inexpensive low-fidelity model evaluations with limited, costly high-fidelity evaluations. We introduce two projection-based multifidelity linear regression approaches with linear and nonlinear features that leverage principal component basis vectors for dimensionality reduction and combine multifidelity data through: (i) a direct data augmentation using low-fidelity data, and (ii) a data augmentation incorporating explicit linear corrections between low-fidelity and high-fidelity data. The data augmentation approaches combine high-fidelity and low-fidelity data into a unified training set and train the linear regression model through weighted least squares with fidelity-specific weights. We introduce a proximity-based weighting scheme with automatic weight selection strategy through cross-validation. Here, the proposed multifidelity linear regression methods are demonstrated on approximating the surface pressure field of a hypersonic vehicle in flight and the temperature field on an aircraft disc braking system. In an ultra low-data regime of no more than twelve high-fidelity samples, multifidelity linear regression achieves approximately 2% – 12% improvement in median accuracy and a higher R 2 score relative to single-fidelity methods at comparable computational cost.

data augmentation

Machine Learning-Driven Reliability Estimation of PV Inverters Considering Alert-Ambient Variability

Weather-induced spatio-temporal degradation limits outdoor PV inverter lifetime and reliability, necessitating advanced data analysis. This study employs a top-down, data-driven approach utilizing multiple machine learning (ML) algorithms to estimate inverter reliability in a 1.4 MW PV power plant, considering factors such as irradiance, humidity, temperature, time of day, and weather conditions. An extensive alert dataset from 17 identical inverters, including alert types, propagation, and frequency, reveals significant correlations with environmental factors and inverter output power, enabling the construction of a performance reliability model. Dual-stage supervised-ML models are evaluated for accuracy, with the ‘classification-regression’ model by an artificial neural network (ANN) tested on the averaged “Alert-Ambient” dataset, which is outperformed by ‘clustering-regression’ models using random forest (RF) and K-Nearest Neighbors (KNN) on individual inverter datasets. K-means clustering applies principal component analysis to reduce dimensions, achieving improved accuracy beyond the 80% achieved by ANN on the averaged dataset. Second-stage regression estimates inverter reliability with a mean square error of 0.0195 on the averaged dataset and as low as 0.002 on individual inverter datasets using RF. Furthermore, these findings highlight the method's suitability for estimating PV inverter output reliability under ambient conditions, essential for digital twin development and related applications.

14 SOLAR ENERGY

A Parametric, Data-Driven, Non-Intrusive Reduced-Order Model Framework for Crystal Plasticity Simulations of Voids

The influence of the internal structure at micrometer length scales on the deformation of polycrystalline materials can be effectively captured using crystal plasticity finite element methods (CPFEM). However, the complexity and nonlinearity of the deformation equations CPFEM solves demand significant computational power and resources to achieve accurate predictions, limiting its broader application. To address this challenge, we have identified a reduced-order representation of the complex data in order to establish a computationally efficient reduced-order models (ROM) and drastically reduce the computational expense of CPFEM. Specifically, in this work, we developed a parametric, data-driven, and non-intrusive ROM framework for CPFEM using proper orthogonal decomposition (POD) and sparse variational Gaussian process (SVGP) regression for single-crystal microstructures under tensile loading conditions. The developed protocol enables one to compress field into a latent/low-dimensional space described by principal component analysis (PCA) via the singular value decomposition (SVD) algorithm. As a result, the high-dimensional data are reduced to a significantly smaller amount of dimensions with POD bases and POD coefficients. Furthermore, we deployed an ensemble of SVGPs—extended from the classical Gaussian process (GP) regression for scalability and handling big data—in a massively parallel manner to train and predict latent POD coefficients using known POD bases from a set of previously obtained simulations results. Lastly, using the predicted POD coefficients, we reconstructed the full-field results and showed reasonable agreement compared with the true values obtained from running CPFEM. The developed framework is validated with a set of CPFEM simulations of a single embedded void in single-crystal aluminum alloy. While the framework is broadly applicable, this work specifically focuses on single-crystal microstructures, a single load case (e.g., tensile), and a specific void geometry (spherical).

Anisotropy

Machine Learning–Augmented Laser-Induced Breakdown Spectroscopy for Spectral Discrimination of Iron Oxalates

Enhanced characterization and phase identification of post-PUREX Pu Oxalates (PuOXA) are pivotal for nonproliferation and pre-detonation nuclear forensics. Despite significant advances in the characterization of PuO 2 samples, little is known about the impact of both the chemical structure and oxidation states of PuOXA (i.e., Pu(III) and Pu(IV)) have on optical emission signatures. Here, we demonstrate the analytical capabilities of laser-induced breakdown spectroscopy (LIBS) applied to Fe(II) and Fe(III) oxalate samples as surrogates for PuOXA, highlighting the discriminating features in the LIBS emission spectra arising from differences in the oxidation states within mixed FeOXA samples. We report the enhancement of spectral feature selection using Principal Component Analysis (PCA), which enables the analytical superiority of machine learning algorithms such as Linear Discriminant Analysis (LDA), Quadratic Discriminant Analysis (QDA), Partial Least Squares Regression (PLSR), Support Vector Regression (SVR), and Random Forest Regression (RFR) over conventional univariate techniques for phase discrimination and chemometric analysis. Cluster analysis revealed how both matrix effects and laser ablation influence cluster separability by introducing spectral artifacts that misdirect the maximization of variance. PCA-selected emission lines were used in the regression models, demonstrating that both univariate and multivariate linear regression models (i.e., PLSR and SVR) can achieve acceptable performance, with machine learning models outperforming conventional calibration regressions. Furthermore, the application of non-linearly activated PCA-selected emission lines illustrates how simplifying the data while retaining captured variance enables the use of less complex and more computationally efficient models. Furthermore, this is particularly evident in the underperformance of RFR, which suffers from increased computational costs and overfitting owing to its high complexity.

Oxalates

Machine learning inversion from small-angle scattering for charged polymers

We develop Monte Carlo simulations for uniformly charged polymers and a machine learning algorithm to interpret the intra-polymer structure factor of the charged polymer system, which can be obtained from small-angle scattering experiments. The polymer is modeled as a chain of fixed-length bonds, where the connected bonds are subject to bending energy, and there is also a screened Coulomb potential for charge interaction between all joints. The bending energy is determined by the intrinsic bending stiffness, and the charge interaction depends on the interaction strength and screening length. All three contribute to the stiffness of the polymer chain and lead to longer and larger polymer conformations. The screening length also introduces a second length scale for the polymer besides the bending persistence length. To obtain the inverse mapping from the structure factor to these polymer conformation and energy-related parameters, we generate a large data set of structure factors by running simulations for a wide range of polymer energy parameters. We use principal component analysis to investigate the intra-polymer structure factors and determine the feasibility of the inversion using the nearest neighbor distance. We employ Gaussian process regression to achieve the inverse mapping and extract the characteristic parameters of polymers from the structure factor with low relative error.

36 MATERIALS SCIENCE

Remote sensing of Pu in uranyl nitrate crystals using reflectance spectroscopy and chemometrics

Remote quantification of Pu(VI) (0–5 mol%) co-crystallized with U in uranyl nitrate hexahydrate (UNH) crystals was achieved in a glove box using reflectance spectroscopy coupled with chemometric modeling. Reflectance spectra were also acquired for Pu(IV) and Np(VI) (0–5 mol%) crystallized with UNH; revealing spectral features consistent with their solution-phase analogs. Principal component analysis revealed Pu(IV/VI) and Np(VI) concentrations as the primary source of variation in the data, informing the development of a supervised partial least squares regression model for Pu(VI). The resulting calibration demonstrated robust performance, with replicate root mean square errors near 10% and quantifiable limits near 0.2 mol% Pu(VI) relative to U. The Pu(VI) remained stable in the crystalline UNH matrix for at least one week with minimal reduction to Pu(IV). Notably, Pu(VI) and Np(VI) incorporation in UNH quenched U(VI) fluorescence while Pu(IV) did not. This study presents a noninvasive, spectroscopic approach for solid-state Pu quantification, with direct implications for material accountability and nuclear nonproliferation monitoring.

Sadergaski, Luke R. [Oak Ridge National Laboratory

Quantifying Temperature Dependence of Pu(IV) Absorbance Spectra for Advanced Online Monitoring of Nuclear Processes

This article presents a systematic study of Pu(IV) absorbance spectral features as a function of temperature to develop an understanding of this parameter’s effect on chemometric models that can be used as online monitoring tools to support nuclear processing. The descriptive and predictive models that provide real-time feedback of these processes are usually constructed with data collected in conditions typical of a laboratory environment, which can differ drastically from a processing environment. To assess the impact of temperature on Pu(IV) absorbance spectra, 11 samples of Pu(IV) were synthesized with varying HNO 3 concentrations ranging from 0.6 to 9.5 M and heated between 15 and 45 °C. Ultraviolet (UV)–visible (vis)–near-infrared (NIR) absorption spectra collected at different HNO 3 concentrations and temperatures revealed that features associated with Pu(IV) are sensitive to temperature at all HNO 3 concentrations and that changes in features depend on HNO 3 concentration. The contributions of temperature and HNO 3 concentration to variation in Pu(IV) spectral features were evaluated using the principal component analysis of spectra that were baseline-corrected with an asymmetric least-squares method. Furthermore, predictive modeling for HNO 3 concentration with partial least-squares regression of UV–vis–NIR spectra highlighted the importance of accounting for temperature in the calibration set to optimize model performance. This methodology constitutes a new, systematic approach to account for the effect of temperature on the absorption spectra of metal ions and is useful for process monitoring applications in many industries.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH