Search NASA⌕ Search

SEARCH · Search NASA

Results for “Regression Analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Extracting the Breakout Distance from the ECOT Trajectories: Gaussian Process Regression Approach

Enhanced Corner Turning (ECOT) experiments provide an important metric of performance of high explosive (HE) formulations. The breakout distance is a single scalar value that characterizes the corner turning efficiency of an HE. Extracting the breakout distance from the raw ECOT results, whether experimental or simulated, is a conceptually straightforward procedure which, however, is non-unique, especially in the presence of noise. More specifically, this procedure involves numerical smoothing and selecting particular values for parameters of this smoothing introduces human bias. In this work, we propose to use the Gaussian process regression to analyze ECOT results. This analysis involves the effective smoothing of the data, thus allowing for accurate extraction of the breakout distance. Most importantly, the parameters of this smoothing can be inferred from the ECOT data itself, rendering the approach effectively parameter-free and thus diminishing the human bias. An additional benefit of the Gaussian process regression, being a statistical inference method, is that not just the value of the breakout distance, but also its confidence interval can be extracted from the data. This report introduces the Gaussian process regression, as applied to ECOT, and demonstrates its usefulness by extracting the breakout distances for a selection of experimental and simulated data.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Jensen–Shannon divergence based novel loss functions for Bayesian neural networks

Bayesian neural networks (BNNs) are state-of-the-art machine learning methods that can naturally regularize and systematically quantify uncertainties using their stochastic parameters. Kullback–Leibler (KL) divergence-based variational inference used in BNNs suffer from unstable optimization and challenges in approximating light-tailed posteriors due to the unbounded nature of the KL divergence. To resolve these issues, we formulate a novel loss function for BNNs based on a new modification to the generalized Jensen–Shannon (JS) divergence, which is bounded. In addition, we propose a Geometric JS divergence-based loss, which is computationally efficient since it can be evaluated analytically. We found that the JS divergence-based variational inference is intractable, and hence employed a constrained optimization framework to formulate these losses. Our theoretical analysis and empirical experiments on multiple regression and classification data sets suggest that the proposed losses perform better than the KL divergence-based loss, especially when the data sets are noisy or biased. Specifically, there are approximately 5% and 8% improvements in accuracy for a noise-added CIFAR-10 dataset and a regression dataset, respectively. There is about 13% reduction in false negative predictions of a biased histopathology dataset. Additionally, we quantify and compare the uncertainty metrics for the regression and classification tasks.

97 MATHEMATICS AND COMPUTING↗

Machine learning-assisted profiling of a kinked ladder polymer structure using scattering

Ladder polymers consisting of fused rings in the backbone have very limited conformational freedom, which results in very different properties from traditional linear polymers. However, accurately determining their size and chain conformations from solution scattering remains a challenge. Their chain conformations of kinked ladder polymers are largely governed by the structures and relative orientations or configurations of the repeat units, unlike conventional polymer chains whose bending angles between repeat units follow a unimodal Gaussian distribution. Meanwhile, traditional scattering models for polymer chains do not account for these unique structural features. This work introduces a novel approach that integrates machine learning with Monte Carlo simulations to construct a model that can describe the geometry of a type of kinked CANAL ladder polymers. We first develop a Monte Carlo simulation model for sampling the configuration space of CANAL ladder polymers, where each repeat unit is modeled as a biaxial segment. Then, we establish a machine learning-assisted scattering analysis framework based on Gaussian Process Regression. Finally, we conduct small-angle neutron scattering experiments on a CANAL ladder polymer solution to apply our approach. Our method uncovers structural features of such ladder polymers that conventional methods fail to capture.

Ding, Lijie [Oak Ridge National Laboratory (ORNL),↗

Tracking the Sun: Pricing and Design Trends for Distributed Photovoltaic Systems in the United States, 2024 Edition [Slides]

Berkeley Lab’s annual Tracking the Sun report describes trends among grid-connected, distributed solar photovoltaic (PV) and paired PV+storage systems in the United States. For the purpose of this report, distributed solar includes residential systems, roof-mounted non-residential systems, and ground-mounted systems up to 5 MW-AC. Ground-mounted systems larger than 5 MW-AC are covered in Berkeley Lab’s companion annual report, Utility-Scale Solar. The latest edition of the report is based on 3.7 million systems installed through year-end 2023, representing close to 80% of systems installed to date. The report describes and discusses key trends related to: -Project characteristics, including system size, module efficiencies, roof-coverage ratios, prevalence of paired PV with storage, use of module-level power electronics, third-party ownership, mounting configurations, panel orientation, and customer segmentation -Median installed-price trends, both nationally and by state -Variability in pricing according to system size, state, installer, equipment type, and other factors, relying on both descriptive analysis and a multi-variate regression to estimate the effects of key pricing drivers for residential systems installed in 2023.

14 SOLAR ENERGY↗

Uncertainty Quantification for Smooth Functional Data with Application to Material Properties

This document outlines a method for processing functional output (i.e., curves) for the ultimate purpose of sampling curves under specified input conditions for use in modeling and simulation uncertainty quantification (UQ) studies. A set of benchmark curves sufficiently representative of the relevant scenario(s) being simulated are provided to the process and formatted as described in Section 1. Principal Component Analysis (PCA) is utilized to discover the components of uncertainty in the benchmark curves and is outlined in Section 2. Section 3 describes the application of uncertainty quantification to the PCA results for the purpose of sampling curves to be used in UQ analysis. Section 4 applies these techniques to an example benchmark dataset. Concluding remarks are provided in the final section.

36 MATERIALS SCIENCE↗

A novel approach for large-scale wind energy potential assessment

Increasing wind energy generation is central to grid decarbonization, yet methods to estimate wind energy potential are not standardized, leading to inconsistencies and even skewed results. This study aims to improve the fidelity of wind energy potential estimates through an approach that integrates geospatial analysis and machine learning (i.e., Gaussian process regression). We demonstrate this approach to assess the spatial distribution of wind energy capacity potential in the Contiguous United States (CONUS). We find that the capacity-based power density ranges from 1.70 MW/km2 (25th percentile) to 3.88 MW/km2 (75th percentile) for existing wind farms in the CONUS. The value is lower in agricultural areas (2.73 ± 0.02 MW/km2, mean ± 95 % confidence interval) and higher in other land cover types (3.30 ± 0.03 MW/km2). Notably, advancements in turbine manufacturing could reduce power density in areas with lower wind speeds by adopting low specific-power turbines, but improve power density in areas with higher wind speeds (>8.35 m/s at 120m above the ground), highlighting opportunities for repowering existing wind farms. Wind energy potential is shaped by wind resource quality and is regionally characterized by land cover and physical conditions, revealing significant capacity potential in the Great Plains and Upper Texas. The results indicate that areas previously identified as hot spots using existing approaches (e.g., the west of the Rocky Mountains) may have a limited capacity potential due to low wind resource quality. Improvements in methodology and capacity potential estimates in this study could serve as a new basis for future energy systems analysis and planning.

Dai, Tao↗

The cluster decomposition of the configurational energy of multicomponent alloys

Abstract The cluster expansion method (CEM) is a widely used lattice-based technique in the study of multicomponent alloys. Despite its prevalent use, a clear understanding of expansion terms is lacking. We present a modern mathematical formalism of the CEM and introduce thecluster decomposition—a unique and basis-independent decomposition for functions of the atomic configuration in a crystal. We identify the cluster decomposition as an invariant ANOVA decomposition; and demonstrate how functional analysis of variance and sensitivity analysis can be used to interpret interactions among species. Furthermore, we show how the mathematical structure of the cluster decomposition enables numerical evaluation that scales with the number of clusters and is independent of the number of species. Overall, our work enables rigorous interpretations of interactions among species, provides opportunities to explore parameter estimation beyond linear regression, introduces a numerical efficient implementation, and enables analysis of cluster expansions based on established mathematical and statistical principles.

Chemistry↗

Taylor-Expansion-Based Robust Power Flow in Unbalanced Distribution Systems: A Hybrid Data-Aided Method

Traditional power flow methods often adopt certain assumptions designed for passive balanced distribution systems, thus lacking practicality for unbalanced operation. moreover, their computation accuracy and efficiency are heavily subject to unknown errors and bad data in measurements or prediction data of distributed energy resources (ders). to address these issues, this paper proposes a hybrid data-aided robust power flow algorithm in unbalanced distribution systems, which combines taylor series expansion knowledge with a data-driven regression technique. the proposed method initiates a linearization power flow model to derive an explicitly analytical solution by modified taylor expansion. to mitigate the approximation loss that surges due to the der integration and bad data, we further develop a data-aided robust support vector regression approach to estimate the errors efficiently. comparative analysis in the 13-bus and 123-bus ieee unbalanced feeders shows that the proposed hybrid algorithm achieves superior computational efficiency, with guaranteed accuracy and robustness against outliers.

data-driven↗

Bayesian Linear Regression for Hugoniot Data

This repository provides the code and datasets used in the paper Bayesian Analysis of Linear Shock Compression Data. This paper analyzes publicly available shock compression datasets on copper, argon, and nickel from Marsh (1980) using Bayesian linear regression, and compares the results with those obtained using bootstrapping methods. References: - Marsh, S. P. (1980). LASL shock Hugoniot data (Vol. 5). Univ of California Press.

Bernstein, JasonA [Lawrence Livermore National Lab↗

Projection-based multifidelity linear regression for data-scarce applications

Surrogate modeling for systems with high-dimensional quantities of interest remains challenging, particularly when training data are costly to acquire. This work develops multifidelity methods for multiple-input multiple-output linear regression targeting data-limited applications with high-dimensional outputs. Multifidelity methods integrate many inexpensive low-fidelity model evaluations with limited, costly high-fidelity evaluations. We introduce two projection-based multifidelity linear regression approaches with linear and nonlinear features that leverage principal component basis vectors for dimensionality reduction and combine multifidelity data through: (i) a direct data augmentation using low-fidelity data, and (ii) a data augmentation incorporating explicit linear corrections between low-fidelity and high-fidelity data. The data augmentation approaches combine high-fidelity and low-fidelity data into a unified training set and train the linear regression model through weighted least squares with fidelity-specific weights. We introduce a proximity-based weighting scheme with automatic weight selection strategy through cross-validation. Here, the proposed multifidelity linear regression methods are demonstrated on approximating the surface pressure field of a hypersonic vehicle in flight and the temperature field on an aircraft disc braking system. In an ultra low-data regime of no more than twelve high-fidelity samples, multifidelity linear regression achieves approximately 2% – 12% improvement in median accuracy and a higher R 2 score relative to single-fidelity methods at comparable computational cost.

data augmentation↗

Dataset for ASME VVUQ Symposium Workshop on Regression of Validation Data to an Application Point

This dataset consists of a collection of Excel spreadsheets that contain output from analysis specified in the workshop. The analysis involves ASME V&V 20-style validation as well as the application of a supplement methodology for regression of validation comparison error and validation uncertainty to application points where experimental data does not exist for comparison. The simulation results and experimental data are provided by the workshop organizers and a NASA report, respectively.

Kirsch, Jared Roelof [Sandia National Laboratories↗

Code Coverage Status of the ARC Code RCT

The Argonne Reactor Code (ARC) software system supports users in their fast reactor design goals by providing neutronic, thermal-hydraulic, and structural analysis capabilities. REBUS plays a pivotal role in the ARC system as the primary fuel cycle analysis capability for fast reactor problems. Over its 60 year history, ARC software usage with REBUS has been applied to numerous fast and thermal spectrum reactor analysis projects with good to excellent comparison against experiments. The RCT code is a later addition and uses the REBUS restart files to define its input. The RCT code was built to provide pin depletion details on EBR-II models and thus many features of RCT were specifically tailored to the needs of EBR-II models. Additional approximations were invoked which are likely only valid for the EBR-II reactor and the particular fuel management that was done for it. The purpose of the present work is to identify a set of test problems for RCT and assess the code coverage for those test problems. The goal is to document what parts of the existing RCT code are touched by the set of test problems and which are not. Because no detailed verification work has been done on RCT, the existing regression testing suite was chosen for the code coverage assessment. The code coverage analysis of RCT was performed with the Code Coverage Tool of the Intel Fortran compiler which requires modifications to the compilation of RCT. The detailed coverage tables are given for each part of RCT. As will be discussed and shown, some parts of the RCT capability that are known to be used by the EBR-II analysis work are not tested by the regression testing suite. These aspects should be resolved before major source code changes are taken for the RCT software. Because REBUS and DIF3D are not subroutines of RCT, the coverage changes in both of those codes is not altered by RCT. The same is true for all of the modules of DIF3D that are used by RCT such as SYSLIB and SEGLIB.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Interpretable, extensible linear and symbolic regression models for charge density prediction using a hierarchy of many-body correlation descriptors

Here, density functional theory (DFT) is routinely used to make electronic structure predictions for high-throughput screening of materials and molecules for technologically relevant areas, like the identification of better catalysts, electronic materials, and drug discovery. However, the DFT formalism is limited by (a) its poor (quadratic-to-quartic) scaling, and (b) the need to perform repeated eigenvalue computations of the electronic Hamiltonian as part of its self-consistent field (SCF) iteration procedure to obtain the converged ground state electron density, ρ (r). Approaches that directly predict ρ (r) of a structure with high accuracy can accelerate conventional SCF calculations and can also be used in linearly scaling methods such as orbital-free DFT. To this end, we present a procedure to predict the ground state electron density of molecular and periodic three-dimensional systems directly from the atomic structure with a particular emphasis on physical interpretability. In our framework, ρ (r) is modeled using many-body correlation descriptors that accurately capture the effects of local atomic arrangements in the neighborhood of a grid point. Our use of a linear regression scheme to fit to charge density data enables transparent analysis of the relative contributions of various types of local atomic correlations. By systematically including increasingly complex correlations, our model is shown to accurately predict ρ (r) for a variety of chemically and electronically diverse systems — amorphous Ge, Al(001) slab, crystalline Ga 2 O 3 , molecular benzene, and polyethylene. We then demonstrate a symbolic regression-based protocol to construct easily computable, interpretable features from lower-order correlations that significantly improves our electron density predictions with effectively no increase in the computational cost.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Quantifying market volume sensitivity to material property modifications in polyhydroxybutyrate: A parametric analysis approach

Polyhydroxybutyrate (PHB), a biodegradable biopolymer, represents a promising alternative to petroleum-based thermoplastics. However, despite consistent market growth, PHB faces persistent commercialization challenges that limit widespread adoption. Existing research has focused predominantly on optimizing PHB production processes, leaving a critical gap in understanding which material property modifications would most effectively enhance market competitiveness. This study addresses this gap by systematically analyzing the relationship between polymer material properties and market performance using U.S. market data from 2008 to 2021 for 21 thermoplastic polymers across 19 material properties. We employed principal component regression to identify property modifications that could maximize market volume while reducing CO 2 emissions. Our parametric analysis revealed that two specific material properties – Hardness Shore A and Sheet Extrusion Temperature – significantly influence PHB marketability across different price points. Market simulations demonstrated that a 10% increase in Hardness Shore A could increase PHB market volume by 431.5 million kg while reducing emissions by 188.7 kg CO 2 . A similar 10% increase to Sheet Extrusion Temperature could yield a 297.5 million kg volume increase and a 99.2 kg CO 2 reduction in emissions. Critically, this approach is agnostic to the specific methods required to achieve these property changes, instead providing material scientists with quantitative, data-driven targets for R&D prioritization. Here, this framework offers a novel methodology for evaluating biopolymer competitiveness and supporting strategic decisions to accelerate PHB market adoption and contribute to decarbonization of the plastics industry.

09 BIOMASS FUELS↗

Poisson-response Tensor-on-Tensor Regression and Applications

We introduce Poisson-response tensor-on-tensor regression (PToTR), a novel regression framework designed to handle tensor responses composed element-wise of random Poisson-distributed counts. Tensors, or multi-dimensional arrays, composed of counts are common data in fields such as inter national relations, social networks, epidemiology, and medical imaging, where events occur across multiple dimensions like time, location, and dyads. PToTR accommodates such tensor responses alongside tensor covariates, providing a versatile tool for multi dimensional data analysis. We propose algorithms for maximum likelihood estimation under a canonical polyadic (CP) structure on the regression coefficient tensor that satisfy the positivity of Poisson parameters and then provide an initial theoretical error analysis for PToTR estimators. We also demonstrate the utility of PToTR through three concrete applications: longitudinal data analysis of the Integrated Crisis Early Warning System database, positron emission tomography (PET) image reconstruction, and change-point detection of communication patterns in longitudinal dyadic data. These applications highlight the versatility of PToTR in addressing complex, structured count data across various domains.

97 MATHEMATICS AND COMPUTING↗

FORCE Regression Testing

Via programs including the Light Water Reactor Sustainability and Integrated Energy Systems, the U.S. Department of Energy has invested in the Framework for Optimization of ResourCes and Economics (FORCE) software framework (Idaho National Laboratory 2024a) for the technical and economic analysis of nuclear-integrated energy systems (IES). Nuclear IES expand the use of nuclear from traditional baseload electricity generation to a flexible and adaptive source of combined heat and power. Nuclear heat can be used in the production of a variety of energy currencies such as hydrogen and ammonia as well as other heat applications including water desalination and district heating. FORCE is designed with the intent to provide interconnected analysis tools that enable the accurate technical and economic assessment of specific nuclear IES configurations for individual energy markets. FORCE consists of three main analysis pathways: HYBRID (Idaho National Laboratory 2024b), which contains high-resolution physical models for IES; Holistic Energy Resource Optimization Network (HERON) (Idaho National Laboratory 2024c), which analyzes IES long-term economic viability; and Optimization of Real-time Capacity Allocation (ORCA) (Idaho National Laboratory 2024d), designed for real-time control of IES via digital twins and optimal decision making, including autonomous and remote operation research. Development of the FORCE ecosystem is guided by three pillars: capability, which assures that the computational requirements of IES analysis are met by the software tools; reliability, which provides for consistent code performance and expected behaviors; and accessibility, which lowers the barrier to entry for using the software and accelerates analysis by users beyond the FORCE primary developers. Reliability of the FORCE ecosystem is established according to the American Nuclear Society?s Nuclear Quality Assurance (NQA-1) program [American Society of Mechanical Engineers 1982], with specific levels of software quality assurance (SQA) within NQA-1 applied to each software tool in FORCE. As the tools within FORCE have matured, some integration algorithms to accurately connect the software tools for holistic analysis have been developed and deployed within the FORCE software repository. In accordance with NQA-1 standards, regression tests are required to guarantee the software performs consistently even when new capabilities are added to the software. In this report, we document the deployment of both unit tests, which test the consistent behavior of small pieces of the FORCE code base, as well as integration tests, which test the consistent performance of full use cases for the FORCE integration algorithms. We further document the encapsulation of these tests within a test harness, which collectively checks for each successful test completion on demand. Finally, we document the automation of the test harness using GitHub Actions [GitHub 2024], which require all tests succeed before any new capability or other changes can be added to the FORCE integration software

97 MATHEMATICS AND COMPUTING↗

Results from a multi-laboratory ocean metaproteomic intercomparison: effects of LC-MS acquisition and data analysis procedures

Metaproteomics is an increasingly popular methodology that provides information regarding the metabolic functions of specific microbial taxa and has potential for contributing to ocean ecology and biogeochemical studies. A blinded multi-laboratory intercomparison was conducted to assess comparability and reproducibility of taxonomic and functional results and their sensitivity to methodological variables. Euphotic zone samples from the Bermuda Atlantic Time-series Study (BATS) in the North Atlantic Ocean collected by in situ pumps and the autonomous underwater vehicle (AUV) Clio were distributed with a paired metagenome, and one-dimensional (1D) liquid chromatographic data-dependent acquisition mass spectrometry analysis was stipulated. Analysis of mass spectra from seven laboratories through a common bioinformatic pipeline identified a shared set of 1056 proteins from 1395 shared peptide constituents. Quantitative analyses showed good reproducibility: pairwise regressions of spectral counts between laboratories yielded R 2 values averaged 0.62±0.11, and a Sørensen similarity analysis of the top 1000 proteins revealed 70 %–80 % similarity between laboratory groups. Taxonomic and functional assignments showed good coherence between technical replicates and different laboratories. A bioinformatic intercomparison study, involving 10 laboratories using eight software packages, successfully identified thousands of peptides within the complex metaproteomic datasets, demonstrating the utility of these software tools for ocean metaproteomic research. Lessons learned and potential improvements in methods were described. Future efforts could examine reproducibility in deeper metaproteomes, examine accuracy in targeted absolute quantitation analyses, and develop standards for data output formats to improve data interoperability. Together, these results demonstrate the reproducibility of metaproteomic analyses and their suitability for microbial oceanography research, including integration into global-scale ocean surveys and ocean biogeochemical models.

59 BASIC BIOLOGICAL SCIENCES↗

Uncertainty-Aware Machine Learning for Small-Angle X-ray Scattering Analysis in Autonomous Experimentation

Small-angle X-ray scattering (SAXS) is a powerful high-throughput characterization tool for probing nanoscale structure in native sample environments, providing real-time morphological information such as nanoparticle size and shape during synthesis. However, automated SAXS data analysis for extracting meaningful structural parameters is non-trivial and remains a bottleneck in closed-loop experimentation towards autonomous materials discovery, which demands fast, reliable, and uncertainty-aware data analysis. Here, we develop a machine-learning approach for automated SAXS analysis tailored to closed-loop nanoparticle synthesis. A Random Forest (RF) regression model is trained on 100,000 synthetic SAXS curves generated from polydisperse spherical nanoparticles with realistic background contributions. Using normalized one-dimensional SAXS intensity profiles as input, the RF model directly predicts nanoparticle radius, size polydispersity, and background parameters, while the ensemble standard deviation across trees provides built-in uncertainty quantification (UQ). On synthetic data, we show that combining fit-quality metrics (R 2 , MAE) with thresholds on prediction uncertainty reliably identifies accurate parameter estimates without access to ground truth. We then apply the trained model to 365 experimental SAXS profiles of citrate-reduced gold nanoparticles synthesized using an automated droplet-flow microreactor with in situ SAXS at a synchrotron beamline, classifying the results into high- and low-confidence subsets based on UQ metrics. Finally, we integrate RF-based SAXS analysis into a simulated closed-loop optimization campaign using Gaussian process Bayesian optimization to minimize nanoparticle polydispersity, benchmarking against conventional automated Levenberg–Marquardt fitting. The RF-guided campaign exhibits substantially faster convergence and lower relative opportunity cost (∼0.07 vs ∼0.3), demonstrating that uncertainty-aware machine-learning SAXS analysis significantly enhances the efficiency and robustness of autonomous nanomaterials synthesis workflows.

Bayesian optimization↗