Search NASA⌕ Search

SEARCH · Search NASA

Results for “Regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Benchmark Tracking System for Performance Monitoring

Benchmarking is essential for high-performance software development, particularly for monitoring performance across code iterations. This project focused on enhancing the benchmarking process for Lamellar, an asynchronous runtime for High-Performance Computing (HPC) systems developed at Pacific Northwest National Laboratory. Prior to this work, benchmark results were difficult to track and compare across code versions, presenting significant challenges in identifying performance regressions and long-term trends. The primary objective was to establish a systematic, reproducible approach for measuring performance and detecting regressions following code commits. Our methodology involved three key components: standardizing benchmark outputs, implementing data versioning, and developing analysis tools. We standardized the benchmark output format to JSON Line records containing specific fields (execution time, hardware specifications, and environmental variables). To address data management challenges, we evaluated several options and eventually chose a git repository dedicated to benchmark data. We developed a suite of Python tools that processed benchmark results, enriched them with metadata, and facilitated search in the repository. The resulting system enables more efficient filtering and comparison of performance metrics across commit histories, hardware configurations, and benchmark variants through a unified query interface. Our implementation reduces computational overhead by first checking for existing results through configuration matching before initiating new benchmark runs, thereby conserving resources. The system has been validated by Lamellar developers. It organizes results by benchmark type and build configurations for efficient retrieval. Future developments include a planned Large Language Model interface for predicting benchmark performance, incorporating the criterion package for statistical analysis, which will enable automated detection of statistically significant performance changes, and integration with continuous integration pipelines. Despite these enhancements being reserved for future work, this project has successfully provided the Lamellar development team with a framework for maintaining consistent performance standards and identifying optimization opportunities across workloads and hardware environments.

97 MATHEMATICS AND COMPUTING↗

Can Error Mitigation Improve Trainability of Noisy Variational Quantum Algorithms?

Variational Quantum Algorithms (VQAs) are often viewed as the best hope for near-term quantum advantage. However, recent studies have shown that noise can severely limit the trainability of VQAs, e.g., by exponentially flattening the cost landscape and suppressing the magnitudes of cost gradients. Error Mitigation (EM) shows promise in reducing the impact of noise on near-term devices. Thus, it is natural to ask whether EM can improve the trainability of VQAs. In this work, we first show that, for a broad class of EM strategies, exponential cost concentration cannot be resolved without committing exponential resources elsewhere. This class of strategies includes as special cases Zero Noise Extrapolation, Virtual Distillation, Probabilistic Error Cancellation, and Clifford Data Regression. Second, we perform analytical and numerical analysis of these EM protocols, and we find that some of them (e.g., Virtual Distillation) can make it harder to resolve cost function values compared to running no EM at all. As a positive result, we do find numerical evidence that Clifford Data Regression (CDR) can aid the training process in certain settings where cost concentration is not too severe. Our results show that care should be taken in applying EM protocols as they can either worsen or not improve trainability. On the other hand, our positive results for CDR highlight the possibility of engineering error mitigation methods to improve trainability.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Consumer-Oriented Energy Use and Range Metrics for Battery Electric Vehicles

The present study was motivated by a need to expand information for consumers offered through the FuelEconomy.Gov website. To that end, a power-based modeling approach has been used to examine the effect of steady-speed driving on estimated range for model year 2020 – 2023 battery electric vehicles (BEVs). This approach allowed rapid study of a broader range of BEV models than could be accomplished through vehicle tests. Publicly accessible certification test results and other data were used to perform a regression between cycle-average tractive power requirements and the resulting electrical power. Importantly, this regression enabled estimation of electric power and energy use over a range of steady highway speeds. These analyses in turn allowed projection of vehicle range at differing speeds. The projections agree within 6% with available 65 MPH manufacturer test data. Analyses of vehicles from model years 2020 – 2023 show that the 5-cycle range and energy use values from the window stickers of new vehicles are reasonable values for steady-speed driving at 65 MPH. Range decreases by a median value of approximately 15% for each 10 MPH increase in speed. The 5-cycle energy use (in kW-Hr / 100 miles) and energy economy (in miles / kW-Hr) derived from the 5-cycle energy use are reasonable estimates for 65 MPH steady speed driving. Energy economy decreases by about 0.5 miles / kW-Hr for every 10 MPH increase in speed.

33 ADVANCED PROPULSION SYSTEMS↗

NuGraph2: A Graph Neural Network for Neutrino Event Reconstruction

Neutrino experiments are set to probe some of the most important open questions in physics, from CP violation and the nature of dark matter. The technology of choice for many of these experiments is the liquid argon time projection chamber (LArTPC). In current LArTPC experiments, reconstruction performance often represents a limiting factor for the sensitivity. New developments are therefore needed to unlock the full potential of LArTPC experiments. NuGraph2 is a state of the art Graph Neural Network for reconstruction of data in LArTPC experiments. NuGraph2 utilizes a heterogeneous graph structure, with separate subgraphs of 2D nodes (hits in each plane) connected across planes via 3D nodes (space points). The model provides a consistent description of the neutrino interaction across all planes. NuGraph2 is a multi-purpose network, with a common message-passing attention engine connected to multiple decoders with different classification or regression tasks. These include the classification of detector hits according to the particle type that produced them (semantic segmentation) and the separation of hits from the neutrino interaction from hits due to noise or cosmic-ray background. Additional decoders are being developed, performing tasks such as the regression of the neutrino interaction vertex position. Performance results will be presented based on publicly available samples from MicroBooNE. These include both physics performance metrics, achieving 95% accuracy for semantic segmentation and 98% classification of neutrino hits, as well as computational metrics for training and for inference on CPU or GPU. The status of the NuGraph integration in the LArSoft software framework will be presented, as well as initial studies about model interpretability and injection of domain knowledge.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

NuGraph2: A Graph Neural Network for Neutrino Event Reconstruction

Neutrino experiments are set to probe some of the most important open questions in physics, from CP violation and the nature of dark matter. The technology of choice for many of these experiments is the liquid argon time projection chamber (LArTPC). In current LArTPC experiments, reconstruction performance often represents a limiting factor for the sensitivity. New developments are therefore needed to unlock the full potential of LArTPC experiments. NuGraph2 is a state of the art Graph Neural Network for reconstruction of data in LArTPC experiments [https://arxiv.org/abs/2403.11872]. NuGraph2 utilizes a heterogeneous graph structure, with separate subgraphs of 2D nodes (hits in each plane) connected across planes via 3D nodes (space points). The model provides a consistent description of the neutrino interaction across all planes. NuGraph2 is a multi-purpose network, with a common message-passing attention engine connected to multiple decoders with different classification or regression tasks. These include the classification of detector hits according to the particle type that produced them (semantic segmentation) and the separation of hits from the neutrino interaction from hits due to noise or cosmic-ray background. Additional decoders are being developed, performing tasks such as the regression of the neutrino interaction vertex position. Performance results will be presented based on publicly available samples from MicroBooNE. These include both physics performance metrics, achieving 95% accuracy for semantic segmentation and 98% classification of neutrino hits, as well as computational metrics for training and for inference on CPU or GPU. The status of the NuGraph integration in the LArSoft software framework will be presented, as well as initial studies about model interpretability and injection of domain knowledge.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Predicting Initial Trans-Membrane Pressure for Optimized Operations in UF Unit Using Random Forest

With the growing scarcity of freshwater, innovative process design mechanisms like Reverse Osmosis (RO) are increasingly gaining attention among water treatment utilities to address the rising demand. Ensuring reliable water production necessitates efficient resource utilization, minimizing downtime in (ultra-filtration) UF systems. Recent advancements in machine learning (ML) have enabled the development of accurate data-driven models for Model Predictive Control (MPC), often requiring minimal prior knowledge of underlying physical processes. In this study, we present predictive regression models based on Random Forest (RF) and Auto-Regressive (AR) approaches to forecast the initial Trans-Membrane Pressure (TMP) for each filtration cycle in data generated by Direct Potable Reuse (DPR) systems. The proposed RF-based model demonstrates superior performance compared to baseline methods, including historical mean, Last Observation Carried Forward (LOCF), and naïve AR models, across various forecasting horizons in terms of root mean square error (RMSE) metric. To evaluate how different classes of process variables contribute to TMP dynamics over time, we examine the feature importance of independent covariates across multiple forecast horizons. This analysis provides insight into the temporal relevance of operational and sensor-derived features, guiding control and monitoring strategies. Additionally, the impact of hyperparameter tuning on TMP prediction performance is studied for both direct and recursive RF modelling approaches across increasing forecast horizons. Accurate prediction of initial TMP is critical for optimizing RO operations, as it enables the development of robust modelling frameworks by accurately estimating membrane fouling trends, thereby enhancing process efficiency and long-term reliability. The demonstrated efficacy of the RF-based approach highlights its potential as a tool for real-time decision-making in water treatment systems, paving the way for advanced process optimization and sustainable water resource management.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)↗

Predicting Initial Trans-Membrane Pressure for Optimized Operations in UF Unit Using Random Forest

With the growing scarcity of freshwater, innovative process design mechanisms like Ultra-filtration(UF) units are increasingly gaining attention among water treatment utilities to address the rising demand. Ensuring reliable water production necessitates efficient resource utilization, minimizing downtime in UF systems. Recent advancements in machine learning (ML) have enabled the development of accurate data-driven models for Model Predictive Control (MPC), often requiring minimal prior knowledge of underlying physical processes. In this study, we present predictive regression models based on Random Forest (RF) and Auto-Regressive (AR) approaches to forecast the initial Trans-Membrane Pressure (TMP) for each filtration cycle in data generated by Direct Potable Reuse (DPR) systems. The proposed RF-based model demonstrates superior performance compared to baseline methods, including historical mean, Last Observation Carried Forward (LOCF), and naïve AR models, across various forecasting horizons in terms of root mean square (RMSE) metric. Accurate prediction of initial TMP is critical for optimizing CCRO operations, as it enables the development of robust modelling frameworks that enhance process efficiency and reliability. The demonstrated efficacy of the RF-based approach highlights its potential as a tool for real-time decision-making in water treatment systems, paving the way for advanced process optimization and sustainable water resource management.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)↗

Analysis of Waste Material Feedstocks Using Laser-Induced Breakdown Spectroscopy and Machine Learning

Predicting properties such as heating value, ash fusion temperature, and mineral ash composition from Laser-Induced Breakdown Spectroscopy (LIBS) data can make gasifiers more flexible to different feedstocks. Understanding these feedstock properties in-situ improves feedstock conversion modelling methods that allow for consistent operation, higher carbon conversion, and reduced fouling and erosion rates. The purpose of this study is to demonstrate methods for model creation that take LIBS data as predictor features and estimate higher order material properties as a function of feedstock material properties. Six samples were chosen to represent a mixture of abundant and carbon rich waste materials. LIBS measurements were performed on these samples for elemental wavelengths and intensity values. Laboratory analytical results were obtained for each sample’s heating value, proximate and ultimate analysis, mineral ash composition, ash fusion temperatures, and viscosity temperatures. Thermal conductivity was measured using a HotDisk TPS 2500S. LIBS measurements were processed and used as predictor features for machine learning (ML) models to predict the sample’s material properties. Predictor feature selection algorithms, particularly minimum redundancy maximum relevance (mRMR), reduced the dimensionality of ML models. Many modelling methods such as Gaussian process regression (GPR), regression tree, neural networks (NN), and support vector machines (SVM) were demonstrated to be effective at predicting higher order properties; however, mRMR with GPR stood out as a clear winning combination.

01 COAL, LIGNITE, AND PEAT↗

Poisson Log-Normal Process for Count Data Prediction

Modeling count data is important in physics and other scientific disciplines, where measurements often involve discrete, non-negative quantities such as photon or neutrino detection events. Traditional parametric approaches can be trained to generate integer-count predictions but may struggle with capturing complex, non-linear dependencies often observed in the data. Gaussian process (GP) regression provides a robust non-parametric alternative to modeling continuous data; however, it cannot generate integer outputs. We propose the Poisson Log-Normal (PoLoN) process, a framework that employs GP to model Poisson log-rates. As in GP regression, our approach relies on the correlations between data points captured via GP kernel structure rather than explicit functional parameterizations. We demonstrate that the PoLoN predictive distribution is Poisson-LogNormal and provide an algorithm for optimizing kernel hyperparameters. Furthermore, we adapt the PoLoN approach to the problem of detecting weak localized signals superimposed on a smoothly varying background - a task of considerable interest in many areas of science and engineering. Our framework allows us to predict the strength, location and width of the detected signals. We evaluate PoLoN's performance using both synthetic and real-world datasets, including the open dataset from CERN which was used to detect the Higgs boson at the Large Hadron Collider. Our results indicate that the PoLoN process can be used as a non-parametric alternative for analyzing, predicting, and extracting signals from integer-valued data.

Saha, Anushka [Rutgers U., Piscataway]↗

Model Residuals as Shields: A Two-Level Formulation to Defend Smart Grids From Poisoning Attacks

The advancement of smart grids presents both vast opportunities and heightened cybersecurity risks. Data-driven defense mechanisms, though designed as a shield against these threats, can fall prey to poisoning attacks. We delve into regression settings, underscoring the imperative to fortify defenses against a spectrum of poison ratios, notably those above 0.5—an issue scarcely addressed in prior studies. Recognizing the susceptibilities of smart grids and their manipulable sensors, we exploit the very intent of poisoning attacks, compromising model accuracy, as our defense mechanism. Our proposed two-level optimization framework discerns between poisoned and authentic data based on model residuals, outperforming or matching existing methods in 72% to 77% of precision and 75% to 80% of recalls across various poisoning attacks, poison ratios, and datasets. Once the authentic data are identified, the trained model is adaptable for a variety of applications. Comprehensive evaluations on different smart grid datasets, pitted against myriad poisoning schemes, validate our methodology’s edge over existing methods. Here, we also shed light on the implications of model misspecification originating from temporal auto-correlation, a common feature in Internet of Things and smart grid data.

Adversarial machine learning (ML)↗

Exploring biofiber properties and their influence on biocomposite tensile properties

Biofibers serve as effective reinforcements for neat polylactic acid (PLA) in biocomposites, offering an attractive opportunity to decarbonize the manufacturing sector of the United States by displacing fossil-based reinforcement fibers such as carbon fibers. Also, biofiber production can stimulate economic growth in rural economies, fueling sustainable development. PLA resins are commonly compounded with biofibers to create biocomposites suitable for additive manufacturing. PLA-biofiber composites often exhibit better overall material properties than neat (pure) PLA, but the associations between biofiber properties and the material properties of their biocomposites remain largely unexplored. Hence, this research delves into a comprehensive exploration of diverse biofibers, scrutinizing their physical and chemical attributes, including size, shape, ash content and biochemical composition. The study meticulously analyzes the flow properties of each biofiber and elucidates the ultimate tensile strengths and Young's modulus of corresponding biocomposite samples. Noteworthy correlations between biofiber and biocomposite tensile properties are uncovered, shedding light on critical interrelationships. The study introduces an approach employing regression models to predict the ultimate tensile strength and Young's modulus of biocomposites. These models, validated with a cross-validation technique, exhibit remarkable predictive accuracy, particularly in estimating ultimate tensile strength. © 2024 Oak Ridge National Laboratory managed by UT-Battelle, LLC and The Author(s). Polymer International published by John Wiley & Sons Ltd on behalf of Society of Chemical Industry.

36 MATERIALS SCIENCE↗

Developing a robust strength model using physically-informed genetic programming

The strength of materials is influenced by a range of external conditions, such as temperature and deformation rate. Consequently, materials that demonstrate substantial variations in their mechanical behavior due to fluctuations in temperature and strain rate require complex strength models to accurately predict material performance in real-world applications. To predict such complex behavior, a robust and flexible strength model is necessary. In this work, we utilize genetic programming-based symbolic regression (GPSR) to develop data-driven strength models that accurately represent the measured stress–strain responses of tin across a wide range of strain, strain rate and temperature regimes. The GPSR models are constrained by physically-informed conditions, which leads to significant improvement in extrapolation. The best model is integrated into a multi-physics code to perform Taylor impact simulations, validating the model’s accuracy and robustness. In conclusion, the model predictions showed excellent agreement with experimental results, particularly when compared to predictions using traditional strength models.

Genetic programming↗

GPR_calculator: An on-the-fly surrogate model to accelerate massive nudged elastic band calculations

We present GPR_calculator, a package based on Python and C++ programming languages to build an on-the-fly surrogate model using Gaussian Process Regression (GPR) to approximate computationally expensive electronic structure calculations. The key idea is to dynamically train a GPR model during the simulation that can accurately predict energies and forces with uncertainty quantification. When the uncertainty is high, the costly electronic structure calculation is performed to obtain the ground truth data, which is then used to update the GPR model. To illustrate the effectiveness of GPR_calculator, we demonstrate its application in Nudged Elastic Band (NEB) simulations of surface diffusion and reactions, achieving 3-10 times acceleration compared to pure ab initio calculations. The source code is available at https://github.com/MaterSim/GPR_calculator.

Gaussian process regression↗

Leveraging hyperspectral phenotyping for accurate, non-destructive prediction of metabolite profiles in poplar under drought stress

Accurately predicting drought tolerance in woody perennial bioenergy crops is critical for sustainable biomass production under fluctuating precipitation. Hyperspectral imaging (HSI) in the visible-near-infrared (VNIR) and shortwave-infrared (SWIR) ranges offers a promising approach for predicting plant biochemical traits, yet its application in metabolite profiling remains underexplored. We integrated VNIR+SWIR HSI with untargeted metabolomics to investigate drought-induced metabolic shifts in Populus leaves from eight Populus genotypes. Metabolite profiling identified 127 compounds, with 73 showing significant drought responses spanning amino acids (AA), carbohydrates (CHO), phenolic glycosides (PG), organic acids (OA), fatty acids and alcohols (FA), terpenes (T), phenolic metabolites (P), and unclassified metabolites. Spectral analysis revealed consistently higher reflectance across VNIR and SWIR wavelengths in drought-stressed plants, corresponding with increased accumulation of AA and reduced CHO and PG levels. Least absolute shrinkage and selection operator (LASSO) regression modeling identified robust spectral predictors of metabolite concentrations, associating VNIR wavelengths (500–700 nm) predominantly with AA and P, whereas SWIR wavelengths (1680–1700 nm) reliably predicted CHO, OA, and T. Several stable spectral-metabolite associations persisted across the two watering regimes (drought vs. well-watered), highlighting their potential as spectral biomarkers for non-destructive stress monitoring. Minimal genotype-specific variation suggests that observed spectral and metabolic responses were driven primarily by environmental factors, likely reflecting limited genetic diversity among the commercial Populus genotypes examined. This work establishes VNIR+SWIR hyperspectral imaging as a powerful, non-destructive phenotyping tool for precision monitoring and targeted improvement of drought resilience in bioenergy crops.

Biochemical trait prediction↗

Quantifying spatial and vertical variations in soil C:N relationships in permafrost-affected landscapes

Permafrost regions are experiencing rapid changes that affect carbon (C) and nitrogen (N) cycles, with implications for vegetation dynamics and gas exchanges with the atmosphere. Soil C:N ratio is a key indicator of organic matter quality, yet spatial estimates of N stocks and C:N ratios lag behind those for C. We used quantile regression forests to compare direct and indirect digital soil mapping approaches for predicting soil C:N ratios at 0–30, 30–60, and 60–100 cm depths across a latitudinal transect in Alaska. The indirect approach – deriving C:N from separately predicted C and N stocks – outperformed direct mapping for the surface layer (0–30 cm), while direct mapping was marginally better at greater depths. However, prediction accuracy decreased with depth for both methods. Temperature and topography were the most important predictors. Both approaches overestimated low and underestimated high C:N ratios, with direct mapping showing greater bias. Our results underscore the challenges of modeling C:N ratios in heterogeneous, data-sparse permafrost soils, but also suggest that indirect mapping holds promise if supported by more extensive datasets.

54 ENVIRONMENTAL SCIENCES↗

"Hidden" hydrothermal technical potential & technoeconomics: Revealing permeability & fluids with more data

Historical hydrothermal estimates have largely relied on temperature or heat flow estimates ignoring the need for natural flowing fluids. More accurate hydrothermal estimates require some indication of permeability and fluids that naturally exist in the subsurface. This paper describes a novel approach that includes proxies of permeability and fluids in hydrothermal estimates by leveraging the relatively data-rich Great Basin. Specifically, nameplate capacities (megawatts) of operating geothermal plants, negative (0 megawatt) locations and 48 geophysical and geologic features are used to used in eXtreme Gradient Boosting (XGBoost) regression to make hydrothermal capacity predictions. Additionally, this work inputs the XGBoost-based hydrothermal predictions into the Renewable Energy Potential (reV) model to quantify technical capacity, its uncertainty and techno-economics. Compared to historical hydrothermal estimates, these predictions adhere to the 37 operating geothermal plants and negative locations. We present a method for subsampling the negative sites to bring the labels into balance that uses the geologic domain knowledge to proportionally represent negatives. Overall, the distributions of the hydrothermal technical capacity and the site levelized cost of energy are respectively much tighter, lower and more accurate than the previous estimates for the Great Basin, as they include geological and geophysical surrogates for permeability and fluids. Percentile (50th and 90th, median and high estimate, respectively) models provide bookends for these metrics.

13 HYDRO ENERGY↗

Quantifying market volume sensitivity to material property modifications in polyhydroxybutyrate: A parametric analysis approach

Polyhydroxybutyrate (PHB), a biodegradable biopolymer, represents a promising alternative to petroleum-based thermoplastics. However, despite consistent market growth, PHB faces persistent commercialization challenges that limit widespread adoption. Existing research has focused predominantly on optimizing PHB production processes, leaving a critical gap in understanding which material property modifications would most effectively enhance market competitiveness. This study addresses this gap by systematically analyzing the relationship between polymer material properties and market performance using U.S. market data from 2008 to 2021 for 21 thermoplastic polymers across 19 material properties. We employed principal component regression to identify property modifications that could maximize market volume while reducing CO 2 emissions. Our parametric analysis revealed that two specific material properties – Hardness Shore A and Sheet Extrusion Temperature – significantly influence PHB marketability across different price points. Market simulations demonstrated that a 10% increase in Hardness Shore A could increase PHB market volume by 431.5 million kg while reducing emissions by 188.7 kg CO 2 . A similar 10% increase to Sheet Extrusion Temperature could yield a 297.5 million kg volume increase and a 99.2 kg CO 2 reduction in emissions. Critically, this approach is agnostic to the specific methods required to achieve these property changes, instead providing material scientists with quantitative, data-driven targets for R&D prioritization. Here, this framework offers a novel methodology for evaluating biopolymer competitiveness and supporting strategic decisions to accelerate PHB market adoption and contribute to decarbonization of the plastics industry.

09 BIOMASS FUELS↗

Absorption dissymmetry factor enhancement: A data-driven approach to unravel the synthesis knobs of chiral 2D perovskites

Chiral 2D metal halide perovskites (MHPs) are promising for spin-optoelectronic applications, yet their absorption dissymmetry factor (g abs ) exhibits significant variability due to complex, co-dependent structural and experimental factors. Here, we established a data-driven framework using Pearson’s correlation, ANOVA, and Gaussian process regression to identify and model key synthesis “knobs” governing these properties. The analysis revealed that solvent choice is the primary factor driving variability. For acetonitrile-based films, g abs was maximized by optimizing annealing temperature and film thickness. Conversely, films from higher boiling point solvents showed complex dependencies on annealing temperature, excitonic integral intensity, and film texture. These statistical correlations provide a roadmap for the rational design of high-performance chiral MHPs and establish a foundation for future machine learning-driven material exploration.

ANOVA↗