Search NASASearch

SEARCH · Search NASA

Results for “Regression model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Machine Learning-Driven Reliability Estimation of PV Inverters Considering Alert-Ambient Variability

Weather-induced spatio-temporal degradation limits outdoor PV inverter lifetime and reliability, necessitating advanced data analysis. This study employs a top-down, data-driven approach utilizing multiple machine learning (ML) algorithms to estimate inverter reliability in a 1.4 MW PV power plant, considering factors such as irradiance, humidity, temperature, time of day, and weather conditions. An extensive alert dataset from 17 identical inverters, including alert types, propagation, and frequency, reveals significant correlations with environmental factors and inverter output power, enabling the construction of a performance reliability model. Dual-stage supervised-ML models are evaluated for accuracy, with the ‘classification-regression’ model by an artificial neural network (ANN) tested on the averaged “Alert-Ambient” dataset, which is outperformed by ‘clustering-regression’ models using random forest (RF) and K-Nearest Neighbors (KNN) on individual inverter datasets. K-means clustering applies principal component analysis to reduce dimensions, achieving improved accuracy beyond the 80% achieved by ANN on the averaged dataset. Second-stage regression estimates inverter reliability with a mean square error of 0.0195 on the averaged dataset and as low as 0.002 on individual inverter datasets using RF. Furthermore, these findings highlight the method's suitability for estimating PV inverter output reliability under ambient conditions, essential for digital twin development and related applications.

14 SOLAR ENERGY

Remote quantification of Cm(III) and HNO 3 by fluorescence spectroscopy and chemometrics

A unique approach to remotely quantify Cm(III) (0–100 µg mL −1 ) in HNO 3 (1–12 M) using steady-state laser fluorescence spectroscopy and multivariate regression models was developed. Photoluminescence is amenable to remote measurements using fiber-optic cables and is sensitive to numerous lanthanide and actinide species. In-line measurements can provide feedback to support complex processing in harsh environments (e.g., hot cells) to help guide and optimize radiochemical separations. In this work, Cm(III) spectra were acquired remotely in a glove box as a function of HNO 3 concentration to better understand spectral characteristics and evaluate the utility of multivariate regression models in this system. Furthermore, the Cm(III) fluorescence peak shape, width, position, and intensity changed significantly as a function of HNO 3 concentration, likely because of the displacement of emission quenching inner-sphere water molecules and complexation with nitrate ions. Despite significant covariance and nonlinearity in the data, a D-optimal design strategy successfully minimized training set sample size and was used to build effective partial least squares regression models for Cm(III) and HNO 3 concentrations without a priori knowledge of solution conditions. Chemometrics for modeling complex fluorescence spectra are promising and may find widespread applicability for online analysis in numerous chemical systems found in the nuclear field.

Actinide

Field Validation of Thermoelectric Generation System at Holcim Cement Plant in Alpena, Michigan

Executive Summary Project Background The Industrial Technology Validation (ITV) program aims to identify and demonstrate the performance of new, emerging, and underutilized energy-saving technologies in the industrial sector to help inform decisions to help accelerate their commercialization and deployment, as well as to help make industries more competitive. This ITV demonstration evaluated a thermoelectric generation (TEG) technology at a cement plant, aiming to reduce energy demand in the cement industry. A median cement plant consumes 5.73 million British thermal units per ton of clinker production (resulting in 0.838 metric tons of carbon dioxide [CO₂] emissions per ton of clinker) (Boyd and Zhang 2011, EPA 2021), equivalent to approximately 6.9 trillion British thermal units (TBtu) per year in energy consumption at a cement plant producing 3,300 tons of clinker per day.¹ Collaborating with Holcim, Advanced Thermovoltaic Systems (ATS) developed and deployed a pilot-scale thermoelectric power system to efficiently capture and convert waste heat to electricity. The system leverages the Seebeck effect to convert temperature differences on two sides of semiconductor cartridges into electrical power (ScienceDirect, n.d.). This generation is realized with minimal moving parts compared to existing waste-heat-to-generation solutions and allows capture from heat sources with temperatures as low as 150°C. This project aimed to validate a scalable solution applicable for capturing medium-temperature waste heat, including ambient losses from other high-temperature processes, and high-temperature sources less suitable for other waste-heat-to-power solutions. By recovering this otherwise wasted heat, this project intends to validate improvements to overall process efficiency through reduction in purchased electricity, thereby reducing operational costs while enhancing resiliency and competitiveness. Description and Scope This study evaluated the performance of a TEG system from ATS as a solution to convert waste heat into useful power at a Holcim cement plant in Alpena, Michigan. This plant is a fully integrated cement plant that has been operating since 1907. The facility operates continuously (24/7/365) with approximately 250 employees and five long dry kilns, yielding a total production capacity of 7,852 tons of cement per day (EPA 2023). Currently, the Alpena plant uses waste heat boilers to convert waste heat from the exhaust of each kiln into steam, which drives steam turbine generators. The ATS TEG is being evaluated for its potential to supplement the steam turbines by capturing the remaining lower grade heat. This technology is also being considered for other Holcim plants where steam turbines are not a viable option. ATS installed a pilot-scale TEG unit with an array of 582 individual thermoelectric semiconductor cartridges, of which 573 were operational. The cartridges are sandwiched between 48 hot plates and 49 cold plates. Each cartridge is designed to generate 20 watts (W) of gross power at a hot-side temperature of 240°C and cold-side temperature of 20°C. As such, the total gross generation capacity of the installed system is 11.5 kilowatts (kW) at design conditions. The system configuration for the evaluation was designed to prioritize convenience of installation and minimize disruption to production at the site, while ensuring that the heat required can be obtained for evaluating the TEG system at various operational conditions. To accomplish this, a portion of the steam supplied to Alpena’s steam turbine generation system was diverted to be used as the heat source for the TEG system, while water was supplied to the cold side of the system from nearby Lake Huron. This configuration was designed for the evaluation of the pilot-scale system to assess the performance at different conditions. A commercial-scale system will likely vary from the pilot system depending on typical configurations, including both scale and application. Future commercial applications of the ATS system would involve integrating the system into the exhaust from kiln preheaters, clinker coolers, or radiant heat capture from kiln shells for the heat source. For the cold source, a range of cooling solutions can be considered, including a mechanical cooling system, depending on the location and the application. To increase the generation capacity for commercial applications, the technology provider is working toward developing a commercial-scale TEG system, which would combine multiple TEG units (each similar in design to the pilot system) together. The scope of this evaluation includes the pilot-scale TEG system and all impacted equipment including pumps, controllers, and power handling equipment. Study Objectives The evaluation's goal was to assess the potential of the ATS TEG system to generate useful electrical power by capturing waste heat from cement production kilns. The objectives of this study are to evaluate and verify the following claims made by ATS regarding the pilot-scale system installed at the Holcim Alpena plant. The following design parameters and claims are also outlined in Table ES- 1 and Table ES- 2: • Gross Power: The thermoelectric system converts heat into power to create gross power, the total measured power generated by the system. The 573 active cartridge pilot-scale system is expected to generate 11.5 kW of gross power at the designed hot-side temperature of 240°C and cold-side temperature of 20°C. Power production is dependent on the temperature difference between the heat source (ultimately from the waste heat) and cold temperature supply source. • Net Power: The net power is the total usable power provided to the site by the TEG system after deducting parasitic power loads from the gross generated power. Supplementary equipment is required to operate the TEG system including pumps, controllers, and, in certain anticipated applications, mechanical cooling, which introduce parasitic loads to system operation. After deducting the parasitic loads from the gross power generation, ATS anticipates achieving a net power generation of 7.5 kW from the pilot-scale system. • Thermal Efficiency: The thermal efficiency is the percent of the total heat transferred to the TEG system that is converted to gross power. Historically, TEGs have a thermal efficiency of 2%–5% (DOE 2008). Prior industrial-scale TEG systems, such as the E1 TEG offered by Alphabet Energy, operated at an efficiency of 2.5% (Lamonica, 2014). ATS anticipates achieving an average efficiency of 4.8% or higher in converting heat energy to usable electricity. • Cartridge Performance: The TEG system comprises 573 active individual semiconductor cartridges, each of which generates a portion of the total power. Cartridge optimization and selection is an important design consideration for potential future TEG system design performance. Therefore, understanding the distribution of gross power and efficiency within the pilot system is vital to understanding what is achievable. At a design hot-side temperature of 240°C and cold-side temperature of 20°C, ATS anticipates a cartridge performance of 20 W of gross power per cartridge at an efficiency of 4.8% per cartridge. In addition to evaluating the claimed performance of the TEG pilot-scale unit, the study estimated the potential annual impacts of a scaled-up commercial system used to capture kiln waste heat over annual operations. The evaluation estimated the gross and net annual electric generation achievable by capturing heat from the two proposed tap-in points: the kiln exhaust and the clinker cooler exhaust; see Section 2.1 for details. Two use cases were examined: • Holcim Alpena: The Holcim Alpena site consists of long dry kilns with superheater boilers, which differs from the rest of Holcim’s cement plant portfolio and results in lower waste heat temperatures. The study estimates gross and net annual generation using the superheater boiler exhaust and clinker cooler exhaust, based on 2023 operational data. • Typical Installation: Common cement plants have preheater kilns with higher exhaust temperatures than Holcim Alpena across a range of production rates. The study estimates gross and net annual generation using the preheater exhaust and clinker cooler exhaust, with a sensitivity analysis to account for the typical range of preheater exhaust temperatures, clinker cooler exhaust temperatures, and clinker production rates. Methodology The evaluation methodology followed a measurement and verification (M&V) strategy based on the International Performance Measurement and Verification Protocol Option B through comprehensive measurements and analyses of the affected systems. Evaluation data was collected from March 9 to March 11, 2024, the test period of the pilot TEG system. During the test period, in coordination with the ITV team, the ATS team adjusted system operations to capture the range of variability expected for each of the variables pertinent to performance of the system. The methodology consisted of two parts: evaluating the performance of the pilot unit's TEG system and estimating the annual TEG impact in terms of gross and net power based on a given waste heat profile. First, the evaluation of the thermoelectric generation performance of the pilot unit relative to the claims was performed by analyzing the collected test data. Gross power of the pilot TEG system was directly measured. Net power was determined by deducting the measured parasitic power from the gross power. The gross power generation was compared to heat transferred to the system by the working fluid (which was heated by steam generated from the kiln waste heat) to calculate the thermal efficiency achieved by the system. Performance of individual semiconductor cartridges within the pilot array was also assessed in terms of measured gross cartridge power and calculated cartridge thermal efficiency. The second part of the evaluation estimated the annual TEG impacts in terms of gross power and net power (calculated from the difference between gross power and parasitic power). This analysis comprised development of mathematical regression models for gross power and parasitic power, with assessment of each model’s goodness-of-fit characteristics to ensure satisfaction of statistical requirements. The models predicted the gross power generation, the parasitic load based on the temperature difference between the hot working fluid and the cold-side fluid (cold water from Lake Huron) entering the system, the volumetric flow rate of the cold-side fluid at the inlet, and the volumetric flow rate of the hot working fluid at the inlet. The annual impact analysis considered a theoretical commercial-scale system sized to capture the available waste heat at a cement plant, consisting of linked pilot-scale units that receive heat from a theoretical gas-to-working-fluid heat exchanger. To estimate annual impacts at the Alpena plant, the gross power and parasitic power regression models were applied to the arrays in the theoretical commercial-scale system. The heat supplied to the unit was calculated based on the kiln run time, annual production, kiln exhaust waste heat, and clinker cooler waste heat derived from 2023 Holcim Alpena kiln operational data. Net power impacts were calculated by deducting the resulting parasitic power from the estimated gross power. Inputs for the model were generated from a combination of hourly data, assumed design considerations for TEG system scale-up from the pilot-scale unit, and assumptions regarding TEG system operations. This analysis was then used as the basis for estimating annual impacts of typical TEG installation at cement plants, by applying sensitivity analyses to key kiln operational characteristics including kiln preheater exhaust temperatures, cooler clinker exhaust temperatures, and plant daily production rates across a range of expected values. Project Results/Findings Table ES- 2 and Table ES- 2 provide a summary of the operating conditions and evaluation results compared to the stated claims from the technology provider. Key takeaways include: • Gross Power: The peak gross power achieved during the testing period was 10.0 kW, compared to the 11.5 kW expected for 573 active cartridges. The claimed gross power was associated with a target hot side of 240°C; however, the system only received a maximum hot-side mean plate temperature of 212°C during the testing period. • Net Power: The pilot-scale unit exceeded the claims for net power, achieving a peak of 7.7 kW net compared to a claim of 7.5 kW. One factor contributing to the higher achieved net power is the relatively high water pressure available through Lake Huron. The pilot TEG system did not require cold-side pumps during the test, whereas most installations would. This reduced the parasitic loads on the system, ultimately contributing to higher net power relative to the gross power. • Thermal Efficiency: The pilot-scale unit outperformed the claimed efficiency, achieving a peak system efficiency of 5.0% thermal efficiency compared to the stated 4.8%. • Cartridge Performance: To compare cartridge performance against claims, the study focused on the third day of testing, which aimed for conditions closest to the design specifications, with a hot side of 240°C and cold-side exit temperature of 6.4°–30°C. On this day, the mean gross power observed in the cartridges within the TEG array was 18.1 W/cartridge, and the peak performance was 34.7 W/cartridge. The estimated mean cartridge efficiency was 5.2%, and the estimated efficiency at peak gross cartridge power was 10%. The regression models developed for gross power generation and parasitic loads were used to estimate the generation impact for given heat input to the TEG from the working fluid (captured from the waste heat) and from the cold loop (Lake Huron) on an hourly basis for a year of operation. Based on this analysis, installation of a commercial-scale TEG system at the Holcim cement plant in Alpena, Michigan, with a waste heat exchanger of 0.85 effectiveness, would generate up to 391 kW of net power, translating to between 920,000 and 1,800,000 kilowatt-hours (kWh) in net electricity per year. Based on typical grid emissions for Alpena, this would avoid estimated net emissions by 752 metric tons of CO₂ annually.² The sensitivity analysis estimated that typical TEG system installations at cement plants could generate an average of 56–1,040 kW of net power, or between 488,000 and 9,110,000 kWh of net energy. This generation potential is most significantly affected by plant production rates and also influenced by preheater and clinker cooler exhaust temperatures. Applying the national average emission rate, typical commercial-scale installations at Holcim plants are projected to avoid between 182 and 3,401 metric tons of CO₂ annually per site. Table ES- 3 shows a summary of the estimated annual impacts.³ While parasitic loads are significant and vary by application, this analysis assumed the use of heating loop pumps and access to Lake Huron as a cold sink. This setup assumed no need for cooling loop pumps due to the available water pressure at the test site. Applications that require cooling towers or additional equipment are likely to experience higher parasitic loads. Therefore, the study’s estimates are most applicable to scenarios with similar parasitic load configurations—namely, access to a high-pressure cold sink. Applicability to other locations may be limited, as differing conditions could necessitate additional pumps and cooling systems, potentially impacting performance significantly.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Event-Based Energy Impact Tracking and Forecasting with Limited Measurements for Rooftop Units

Packaged air conditioning units and heat pumps, also known as rooftop units (RTUs), are responsible for almost 133 billion kWh of electricity usage annually on site for space cooling U.S. commercial buildings. In addition, the use of heat pumps is a trend we expect to accelerate as buildings transition from fossil fuel-based heating to electricity as a key step for decarbonizing the U.S. commercial buildings sector. However, the operation conditions and energy use of RTUs and heat pumps are usually not well monitored as they are not commonly integrated with building automation systems and lack exposed sensing and control points. To fill this gap, this paper proposes a framework for tracking and forecasting energy impacts resulting from degradation of performance and improved performance for unit servicing using limited data. The proposed framework makes use of a constrained dataset, specifically measurements of the outdoor air temperature and the power demand of individual RTUs, to track and forecast changes in energy use associated with changes in performance over various temporal horizons ranging from days to weeks. Following the detection of an RTU fault, performance degradation, or performance improvement, the framework employs a prediction model to assess the cumulative energy impact. We demonstrate the effectiveness of the method with field-collected data for servicing and degradation examples and compare the predicting accuracy of Gradient Boosting Decision Tree (GBDT) Regression models to Support Vector Regression and Linear Regression models. The results show that GBDT achieved the best accuracy for time-series validation datasets for the servicing and degradation cases, and the prediction model was able to track the cumulative energy impacts of events. The proposed framework can inform building owners of the cumulative change in energy usage of RTUs associated with performance degradation, performance improvement, or a fault.

packaged air conditioners, packaged heat pumps, ro

Fusion of Experiments and Simulations for Real-Time Identification of Pipeline Defects

In this study, we explored fusion of experiments and simulations for real time identification of pipeline defects across physical and non-physical domains. The challenges associated to data processing were addressed and a combined classification models was presented via CNN models. In addition, regression model based on XGBOOST is built to determine the defect location and defect dimension from data-driven features of guided wave signals captured by SMS fiber optic sensor.

deep learning

Fusion of Experiments and Simulations for Real-Time Identification of Pipeline Defects

In this study, we explored fusion of experiments and simulations for real time identification of pipeline defects across physical and non-physical domains. The challenges associated to data processing were addressed and a combined classification models was presented via CNN models. In addition, regression model based on XGBOOST is built to determine the defect location and defect dimension from data-driven features of guided wave signals captured by SMS fiber optic sensor.

deep learning

Speedup of UEDGE Parameter Scans Using Machine-Learning Optimized OpenMP Parallelization and a Continuation Solver

This article presents the OpenMP parallelization of the preconditioning Jacobian assembly and right‐hand side residual evaluation in UEDGE. A continuation algorithm, utilizing the internal NKSOL implicit Jacobian‐Free Newton‐Krylov solver to efficiently scan physical parameters, is also presented. The implemented parallelization reduces the computational time for a benchmark scan run on 32 threads by compared to the serial version when using trained random forest regression models to identify the optimal decomposition of the system of equations. Random forest regression models applied to the UEDGE time‐dependent and continuation solver algorithms did not yield meaningful improvement in computational performance. A benchmark DIII‐D gas injection rate scan in the 0.35–0.75 kA interval, performed on a test cluster using the parallelized code and continuation solver, produced 1066 steady‐state solutions with a 22 s average wall‐clock computational time per steady‐state solution.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Predicting Damages to Remainder Parcels in Right-of-Way Acquisitions for Expanding Transportation Infrastructure: Using a Truncated Finite-Mixture Model

Right-of-way acquisition is a critical component of transportation infrastructure development. Transportation infrastructure projects cannot proceed without proper right-of-way acquisition or may face significant delays. State Departments of Transportation frequently acquire parcels of land for roadway expansion projects. A majority of these acquisitions can be partial takings, referring to a portion of a parcel that is acquired. The remainder of the property usually suffers economic changes due to the partial acquisition, which can be calculated as damage percentages. The damage percentage represents the extent to which the remaining land or property value has been diminished due to the acquisition. It reflects the remaining property value percentage that may have been lost or compromised due to the acquisition. Here, this study aims to provide a robust model to estimate damage percentages to the remainder parcels that may help state Departments of Transportation appraisers make early predictions about the damages in cases involving partial takings. The research uses 509 appraisal reports from the Tennessee Department of Transportation to identify the key parcel attributes that influence the percentage of damages. Three regression models are developed: a linear regression model, a finite-mixture model (FMM), and a truncated FMM with two latent classes. The modeling results show that the truncated FMM with two classes outperforms the other models. To validate the models, actual sales data is collected and analyzed for 59 properties, and the results suggest that the model predictions are fairly accurate. A predictive tool is developed based on the models to help appraisers anticipate right-of-way damages under different scenarios and can provide early predictions about the damages.

42 ENGINEERING

RxnRover/amlro

AMLRO (Active Machine Learning Reaction Optimizer) is an open-source framework designed to accelerate chemical reaction optimization using active learning with classical machine learning regression models. AMLRO integrates space-filling sampling strategies (e.g., Sobol and Latin Hypercube sampling) with iterative model training, prediction, and experiment selection to efficiently navigate complex reaction spaces. The platform supports multiple regression models, flexible multi-objective definitions, and user-defined parameter bounds, enabling data-efficient optimization from small initial datasets. AMLRO is designed for ease of use by experimentalists and can operate as a standalone decision-support tool or be integrated into closed-loop automated experimentation workflows.

Kulathunga, Dulitha Prasanna [Iowa State Universit

Predictive modeling of Néel temperature in austenitic alloys using CALPHAD and data analytics

The Néel temperature is a crucial yet often overlooked parameter in calculating the stacking fault energy (SFE) of austenitic alloys. Several empirical equations have been proposed to estimate the Néel temperature of austenitic alloys, which are then used to calculate the SFE and explain deformation mechanisms. However, these empirical equations, typically derived using linear regression algorithms, are often simplistic and may fail to capture the complex interactions among multiple alloying elements that influence the Néel temperature. Moreover, their applicability is usually limited to specific compositional ranges. In this study, we propose a CALPHAD based approach and develop a surrogate decision tree based regression model capable of capturing the interactions among multiple alloying elements to predict the Néel temperature. Predictions from both the CALPHAD approach and the regression model show close agreement with experimental measurements reported in the literature. In conclusion, the implications of accurate Néel temperature predictions on the calculated SFE and deformation mechanisms are also discussed.

36 MATERIALS SCIENCE

The Effect of Updraft Entrainment on Convective Cell Deepening in Realistic Large-Eddy Simulations

Entrainment of surrounding cooler and drier air into convective updrafts is one of the key processes that influence deep convection initiation and growth. Numerous studies have investigated the effect of entrainment on isolated convective cloud growth in idealized simulations, but the importance of this effect in realistic conditions with many interacting convective clouds remains uncertain. We examine the impact of entrainment on the depth reached by convective clouds in realistic large-eddy simulations (LES) over central Argentina during the Cloud, Aerosol, and Complex Terrain Interactions (CACTI) field campaign. Cloudy updrafts and their associated properties are assigned to convective cells tracked with radar reflectivity signatures. Several thousand convective cells are tracked over two high convective available potential energy (CAPE) and two low CAPE cases that support cells of varying depths and intensities. Entrainment is calculated explicitly as the fluxes of air into the outer surface of each cloudy updraft. Single-predictor logistic regression models are used to determine the relative importance of updraft, near-updraft, and preconvective initiation atmospheric conditions in predicting whether convective cells become deep. We then build a multiple-predictor regression model pairing important updraft and meteorological metrics with fractional entrainment rate. The probability of cells transitioning to deep convection is most sensitive to ambient 600-hPa relative humidity (42% of total metric contribution to cloud depth predictability), followed by low-level CAPE (28%), cloud-base updraft width (19%), and fractional entrainment (11%). Thus, the initial width of the updraft along with potential buoyancy and its dilution through the midtroposphere collectively determine whether deep convection will result from shallower clouds.

54 ENVIRONMENTAL SCIENCES

Computationally efficient and error aware surrogate construction for numerical solutions of subsurface flow through porous media

Limiting the injection rate to restrict the pressure below a threshold at a critical location can be an important goal of simulations that model the subsurface pressure between injection and extraction wells. The pressure is approximated by the solution of Darcy’s partial differential equation for a given permeability field. The subsurface permeability is modeled as a random field since it is known only up to statistical properties. This induces uncertainty in the computed pressure. Solving the partial differential equation for an ensemble of random permeability simulations enables estimating a probability distribution for the pressure at the critical location. These simulations are computationally expensive, and practitioners often need rapid online guidance for real-time pressure management. An ensemble of numerical partial differential equation solutions is used to construct a Gaussian process regression model that can quickly predict the pressure at the critical location as a function of the extraction rate and permeability realization. The Gaussian process surrogate analyzes the ensemble of numerical pressure solutions at the critical location as noisy observations of the true pressure solution, enabling robust inference using the conditional Gaussian process distribution. Our first novel contribution is to identify a sampling methodology for the random environment and matching kernel technology for which fitting the Gaussian process regression model scales as O ( n log n ) instead of the typical O ( n 3 ) rate in the number of samples n used to fit the surrogate. The surrogate model allows almost instantaneous predictions for the pressure at the critical location as a function of the extraction rate and permeability realization. Our second contribution is a novel algorithm to calibrate the uncertainty in the surrogate model to the discrepancy between the true pressure solution of Darcy’s equation and the numerical solution. Finally, although our method is derived for building a surrogate for the solution of Darcy’s equation with a random permeability field, the framework broadly applies to solutions of other partial differential equations with random coefficients.

54 ENVIRONMENTAL SCIENCES

HydraGNN_Predictive_GFM_2024 - Ensemble of predictive graph foundation models for ground state atomistic materials modeling

We provide the ensemble of fifteen pre-trained graph foundation models (GFMs) for atomistic materials modeling applications. Each one of the fifteen GFMs has been trained on five open-source datasets that (once aggregated) amount to over 154 million atomistic structures, which cover over two-thirds of the natural elements of the periodic table and that comprises a broad set of organic and inorganic compounds. This vast set of atomistic structures comprises ground state configurations that are dynamically stable (i.e., equilibrated structures with atomic forces approximately close to zero values) as well as dynamically unstable structures (i.e., non-equilibrium structures with non-negligible non-zero values of atomic forces). The ensemble of datasets aggregated does NOT include excited states. The datasets have been curated to remove atomistic structures with spectral norm of the force tensor above 100 eV/angstrom. Moreover, a linear term of the energy was computed for each dataset using a linear regression model that uses the chemical concentration of each natural element as regressor. The linear term predicted by the linear regression model has been subtracted from each original energy value to perform a re-alignment of the energy values across different electronic structures approximation theories performed to generate the diverse multi-source, multi-fidelity datasets. The folder "ADIOS_files" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "ADIOS_files" directory contains 6 sub-directories named as follows: - ANI1x-v3.bp - MPTrj-v3.bp - OC2020-20M-v3.bp - OC2020-v3.bp - OC2022-v3.bp - qm7x-v3.bp Each sub-directory contains the pre-processed datasets converted in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used to the development, training, and performance testing of the ensemble go predictive graph foundation models. Each GFM was developed using HydraGNN (https://github.com/ORNL/HydraGNN) as underlying graph neural network (GNN) architecture. The multi-task learning (MTL) capability of HydraGNN was used to simultaneously train the GFMs on labeled values for direct predictions of energy (a total system property of an atomistic structure that measures the chemical stability) and atomic forces (an atomic level property of an atomistic structure that measures the dynamical stability). The hyper parameters of the GFM have been tuned using scalable hyperparameter optimization (HPO) algorithms implemented in the software DeepHyper (https://github.com/deephyper/deephyper). The pre-training of each HPO trial was performed using distributed data parallelism (DDP) to scale the training across 128 compute nodes of the exascale OLCF supercomputer Frontier. Each HPO trial was trained only for 10 epochs and an early stopping was performed to avoid wasting significant computational resources on GNN architectures that were clearly underperforming. For each HPO trial, the 'omnistat' tool developed by (AMD Research - Advanced Micro Device) was used to measure the total energy consumption in kWh. The ensemble of GFMs was obtained by selecting the fifteen best performing HPO trials. Four models have been selected for their clear advantage in accuracy, and these are the GFMs with IDs 229, 156, 147, 260. Additional eleven models have been selected based on judicious balance between accuracy and energy consumption needed for training, and these are the GFMs with IDs 165, 78, 137, 1, 175, 171, 181, 67, 179, 167, 351. Each selected GFM of the ensemble was continued to cumulate a total of at most 30 epochs. In some cases, the total number of epochs actually performed was les than 30 due to two combined factors: (1) the size of the GFM (i.e., the number of model parameters to train) and (2) the total wall-clock time for which the computational resources could be allocated on OLCF-Frontier. The "Ensemble_of_models" directory contains 15 sub-directories named as follows: - gfm_0.229 - gfm_0.156 - gfm_0.147 - gfm_0.260 - gfm_0.165 - gfm_0.78 - gfm_0.137 - gfm_0.1 - gfm_0.175 - gfm_0.171 - gfm_0.181 - gfm_0.67 - gfm_0.179 - gfm_0.167 - gfm_0.351 Each one of these sub-directories refers to one of the fifteen HPO trials that have been selected to continue the pre-training with at most 30 epochs. With each sub-directory associated with a specific HPO trial, the following files can be found: - config.json: file for argument parsing to develop and train an HydraGNN architecture - gfm_0.ID_epoch_N.pk: file with model parameters for HPO ID trial after N epochs of training The ensemble of fifteen GFM architectures was used for (1) ensemble averaging to stabilize the predictions of energy and atomic forces after pre-training for post-processing analysis and (2) ensemble uncertainty quantification (UQ). The code used to develop, pre-train, and load the pre-trained models for post-processing analysis is available on the ORNL-GitHub at the following link: https://github.com/ORNL/HydraGNN/tree/Predictive_GFM_2024

36 MATERIALS SCIENCE

Monitoring Sulfuric Acid and Temperature Using Raman Spectroscopy and Multivariate Chemometrics

Multivariate regression models were optimized for the quantification of sulfuric acid (H 2 SO 4 ) [0–8 M] and temperature (20 °C–80 °C) in the presence of ammonium sulfate ((NH 4 ) 2 SO 4 [0–0.6 M]) using Raman spectroscopy. Optical vibrational spectroscopy is a useful nondestructive technique for the in situ analysis of complex chemical systems notoriously difficult to monitor in situ and in real-time. Multivariate analysis, a chemometrics method, can be paired with these nondestructive optical methods for determining analyte concentration and speciation in complex solutions, such as dissociated species in polyprotic acids, e.g., H 2 SO 4 . The effect of temperature is often overlooked although it can have a major influence on speciation and the corresponding Raman spectra. Here, in this study, partial least squares regression models were optimized for the quantification of H 2 SO 4 and its two deprotonated forms as a function of temperature. Measuring bisulfate as a function of temperature is particularly challenging owing to changes in the second dissociation constant. A designed training set effectively minimized the sample set size and trained a robust predictive model with percent root mean square error of <3% for H 2 SO 4 . The practical strategy employed here was demonstrated to be effective for building chemometric models that directly account for dynamic temperatures with static samples and is shown to be amenable to flow cell analysis applications with a simple calibration transfer for process monitoring applications.

D-optimal design

Computer-aided design of stability enhanced nicotinamide cofactor biomimetics for cell-free biocatalysis

Cell-free biocatalysis (CFB) is an efficient and environmentally friendly method to synthesize molecules such as pharmaceuticals, biochemicals, and biofuels through the in vitro use of enzyme cascades. These enzymes often require redox cofactors to drive chemical reactions. Natural redox cofactors (NAD(P)H) are expensive to isolate, motivating synthetic nicotinamide cofactor biomimetics (NCBs) as a cost-effective solution. A select handful of NCBs have been identified as potential NAD(P)H alternatives with comparable or improved redox capabilities, however, they display a tendency to degrade in common buffers. In this study, a library of 132 NCB candidates is systematically generated, over 85% of which have not been characterized in the literature, to expand the diversity of currently explored NCBs. The decomposition mechanism of NCBs in phosphate is evaluated using density functional theory (DFT), revealing protonation at the nicotinamide C5 position as a reporter of cofactor stability. Based on this result, we trained a linear regression model on DFT calculated descriptors to predict NCB stability in phosphate buffer, achieving mean absolute error (MAE) and root mean squared error (RMSE) values within computational accuracy. Analysis of key atomic descriptors and qualitative trends in our dataset informed the design of novel NCB candidates we propose with optimized stability. This work enables researchers to predict the relative stability of NCBs before synthesis, thereby streamlining the process to make CFB more affordable and viable at industry scales.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Data for Propagation Method and Planting Density Influence Canopy Developmental Transition and Biomass Productivity in Miscanthus × giganteus

Understanding how establishment practices influence the mechanisms underlying Miscanthus × giganteus (miscanthus) productivity and canopy development is critical for optimizing management. Data was collected during the juvenile (2011–2013) and mature (2024) phases of a long-term field experiment established in Urbana, Illinois, to evaluate the effects of propagation method (plug propagation [PP] and rhizome propagation [RP]), planting density (1.0, 0.75, and 0.25 plants m⁻²), and nitrogen application (0 and 67 kg N ha⁻¹) on end-of-season biomass yield, tiller mass, tiller density, and tiller height. Linear regression models identified the dominant predictors of yield across stand ages and management regimes. Planting density, nitrogen (N) application, and propagation method significantly influenced early yield and canopy development. During the juvenile phase, biomass yield was driven by tiller density due to canopy expansion; in the mature phase, yield became driven by tiller mass. The PP plots produced higher tiller density than the RP plots, resulting in faster canopy closure and higher juvenile-phase yields. Rhizome-propagated (RP) plots produced lower tiller density, but individual tillers were 3.3–6.4 g tiller−1 heavier than PP tillers. After the canopy reached equilibrium, the PP and RP yields were similar because greater RP tiller mass compensated for its lower tiller density. Higher planting density resulted in greater yield and tiller density during the second year (2012), but this effect was absent from the third year (2013) onward. In the juvenile phase, N fertilization enhanced yield by 1.6–3.4 Mg ha−1. Initiating fertilization in 2013 on unfertilized plots produced biomass similar to that in fertilized plots, suggesting yield recovery in the mature phase. These findings revealed that establishment strategies, including propagation method and planting density, influence juvenile miscanthus canopy development and productivity, transitioning from tiller-density- to mass-dominated yields, but not mature phase productivity.

Miscanthus

Climate models project increasing precipitation in the US Southwest in 2024–2100

We analyzed the measured precipitation in the southwestern US and found that, from 1900 to 2024, precipitation decreased at an average rate of − 1.2 cm per year per 100 years. The most significant precipitation decreases occurred in the three-month period from February to April. At the same time, precipitation during the southwestern monsoon season (July, August, and September) remained relatively stable. The ensemble mean of all CMIP6 (Coupled Model Intercomparison Project phase 6) climate models, along with the regression model incorporating anthropogenic aerosol (AER) and Pacific Decadal Oscillation (PDO) as predictors, projects precipitation to increase from 2024 to 2100. This projected precipitation increase relies on the anticipated decrease in emissions of anthropogenic aerosols, which is associated with the transition from fossil fuel burning to renewable and nuclear energy sources.

54 ENVIRONMENTAL SCIENCES