Search NASA⌕ Search

SEARCH · Search NASA

Results for “Regression analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Modeling of the metal–insulator transition temperature in alio-valently doped VO 2 through symbolic regression

The correlated semiconductor vanadium dioxide (VO 2 ) exhibits an insulator–metal transition (IMT) near room temperature, which is of interest in various device applications. Precise IMT temperature control is crucial to determine the use cases across technologies such as thermochromic windows, actuators for robots or neuronal oscillators. Doping the cation or anion sites can modulate the IMT by several tens of degrees and control hysteresis. However, modeling the effects of control parameters (e.g., doping concentration, type of dopants) is challenging due to complex experimental procedures and limited data, hindering the use of traditional data-driven machine learning approaches. Symbolic regression (SR) can bridge this gap by identifying nonlinear expressions connecting key input parameters to target properties, even with small data sets. In this work, we develop SR models to capture the IMT trends in VO 2 influenced by different dopant parameters. Using experimental data from the literature, our study reveals a dual nature of the IMT temperature with varying tungsten (W) doping concentrations. The symbolic model captures data trends and accounts for experimental variability, providing a complementary approach to first-principles calculations. Our feature-driven analysis across a broader class of dopants informs selectivity and provides qualitative insights into tuning phase transition properties valuable for neuromorphic computing and thermochromic windows.

36 MATERIALS SCIENCE↗

Predicting Drug Effects from High-dimensional Asymmetric Drug Data Sets using Graph Neural Networks: A Comprehensive Analysis of Multi-target Drug Effect Prediction

Graph neural networks (GNNs) have emerged as one of the most effective Machine learning (ML) techniques for drug effect prediction from drug molecular graphs. Despite having immense potential, GNN models lack performance when using data sets that contain high dimensional asymmetrically co-occurrent drug effects as targets with complex correlations between them. Training individual learning models for each drug effect and incorporating every prediction result for a wide spectrum of drug effects is beyond practicality. Such an implication provides a testbed to address this challenge as multi-target prediction problems, aiming to predict all drug effects at a time. We develop standard and hybrid graph neural networks (GNNs)to perform two separate tasks that are multi-regression for continuous values and multi-label classification for categorical values contained in our data sets. Since this step makes the target data even more sparse and introduces asymmetric label co-occurrence, the learning of multi-label classification models becomes difficult and heavily impacts the GNN's performance. To address these challenges, we propose a new data oversampling technique to improve multi-label classification performances on all the given imbalanced molecular graph data sets. Using the technique, we improve the data imbalance ratio of the drug effects better than before while protecting the data set's integrity. Finally, we evaluate multi-label classification performance using the best-performant hybrid GNN model on all the oversampled data sets obtained from the proposed oversampling technique. These results outperform those of other ML models including GNN models when they are trained on the original data sets or oversampled data sets using MLSMOTE (a well-known oversampling technique) in all evaluation metrics precision, recall, and F1 score by a significant margin.

Bose, Avishek [ORNL]↗

On-chip probabilistic inference for charged-particle tracking at the sensor edge

Modern scientific instruments operate under increasingly extreme constraints on bandwidth, latency, and power. Inference at the sensor edge determines experimental data collection efficiency by deciding which information to save for further analysis. Particle tracking detectors at the Large Hadron Collider exemplify this challenge: pixelated silicon sensors generate rich spatiotemporal ionization patterns, yet most of this information is discarded due to data-rate limitations. Concurrently, advancements in co-design tools provide rapid turn-around for incorporating machine learning into application-specific integrated circuits, motivating designs for particle detectors with new integrated technologies. We demonstrate that neural networks embedded in the front-end electronics can infer charged-particle kinematic parameters from a single silicon layer. We regress hit positions and incident angles with calibrated uncertainties, while satisfying stringent constraints on numerical precision, latency, and silicon area. Our results establish a path toward probabilistic inference directly at the edge, opening new opportunities for intelligent sensing in high-rate scientific instruments.

Das, Arghya Ranjan [Purdue U.] (ORCID:000000018451↗

Uncertainty Quantification and Sensitivity Analysis of Non-Nuclear Advanced Controls Testbed Reactor Mockup

The research presented in this report describes our progress in applying stochastic methods and uncertainty quantification, parametric study, and variance-based sensitivity analysis (also known as Sobol sensitivity analysis) to a full-core model of a nuclear thermal propulsion (NTP) system simulated with Griffin, with the goal of developing a reduced order (surrogate) model which can be rapidly sampled while perturbing multiple input parameters. In this NTP system, reactivity and power feedback affect the rotation of control drums, which are controlled by a hybrid proportional, integral and derivative (PID) controller, actuated by the power demand and reactivity feedback from the numerical model. This model uses reactor kinetic feedback (mean generation time and $\beta$ from a transient Griffin simulation executed with the improved quasi-static method to provide the kinetic parameters) as inputs to functions which control the CD rotation angle. Using a number of stochastic method approaches, we developed a dual purpose training-surrogate model of the NTP system using polynomial regression. The trained model can be rapidly sampled while simultaneously perturbing various input parameters of the model, such as coefficients on the PID control, or temperature (directly affect the neutron cross section). The surrogate model delivers accurate results orders-of-magnitude faster (minutes, not days) than the base model. Once the base model has been trained, distributions of the uncertain parameters can be changed at will to investigate the effects of perturbing multiple inputs and their effect on the output. For example, coefficients used in the PID control system may vary due to some physical interference, or there may be uncertainty in the temperature of the neutron cross sections in various regions of the reactor. A distribution can be placed on these parameters and operational boundaries can be determined. The goal of this work is to support development of an advanced control system to operate CDs in a functioning NTP system.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Quantifying Uncertainty in HPC Job Queue Time Predictions

High Performance Computing (HPC) has developed at an unprecedented pace in recent decades. This growth has demanded corresponding development in the area of HPC Operational Data Analytics (ODA), which encompasses a wide range of data analysis techniques, ML/AI efforts, tools, and visualizations. Published studies in ODA offer a variety of practical ways to inform HPC users, administrators, procurement managers, and other stakeholders. Uncertainty analysis, however, is rare in the related published literature. For instance, we identify only 1 out of 14 existing studies focused on job queue time prediction that investigates the uncertainty aspect of their proposed predictions. We recognize the utmost importance uncertainty quantification can have in such predictive analytics solutions, with consequences in how users interpret information they receive, and attempt to bridge this gap. With the goal of improving access to such insights, we develop a process for determining upper and lower bounds of the predicted queue times of a regression model at a specified confidence level. Our current research is focused on the uncertainty in predicting job queue times, yet our approach may be employed in predicting other metrics.

HPC↗

Application of Partial Least Squares Approaches to Pyroprocessing ER Data

Multivariate approaches show promise for application to process monitoring for safeguards of pyroprocessing. Past MPACT work explored the application of Principal Component Analysis (PCA) to detect off-normal conditions in pyroprocessing electrorefiner (ER) data from in the Hot Fuel Examination Facility (HFEF) at Idaho National Laboratory (INL) known as the Scalable Pyrochemical Recycling testbed (SPyRe) ER. PCA, however, does not consider the output variables. In FY24, multivariate analysis was extended from PCA to Partial Least Squares (PLS) analysis. PLS maximizes the variance between both the input signals and output variables. In the case of this work, PLS was applied in two different manners: Predictive PLS and Discriminant PLS. Predictive PLS maximizes the covariance between the process variables of the ER and the measured U concentration from in-situ voltammetry. Discriminant PLS maximizes the covariance between the process variables and a set of training process “states” such as known off-normal conditions. By projecting into the latent variable space in PLS, the process variables can be regressed onto the outputs and predictions can be made for new data sets. In this work, by applying predictive PLS, a penalized non-linear PLS approach was able to make predictions of concentration based on test and training data and detect when operations were off-normal. However, the predictive PLS does not classify the signals to which off-normal operations are attributable. Discriminant PLS can be used to classify off-normal operations but is inadequate to properly classify specific off-normal classes like power supply faults when the Discriminant PLS model is only specifically trained to detect that off-normal class. When all faults are trained against the observation data, all three operational classes are accurately classified and distinguished. Thus, future application of latent variable techniques should not select any given method, but should use a mixture of PCA, Predictive PLS, and Discriminant PLS.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Predicting Initial Trans-Membrane Pressure for Optimized Operations in UF Unit Using Random Forest

With the growing scarcity of freshwater, innovative process design mechanisms like Reverse Osmosis (RO) are increasingly gaining attention among water treatment utilities to address the rising demand. Ensuring reliable water production necessitates efficient resource utilization, minimizing downtime in (ultra-filtration) UF systems. Recent advancements in machine learning (ML) have enabled the development of accurate data-driven models for Model Predictive Control (MPC), often requiring minimal prior knowledge of underlying physical processes. In this study, we present predictive regression models based on Random Forest (RF) and Auto-Regressive (AR) approaches to forecast the initial Trans-Membrane Pressure (TMP) for each filtration cycle in data generated by Direct Potable Reuse (DPR) systems. The proposed RF-based model demonstrates superior performance compared to baseline methods, including historical mean, Last Observation Carried Forward (LOCF), and naïve AR models, across various forecasting horizons in terms of root mean square error (RMSE) metric. To evaluate how different classes of process variables contribute to TMP dynamics over time, we examine the feature importance of independent covariates across multiple forecast horizons. This analysis provides insight into the temporal relevance of operational and sensor-derived features, guiding control and monitoring strategies. Additionally, the impact of hyperparameter tuning on TMP prediction performance is studied for both direct and recursive RF modelling approaches across increasing forecast horizons. Accurate prediction of initial TMP is critical for optimizing RO operations, as it enables the development of robust modelling frameworks by accurately estimating membrane fouling trends, thereby enhancing process efficiency and long-term reliability. The demonstrated efficacy of the RF-based approach highlights its potential as a tool for real-time decision-making in water treatment systems, paving the way for advanced process optimization and sustainable water resource management.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)↗

Prediction of electric and magnetic fields from spectral data using machine learning algorithms for Doppler-free saturation spectroscopy diagnostics

The prediction of electric and magnetic field amplitudes from atomic spectral data is critical for plasma control in fusion devices such as tokamaks. Conventional approaches that rely on physics-based models are computationally expensive and unsuitable for real-time applications. In this work, we develop and benchmark three machine learning algorithms—simulation-based inference (SBI), fully connected neural networks (FCNN), and histogram-based gradient boosting regression (GBR-Hist)—to infer field intensities directly from Doppler-free saturation spectroscopy (DFSS) spectra. Synthetic datasets of spectra were generated using the EZSSS code and evaluated both with and without added Poisson noise to mimic experimental conditions. We find that SBI achieves the highest accuracy and robustness, FCNN provides a strong balance of accuracy and computational efficiency for real-time applications, and GBR-Hist offers the fastest inference but is more sensitive to noise. Furthermore, these results demonstrate the potential of machine learning to accelerate DFSS analysis and enhance its utility for plasma diagnostics and control.

Doppler-free saturation spectroscopy↗

Climate change and federal aid disbursements after Hurricane Harvey: an extreme event attribution analysis

The role climate change plays in increasing the burden placed on governments and insurers to pay for recovery has not been extensively explored and is the focus of this study. This study examines the impacts of climate change attributed flooding on federal disaster aid disbursement in Harris County, Texas following Hurricane Harvey in 2017. Our approach uses flood models to estimate the amount of flood damages attributable and not attributable to climate change under two climate change attribution scenarios from peer reviewed studies: 20% and 38% increases in rainfall associated with the hurricane due to climate change. These estimates are combined with census tract-level disbursement data for FEMA’s National Flood Insurance Program (NFIP) and the Individual Assistance (IA) part of the Individuals and Households Program. We employ spatial lag regression models with direct and spatial spillover effects to analyze the relationship between a tract’s flood damages—both attributed and not attributed to climate change—and federal disaster aid. We find that both types of flood damage shape federal aid disbursements, but that climate change attributed damages tend to have larger effect sizes (elasticities) especially for IA. Specifically, for a 1% increase in additional climate change attributed damages per household in a census tract (under the 20% scenario), expected NFIP levels in that census tract are 0.26% higher and IA levels are 0.3% higher. Implications center on federal funding in an era of climate change.

FEMA↗

Verbal Learning and Memory Deficits across Neurological and Neuropsychiatric Disorders: Insights from an ENIGMA Mega Analysis

Deficits in memory performance have been linked to a wide range of neurological and neuropsychiatric conditions. While many studies have assessed the memory impacts of individual conditions, this study considers a broader perspective by evaluating how memory recall is differentially associated with nine common neuropsychiatric conditions using data drawn from 55 international studies, aggregating 15,883 unique participants aged 15–90. The effects of dementia, mild cognitive impairment, Parkinson’s disease, traumatic brain injury, stroke, depression, attention-deficit/hyperactivity disorder (ADHD), schizophrenia, and bipolar disorder on immediate, short-, and long-delay verbal learning and memory (VLM) scores were estimated relative to matched healthy individuals. Random forest models identified age, years of education, and site as important VLM covariates. A Bayesian harmonization approach was used to isolate and remove site effects. Regression estimated the adjusted association of each clinical group with VLM scores. Memory deficits were strongly associated with dementia and schizophrenia (p < 0.001), while neither depression nor ADHD showed consistent associations with VLM scores (p > 0.05). Differences associated with clinical conditions were larger for longer delayed recall duration items. By comparing VLM across clinical conditions, this study provides a foundation for enhanced diagnostic precision and offers new insights into disease management of comorbid disorders.

Neurosciences & Neurology↗

Sex Differences in Odds of Brain Metastasis and Outcomes by Brain Metastasis Status after Advanced Melanoma Diagnosis

Sex differences in cancer are well-established. However, less is known about sex differences in diagnosis of brain metastasis and outcomes among patients with advanced melanoma. Using a United States nationwide electronic health record-derived de-identified database, we evaluated patients diagnosed with advanced melanoma from 1 January 2011–30 July 2022 who received an oncologist-defined rule-based first line of therapy (n = 7969, 33% female according to EHR, 35% w/documentation of brain metastases). The odds of documented brain metastasis diagnosis were calculated using multivariable logistic regression adjusted for age, practice type, diagnosis period (pre/post-2017), ECOG performance status, anatomic site of melanoma, group stage, documentation of non-brain metastases prior to first-line of treatment, and BRAF positive status. Real-world overall survival (rwOS) and progression-free survival (rwPFS) starting from first-line initiation were assessed by sex, accounting for brain metastasis diagnosis as a time-varying covariate using the Cox proportional hazards model, with the same adjustments as the logistic model, excluding group stage, while also adjusting for race, socioeconomic status, and insurance status. Adjusted analysis revealed males with advanced melanoma were 22% more likely to receive a brain metastasis diagnosis compared to females (adjusted odds ratio [aOR]: 1.22, 95% confidence interval [CI]: 1.09, 1.36). Males with brain metastases had worse rwOS (aHR: 1.15, 95% CI: 1.04, 1.28) but not worse rwPFS (adjusted hazard ratio [aHR]: 1.04, 95% CI: 0.95, 1.14) following first-line treatment initiation. Among patients with advanced melanoma who were not diagnosed with brain metastases, survival was not different by sex (rwOS aHR: 1.06 [95% CI: 0.97, 1.16], rwPFS aHR: 1.02 [95% CI: 0.94, 1.1]). This study showed that males had greater odds of brain metastasis and, among those with brain metastasis, poorer rwOS compared to females, while there were no sex differences in clinical outcomes for those with advanced melanoma without brain metastasis.

60 APPLIED LIFE SCIENCES↗

Advancing the quantification of aerosol-cloud interactions with the CALIPSO-CloudSat-Aqua/MODIS record

Aerosol-cloud-precipitation interactions are assessed over the non-polar ocean using more than 11 years of combined Aqua-MODIS, CALIPSO-CALIOP, and CloudSat products. The analysis first shows the benefit of incorporating vertically resolved aerosol extinction coefficient (σext) in aerosol-cloud interactions (ACI) assessments, demonstrating that: σext vertically collocated with the cloud layer () correlates best with cloud droplet number concentration (Nd), column-integrated aerosol optical depth (AOD) cannot explain the Nd variability in the extratropics, and the S-shape of the AOD-Nd relationship reported in previous studies is not replicated when using instead of AOD, with a Nd- linearity more consistent with in-situ studies over the ocean. ACI metric, estimated as the log-scale regression between CALIOP and MODIS Nd reveals that the eastern Pacific is the region with the strongest ACI, followed by the Southern Ocean. The susceptibility of clouds to changes in their liquid water path (LWP) and frequency of precipitation followed a 2-step calculation by combining the Nd- regression (ACI) with the regression between these macrophysical variables and Nd. LWP susceptibility is negative (LWP decreases with aerosol loading) and statistically significant over the eastern Pacific, eastern Atlantic, and extratropics. In contrast, vast areas of the tropical and subtropical ocean feature negligible changes in LWP with aerosol. Precipitation frequency susceptibility is negative, but the values are only significant over the coastal eastern Pacific and Atlantic. The findings suggest that previous modeling assessments relying on AOD may need to be revisited by taking advantage of the synergy between passive and active sensors.

Li, Zhujun↗

Investigating the Determinants of Household Capabilities Burden During Power Outages: The Case of Winter Storm Uri

Existing research primarily uses census data to identify the vulnerability of communities to hazards. These vulnerability indices provide aggregated data and are not hazard-specific nor well-validated with post-event data. In contrast, our study uses household survey data (n=1065) to understand which Texan households suffered the greatest loss of their capabilities due to power outages and other utility service disruptions during Winter Storm Uri. Inspired by the Capabilities Approach, our measures of burden include the number of household capability types disrupted during the outages (e.g., cooking, heating, refrigeration), the severity of impact for each disrupted capability, and the additional time and financial costs of coping with these disruptions. We perform a clustering analysis, and find two distinct groups in our data, consisting of ‘lesser burden' and ‘heavier burden' households. Results indicate that the households experiencing the heaviest capabilities burden were most likely to experience longer power outages and the loss of water services. They were also more likely to have a Hispanic-Latino household member, lack access to a generator, live in a rented home, have larger households with more young children, fewer adults over 65, lower household incomes, been impacted by the COVID-19 pandemic, and more family characteristics that made life harder. We also fit a logistic regression model to assess the role of outage, household, and community characteristics in predicting differences in capabilities burden. Our results offer insights into enumerating the consequences of utility service disruptions on households, which can inform more targeted and equitable resilience strategies.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Machine learning in materials research: Developments over the last decade and challenges for the future

The number of studies that apply machine learning (ML) to materials science has been growing at a rate of approximately 1.67 times per year over the past decade. In this review, I examine this growth in various contexts. First, I present an analysis of the most commonly used tools (software, databases, materials science methods, and ML methods) used within papers that apply ML to materials science. The analysis demonstrates that despite the growth of deep learning techniques, the use of classical machine learning is still dominant as a whole. It also demonstrates how new research can effectively build upon past research, particular in the domain of ML models trained on density functional theory calculation data. Next, I present the progression of best scores as a function of time on the matbench materials science benchmark for formation enthalpy prediction. In particular, a dramatic improvement of 7 times reduction in error is obtained when progressing from feature-based methods that use conventional ML (random forest, support vector regression, etc.) to the use of graph neural network techniques. Finally, I provide views on future challenges and opportunities, focusing on data size and complexity, extrapolation, interpretation, access, and relevance.

36 MATERIALS SCIENCE↗

Semi-empirical model for Henry’s law constant of noble gases in molten salts

Henry’s law constant, which describes the proportionality of dissolved gas to partial pressure of free gas in liquid–gas equilibrium systems, can also be applied to mass transport applications. In this work, we investigated an approach for determining the solubility of noble gases in a molten salt liquid utilizing the equilibrium concept of Henry’s gas constant. Henry’s gas constant is described as a mathematical function dependent on the van der Waals radius of the noble gas and the temperature of the molten salt. The alteration in Gibbs free energy encompasses contributions from both surface and volume energies. Enthalpy and entropy are deduced from these surface and volume energies in the Gibbs free energy formulation. A comparative analysis was conducted between the conventional method and our proposed model. Moreover, useful chemical properties can be determined from examination of surface and volume energies. Our findings provide an accurate and general theory of Gibbs free energy that can be validated experimentally based on the model proposed herein. This work unifies the prediction of Henry gas constant and subsequently the entropy and enthalpy calculation for noble gases in a molten salt solution to a single functional form using van der Waals radius of the gas and temperature of the system. This functional form is then used to perform a multiple regression method to find two parameters corresponding to the surface energy and volume energy. These two parameters are consistent between all combinations of noble gas and molten salt.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Data-Driven Kinetic Reaction Networks for Separation Chemistry

Understanding complex, multistep chemical reactions at the molecular level is a major challenge whose solution would greatly benefit the design and optimization of numerous chemical processes. The separation of rare-earth (4f) and actinide (5f) elements is an example where improving our chemical understanding is important for designing and optimizing new chemistries, even with a limited number of observations. Here, in this work, we leverage data-driven artificial intelligence and machine-learning approaches to develop kinetic reaction networks that describe the liquid–liquid extraction mechanism of uranium using N,N-di-2-ethylhexyl-isobutyramide (DEHiBA). Specifically, we compare and contrast the properties of two classes of models: (1) purely data-driven models that are regularized using chemistry-agnostic, L1 regression and (2) chemistry-informed models that are regularized using relative reaction energies provided by quantum mechanical calculations. We observe that purely data-driven models are unbiased, simple, and accurate in their predictions of experimental measurements when provided with sufficient data but are difficult to fully constrain and interpret. In contrast, chemistry-informed models exhibit significantly improved chemical interpretability and consistency, providing a detailed description of the separation process while achieving high accuracy through ensemble averaging. Overall, the dominant species predicted to be extracted into the organic phase is UO 2 (NO 3 ) 2 (DEHiBA) 2 , agreeing with experimental slope analysis, thermodynamic modeling, EXAFS, and crystal structures. This work demonstrates that leveraging the fundamental structure of the problem can lead to efficient learning schemes that provide both accurate predictions and chemical insights at a low computational cost.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Systems Analysis of Biomass and Coal Co-firing Power Plants with Deep Carbon Capture Toward Net-zero Emissions

Achieving a net-zero emission economy in the United States requires integrating diverse low-carbon and negative-emission technologies into the existing fossil fuel-dominant power fleet. Potential technologies from the low-carbon portfolio include renewable power, fossil power with carbon capture and storage (CCS), bioenergy with CCS (BECCS), and direct air capture (DAC). Renewable power is a clean energy source but has to pair with costly battery storage to provide dispatchable electricity. Fossil power with CCS offers dispatchable electricity yet still relies on DAC to offset residual emissions, even when deploying deep CCS with more than 90% CO2 capture. Coal-biomass co-firing with CCS, a subset of BECCS, is a reliable energy production technology that can be retrofitted from existing electricity generation units (EGUs). Power plant retrofit maximizes the use of the current U.S. coal power fleet without the need for large-scale deployment of new renewable power, battery storage, or DAC. Retrofitting coal-biomass co-firing with deep CCS in EGUs is a promising option, but not a universal solution. Biomass co-firing at a power plant introduces economic challenges and indirectly poses pressure on land and water resources. Meanwhile, retrofitting deep CCS affects plant efficiency and raises electricity generation costs. Overall, the technical feasibility and economic viability of plant retrofits vary across EGUs, as they are contingent upon the regional availability of biomass, unit-specific characteristics, site-specific fuel supply costs, and adjacent CO2 storage potential. Government incentives like 45Q can improve the retrofit viability, though the impact requires further quantification. A comprehensive analysis at the unit level is essential to address the question regarding the fate of the U.S. coal-fired electricity generation fleet toward the net-zero emission goal. This study conducts a systematic techno-economic-environmental assessment of EGUs to identify the viability of biomass co-firing and deep CCS retrofits in the U.S. coal-fired power fleet. Specifically, it characterizes the techno-economic performance of deep carbon capture, estimates life cycle greenhouse gas (GHG) emissions, and conducts a fleet-level assessment on retrofit viability. The key objectives are (1) to estimate the unit-specific performance and retrofitted cost under various biomass co-firing levels and CO2 capture rates; (2) to determine the possibility of reaching net-zero emission at the fleet level; (3) to quantify the cumulative capacities that are suitable for plant retrofits under current and future biomass supply scenarios; and (4) to improve the understanding of policy impacts on such retrofits to help the power sector’s transition to a net-zero economy. Techno-economic Model of Deep Carbon Capture. This study develops the performance and economic models for Monoethanolamine-based post-combustion CO2 capture at 95–99% capture rates. The process is simulated in Aspen Plus, analyzing the performance of carbon capture technology by varying the plant sizes, solvent lean loading, CO2 concentrations, and flue gas inlet temperature. Based on the key inputs and output parameters of CO2 capture, a reduced-order performance model of deep carbon capture is formulated. In addition, an engineering-economic model integrating the performance metrics is developed to estimate the capital as well as operation and maintenance (O&M) costs. Capital cost estimations follow the framework of the Integrated Environmental Control Model (IECM) and incorporate data regressions from three technical reports by IECM, the National Energy Technology Laboratory (NETL), and the National Renewable Energy Laboratory. The O&M cost estimation utilizes the actual inventory consumption rate and labor requirements. Both performance and cost models are embedded into IECM v13.0-beta, a fossil-fuel power plant modeling tool. Life Cycle Assessment of Power Plants. This study estimates the GHG emissions of power plants through life cycle assessment (LCA). The LCA scope includes fuel supply, combustion-based power generation, and CO2 transport and storage. The fuel-based life cycle module is designed following the framework of the NETL Unit Process Library and CO2U LCA Guidance Toolkit. The module is then incorporated into IECM v13.0-beta. The process-based LCA is applied to estimate the GHG emissions of coal and biomass supply, coal- and coal-biomass co-firing power plant operation, as well as CO2 pipeline transport and geographical sequestration. An uncertainty analysis is conducted to quantify the variability and uncertainty associated with the LCA using the Latin Hypercube Sampling (LHS) method. Fleet-level Assessment. This study evaluates the technical and economic feasibility of selected coal-fired EGUs, examines the role of tax credits in retrofit viability, and assesses the competitiveness of retrofitted units against other low-carbon options. Unit screening identifies EGUs for the study, focusing on new, efficient baseload units with air pollution controls. The power plant databases are then established to organize unit-specific information on performance and operating conditions from the relevant public databases. Biomass for co-firing retrofits is selected based on home and neighboring county availability, ensuring sustained operation with at least a 5% co-firing level. The CO2 storage site is determined by state-level storage potential, with ArcGIS Pro and NETL CO2 Saline Storage Cost Model used to identify the optimal balance between the nearest transport distances and affordable storage costs. The latest IECM v13.0-beta is then employed to configure and evaluate the eligible EGUs with or without the deployment of deep CCS and biomass co-firing. A supply curve is established to illustrate the cumulative installed capacity suitable for retrofits at different cost levels. A sensitivity analysis on tax credits for carbon sequestration is performed. Finally, a unit-level cost comparison is conducted among retrofitted plants, renewable power with battery storage, and abated fossil fuels with DAC. Expected Results. This study evaluates the technical, economic, and environmental metrics of each EGU across an array of CO2 capture rates and biomass co-firing level scenarios. Unit-level comparisons will identify critical factors influencing technical performance. The supply curves with and without tax incentives will provide insights into the impact of tax credits on biomass co-firing and CCS deployment. The cost comparisons with renewables and DAC-retrofit will assess the competitiveness of the retrofitted units. Life cycle emissions from each unit will be assessed to identify the scenarios under which net-zero emissions can be achieved. These analyses are expected to determine the total coal-fired capacity suitable for serving as a low-carbon energy source with or without tax incentives. The study results are novel in identifying optimal unit-specific strategies for producing carbon-neutral power, whether through retrofitting EGUs with deep CCS, biomass co-firing, DAC, or installing renewable power with battery. The findings will provide insight into nationwide efforts to ensure reliable, affordable, and low-carbon electricity. It also will inform investment decisions and policies in the deployment of deep carbon capture and negative emission technologies for a net-zero energy future.

Biomass Co-firing↗

Deep learning model for fast, science-based forecasting of fluid migration along faults in geologic carbon storage scenarios

Effective long-term geologic storage depends on robust site selection and credible, science-based forecasting of subsurface behavior to ensure storage integrity. For this work, we develop a deep learning–based reduced-order model (ROM) to quantify potential carbon dioxide (CO₂) and brine migration through geological faults. The ROM combines a Transformer model for binary classification and a Stacked Ensemble for regression, trained on a comprehensive dataset generated from 1400 physics-based reservoir simulations. Key geologic and operational parameters—including fault geometry, reservoir structure, and injection conditions—were systematically varied to capture a wide range of fluid migration scenarios. The ROM accurately predicts the onset of migration, cumulative migration volumes of both CO₂ and brine, and associated migration rates, as compared to an independent set of validation simulations, while significantly reducing computational cost compared to traditional simulation methods. Model performance was evaluated across diverse fault configurations, revealing that shallow reservoir geometry and fault angle are among the most influential factors governing migration behavior. Sensitivity analysis using SHapley Additive exPlanations (SHAP) provided interpretability, revealing distinct patterns in how geological and operational features drive transient versus cumulative migration outcomes. The ROM’s ability to rapidly simulate fault migration scenarios enables efficient sensitivity analyses, scenario evaluations, and decision support for site selection and monitoring design. This approach enhances the safety, scalability, and long-term operational performance of geologic carbon storage (GCS) systems by providing a robust, interpretable tool for predicting subsurface fluid migration and assessing fault-related migration potential.

42 ENGINEERING↗