Search NASASearch

SEARCH · Search NASA

Results for “Random Forest”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Newton-Raphson AC Power Flow Convergence Based on Deep Learning Initialization and Homotopy Continuation

Power flow forms the basis of many power system studies. With the increased penetration of renewable energy, grid planners tend to perform multiple power flow simulations under various operating conditions and not just selected snapshots at peak or light load conditions. Getting a converged AC power flow (ACPF) case remains a significant challenge for grid planners especially in large power grid networks. This paper proposes a two-stage approach to improve Newton-Raphson ACPF convergence and was applied to a 6102 bus Electric Reliability Council of Texas (ERCOT) system. The first stage utilizes a deep learning-based initializer with data re-training. Here a deep neural network (DNN) initializer is developed to provide better initial voltage magnitude and angle guesses to aid in power flow convergence. This is because Newton-Raphson ACPF is quite sensitive to the initial conditions and bad initialization could lead to divergence. The DNN initializer includes a data re-training framework that improves the initializer's performance when faced with limited training data. The DNN initializer successfully solved 3,285 cases out of 3,899 non-converging dispatch and performed better than random forest and DC power flow initialization methods. ACPF cases not solved in this first stage are then passed through a hot-starting algorithm based on homotopy continuation with switched shunt control. The hot-starting algorithm successfully converged 416 cases out of the remaining 614 non-converging ACPF dispatch. In conclusion, the combined two-stage approach achieved a 94.9% success rate, by converging a total of 3,701 cases out of the initial 3,899 unsolved cases.

Deep learning

Machine Learning-Driven Reliability Estimation of PV Inverters Considering Alert-Ambient Variability

Weather-induced spatio-temporal degradation limits outdoor PV inverter lifetime and reliability, necessitating advanced data analysis. This study employs a top-down, data-driven approach utilizing multiple machine learning (ML) algorithms to estimate inverter reliability in a 1.4 MW PV power plant, considering factors such as irradiance, humidity, temperature, time of day, and weather conditions. An extensive alert dataset from 17 identical inverters, including alert types, propagation, and frequency, reveals significant correlations with environmental factors and inverter output power, enabling the construction of a performance reliability model. Dual-stage supervised-ML models are evaluated for accuracy, with the ‘classification-regression’ model by an artificial neural network (ANN) tested on the averaged “Alert-Ambient” dataset, which is outperformed by ‘clustering-regression’ models using random forest (RF) and K-Nearest Neighbors (KNN) on individual inverter datasets. K-means clustering applies principal component analysis to reduce dimensions, achieving improved accuracy beyond the 80% achieved by ANN on the averaged dataset. Second-stage regression estimates inverter reliability with a mean square error of 0.0195 on the averaged dataset and as low as 0.002 on individual inverter datasets using RF. Furthermore, these findings highlight the method's suitability for estimating PV inverter output reliability under ambient conditions, essential for digital twin development and related applications.

14 SOLAR ENERGY

Switchgrass Steroidal Saponins Reduce Fungal Disease but Decrease Yeast Fermentation Yield

Increasing the production of bioproducts from lignocellulosic feedstocks requires improvement in both field production and biorefinery efficiency. When plant traits arise that improve field production but decrease biofuel yield, these trade-offs can represent challenges in the entire production process. To examine trade-offs between field and production traits, we examined factors underlying switchgrass resistance to fungal rust pathogens in field conditions and factors that impede yeast fermentation in the lab using repeated measurements on a switchgrass genetic diversity panel. We found that the same switchgrass genotypes that showed high fungal pathogen resistance also showed recalcitrance to yeast fermentation. These switchgrass genotypes were mostly from the Atlantic genetic group, which had high levels of specialized metabolites of the saponin class. Among 1589 metabolites identified through metabolomics, we found that saponins were among the most likely to explain variation in both rust infection and fermentation yield using random forest feature selection, and that only four of these were sufficient to explain 57.9% of the variation in rust susceptibility. Through follow-up testing in recalcitrant biomass, we found that the bacterium Zymomonas mobilis does not suffer the same inhibition as the yeast Saccharomyces cerevisiae, and that the addition of ergosterol (thought to be the fungal cellular target of saponin inhibition) rescues yeast fermentation. Several lines of evidence point to a central role for saponins as key metabolites protecting switchgrass from fungal pathogens and interfering with yeast fermentation, underscoring an ongoing need for collaboration between plant breeders and biofuel production scientists.

VanWallendael, Acer [North Carolina State Universi

Evaluation of GlassNet for physics-informed machine learning of glass stability and glass-forming ability

Glassy materials form the basis of many modern applications, including nuclear waste immobilization, touch-screen displays, and optical fibers, and also hold great potential for future medical and environmental applications. However, their structural complexity and large composition space make design and optimization challenging for certain applications. Of particular importance for glass processing and design is an estimate of a given composition's glass-forming ability (GFA). However, there remain many open questions regarding the underlying physical mechanisms of glass formation, especially in oxide glasses. It is apparent that a proxy for GFA would be highly useful in glass processing and design, but identifying such a surrogate property has proven itself to be difficult. While glass stability (GS) parameters have historically been used as a GFA surrogate, recent research has demonstrated that most of these parameters are not accurate predictors of the GFA of oxide glasses. Here, in this work, we explore the application of an open-source pre-trained neural network model, GlassNet, that can predict the characteristic temperatures necessary to compute GS with reasonable performance and assess the feasibility of using these physics-informed machine learning (PIML)-predicted GS parameters to estimate GFA. In doing so, we track the uncertainties at each step of the computation—from the original ML prediction errors to the compounding of errors during GS estimation, and finally to the final estimation of GFA. While GlassNet exhibits reasonable accuracy on all individual properties, we observe a large compounding of error in the combination of these individual predictions for the PIML prediction of GS, finding that random forest models offer similar accuracy to GlassNet. We also break down the performance of GlassNet on different glass families and find that the error in GS prediction is correlated with the error in crystallization peak temperature prediction. Lastly, we utilize this finding to assess the relationship between top-performing GS parameters and GFA for two ternary glass systems: sodium borosilicate and sodium iron phosphate glasses. We conclude that to obtain true ML predictive capability of GFA, significantly more data needs to be collected.

36 MATERIALS SCIENCE

Plastics Environmental Risk Calculator

SF-24-074 The Plastics Environmental Risk Calculator (PERC) was developed in Microsoft Excel and estimates the environmental distribution and lifetime of new biobased and conventional plastics from commonly measured properties of the plastics. A Random Forest regression model embedded in the calculator calculates plastic degradation rates and lifetimes as a proxy for environmental risk. Model default assumptions may be overwritten by the user.

Beckman, Kevin [Argonne National Laboratory (ANL),

Hydroboost

HydroBoost is the most realistic revenue optimization tool for the hybridization of hydropower and battery energy storage systems to date. The innovative representation of how operators actually schedule hydropower in practice results in more realistic predictions of revenue and operations. Unlike other optimization tools, HydroBoost generates forecast energy prices with uncertainty to use in the optimization. This allows HydroBoost to give users a range of potential revenue with an upper bound using the perfect foresight pricing and a lower bound using a naive persistence forecast model. Additional forecast can be generated and used in the optimization, such as additive models, random forest, and neural networks to give further insight into potential revenue. HydroBoost has been designed to be applicable for both run-of-river and reservoir storage sites. The primary focus is on the day-ahead market and requires year-long data with an hour time-step. All time-series input and constraints are contained in an Excel worksheet for convince. The user will run the forecasting generation first with a Python script to give the optimization model the necessary requirements. Next the optimization is ran using Julia and results are generated and stored into a directory as csv files. HydroBoost includes an additional module to generate figures based on the results of the optimization simulation. The results help analyze the results and users to draw insights into how the hydro and battery systems are operated and the revenue each is producing. Additionally, the difference between the perfect foresight model and models that include forecast can easily be inspected.

Phillips, TylerB. [Idaho National Laboratory (INL)

Decayheatml

This code is designed to predict and analyze the decay heat generated in molten salt reactors (MSRs) using a hybrid approach that combines machine learning and segmented polynomial fitting. The accurate prediction of decay heat is essential for reactor safety and the optimization of spent fuel storage. The code operates through several key components: 1) Data Architecture: It incorporates a modular data architecture that handles various MSR-specific operational parameters such as power density, humidity content, and air ingress. These parameters are sampled using Sobol sequences to ensure comprehensive coverage of operational uncertainties. 2) Machine Learning Framework: The code employs a diverse set of machine learning models, including polynomial regression, decision trees, random forests, gradient boosting, support vector regression, k-nearest neighbors, multi-layer perceptrons, and symbolic regression. These models are trained to predict decay heat over a wide temporal range, from immediate shutdown up to 10,000 years. 3) Region-Optimized Training: The temporal domain is divided into multiple regions, each modeled separately to capture distinct decay heat characteristics across different time scales. This approach significantly improves the accuracy and interpretability of predictions. 4) Segmented Polynomial Interpretation (SPI): The SPI method translates machine learning predictions into piecewise polynomial equations. These equations are physically interpretable and can be directly integrated into existing engineering workflows and safety analyses. 5) Front-End Interfaces: The code includes both a Jupyter notebook interface for research development and a Streamlit web application for operational deployment. These interfaces allow users to interactively explore decay heat predictions, adjust operational parameters, and visualize results in real-time. 6) Applications: The framework supports various applications, including safety system validation and spent fuel container optimization. It enables real-time evaluation of worst-case decay heat scenarios, informing the design of passive safety systems and optimizing container designs for long-term storage. Overall, this code provides a robust, accurate, and user-friendly tool for predicting decay heat in MSRs, enhancing reactor safety, and optimizing spent fuel management.

Retamales, Mauricio Eduardo Tano [Idaho National L

Calibration and Rapid-Adoption Forecasting Techniques

CRAFT (Calibration and Rapid-Adoption Forecasting Techniques) CRAFT is a Python-based project for processing, analyzing, and modeling atmospheric or environmental data. It uses machine learning techniques, specifically Random Forest Regression, to create emulators for various environmental variables such as gross primary production and soil water content. It then uses these emulators to robustly test the parameter space of mechanistic models to provide posterior estimations of the free parameters.

Robins, Zachary

VHClass

The code is used to predict the taxonomic source of an antibody heavy chain sequence. The code assigns a binary label to the input set of sequences - camelid or human. This prediction is generated using a random-forest based classification algorithm which is the backbone of the code. A complementary code splits the antibody sequence into antibody features - framework regions and CDR regions.

Davis, Anastasiia

Dependence of Convective Cloud Microphysical Properties on Environmental Conditions during the TRACER and ESCAPE Field Campaigns: A Synergistic Approach of Observations, Machine Learning and Parcel Models

The sensitivity of convective clouds to aerosols and their interactions with environment, combined with limited observational constraints in parameterizations, introduces significant uncertainties in atmospheric models. Here, this study investigates the dependence of convective cloud microphysical properties on environmental conditions using a synergistic approach that combines unique observations from the TRACER and ESCAPE field campaigns, machine learning techniques, and parcel model simulations with a super-droplet microphysics scheme. A random forest algorithm identifies in-situ vertical velocity (w), temperature (T), and surface fine-mode aerosol mass concentration as the three most important environmental conditions influencing cloud properties including liquid water content (LWC), number concentration for particles with D max < 50 μm (N c ,<50), 50 μm ≤ D max ≤ 3000 μm (N c,50–3000 ), and droplet effective diameter (D e ). Results show that LWC, N c,<50 , and N c,50–3000 significantly increase with w in updrafts. Across w bins, as T decreases, LWC, D e , and N c,50–3000 increase, while N c,<50 decreases, which are closely linked to the distance above cloud bases. Warmer cloud bases yield higher LWC, greater N c,50–3000 , and smaller N c,<50 , while polluted environments produce greater N c,<50 . Parcel model simulations successfully replicate these observed dependencies. The simulation results indicate that warmer cloud bases enhance condensation generating larger droplets, and differences in droplet sizes are then amplified through collision-coalescence, resulting in a greater N c,50–3000 . Polluted conditions result in a greater N c,<50 primarily due to enhanced cloud condensation nuclei activation despite increased collision-coalescence rates compared to pristine conditions. This study provides observed quantitative patterns characterizing cloud microphysical properties as a function of key environmental parameters, offering valuable constraints for improving physics parameterizations and numerical models.

54 ENVIRONMENTAL SCIENCES

Dependence of Deep Convective Cell Properties on Meteorological and Aerosol Conditions during TRACER

Deep convective cells significantly influence Earth’s energy balance and water cycle. However, their accurate representation in numerical models remains challenging due to their small spatiotemporal scales and limited observational constraints. This study examines over ∼400 deep convective cells near Houston, observed by a dual-polarization C-band radar during the Tracking Aerosol Convection Interactions Experiment (TRACER) intensive observation period (June–September 2022). Cells are categorized by lifetime into short-lived (<40 min), intermediate-lived (40–80 min), and long-lived (80+ min) groups. Long-lived cells were broader (∼13.2 km at 2–4-km height) and deeper (∼11.4 km) than short-lived cells (∼6.4-km width, ∼7.31-km height). Using random forest (RF) modeling and correlation analyses, precipitable water vapor (PWV), 2–6-km lapse rate, 0–8-km bulk shear, and fine aerosol mass concentration (Mass_f) are identified as key predictors of cell lifetime. Higher PWV is associated with significantly longer convective cell lifetimes compared to the low-PWV group, particularly within low 2–6-km temperature lapse rate (LR_26km), moderate-to-higher 0–8-km bulk shear (BS_08km), and low-to-moderate Mass_f environments. RF analysis also identifies low-level (0–2 km) equivalent potential temperature, PWV, Mass_f, and surface latent heat flux as key predictors for cell width and height. Short-lived cells have higher aerosol number concentrations (500–1000-nm size range), linked to onshore wind conditions and marine aerosols; however, their low concentration suggests the sensitivity may reflect associated meteorological regimes rather than a direct aerosol effect. Long-lived cells have higher concentrations of organic and sulfate aerosols, while short-lived cells exhibit higher black carbon concentrations. These results highlight the intricate dependence of convective cell lifetimes and structure on environmental moisture, thermodynamics, wind shear, and aerosol characteristics.

54 ENVIRONMENTAL SCIENCES

Human limits in machine learning: prediction of potato yield and disease using soil microbiome data

Abstract Background The preservation of soil health is a critical challenge in the 21st century due to its significant impact on agriculture, human health, and biodiversity. We provide one of the first comprehensive investigations into the predictive potential of machine learning models for understanding the connections between soil and biological phenotypes. We investigate an integrative framework performing accurate machine learning-based prediction of plant performance from biological, chemical, and physical properties of the soil via two models: random forest and Bayesian neural network. Results Prediction improves when we add environmental features, such as soil properties and microbial density, along with microbiome data. Different preprocessing strategies show that human decisions significantly impact predictive performance. We show that the naive total sum scaling normalization that is commonly used in microbiome research is one of the optimal strategies to maximize predictive power. Also, we find that accurately defined labels are more important than normalization, taxonomic level, or model characteristics. ML performance is limited when humans can’t classify samples accurately. Lastly, we provide domain scientists via a full model selection decision tree to identify the human choices that optimize model prediction power. Conclusions Our study highlights the importance of incorporating diverse environmental features and careful data preprocessing in enhancing the predictive power of machine learning models for soil and biological phenotype connections. This approach can significantly contribute to advancing agricultural practices and soil health management.

Aghdam, Rosa

Implementation of stacked ensemble machine learning for the detection of surrogate plutonium contamination in soil via LIBS

Supervised machine learning methods have demonstrated increased utility for the quantification of lanthanide and actinide elements in atomic spectroscopy applications. This study implements laser-induced breakdown spectroscopy (LIBS) for the identification of plutonium surrogate material (CeO 2 ) in soil matrices by training supervised machine learning methods on the recorded spectral data. A bagged ensemble using Random Forest yields the highest sensitivity predictions with a detection limit of 0.015 wt.% CeO 2 . However, high precision in Ce content prediction required the use of a stacked ensemble regression, which provided the superlative Ce quantification model with an error of 0.107% and a detection limit of 0.022 wt.%. Furthermore, the high performance of the stacked ensemble demonstrates its potential to enhance the accuracy and sensitivity of nuclear contaminant detection using field-deployable spectroscopic analyzers in real-world scenarios.

47 OTHER INSTRUMENTATION

Machine learning identifies novel signatures of antifungal drug resistance in Saccharomycotina yeasts

Antifungal drug resistance is a major challenge in fungal infection management. Numerous genomic changes are known to contribute to acquired drug resistance in clinical isolates of specific pathogens, but whether they broadly explain natural resistance across entire lineages is unknown. We leveraged genomic, ecological, and phenotypic trait data from naturally sampled strains from nearly all known species in subphylum Saccharomycotina to examine the evolution of resistance to eight antifungal drugs. The phylogenetic distribution of drug resistance varied by drug; fluconazole resistance was widespread, while 5-fluorocytosine resistance was rare, except in Lipomycetales. A random forest algorithm trained on genomic data predicted drug-resistant yeasts with 54–75% accuracy. Fluconazole resistance was consistently predicted with the highest accuracy (75.2%). Furthermore, fluconazole resistance prediction accuracy was similar between models trained on genome-wide variation in the presence and number of InterPro protein annotations across Saccharomycotina (75.2%) and those trained on amino acid sequence alignment data of Erg11, a protein known to be involved in fluconazole resistance (74.3-74.9%). Interestingly, the top Erg11 residues for predicting fluconazole resistance across Saccharomycotina do not overlap with, are not spatially close to, and are less conserved than those previously linked to resistance in clinical isolates of Candida albicans. In silico deep mutational scanning of the C. albicans Erg11 protein reveals that amino acid variants implicated in clinical cases of resistance are almost universally destabilizing while variants in our most informative residues are energetically more neutral, explaining why the latter are much more common than the former in natural populations. Importantly, previous experimental analyses of C. albicans Erg11 have shown that amino acid variation in our most informative residues, despite having never been directly implicated in clinical cases, can directly contribute to resistance. Our results suggest that studies of natural resistance in yeast species never encountered in the clinic will yield a fuller understanding of antifungal drug resistance.

Harrison, Marie-Claire [Vanderbilt Univ., Nashvill

Urbanization and malaria have a contextual relationship in endemic areas: A temporal and spatial study in Ghana

In West Africa, malaria is one of the leading causes of disease-induced deaths. Existing studies indicate that as urbanization increases, there is corresponding decrease in malaria prevalence. However, in malaria-endemic areas, the prevalence in some rural areas is sometimes lower than in some peri-urban and urban areas. Therefore, the relationship between the degree of urbanization, the impact of living in urban areas, and the prevalence of malaria remains unclear. This study explores this association in Ghana, using epidemiological data at the district level (2015–2018) and data on health, hygiene, and education. We applied a multilevel model and time series decomposition to understand the epidemiological pattern of malaria in Ghana. Then we classified the districts of Ghana into rural, peri-urban, and urban areas using administratively defined urbanization, total built areas, and built intensity. We converted the prevalence time series into cross-sectional data for each district by extracting features from the data. To predict the determinant most impacting according to the degree of urbanization, we used a cluster-specific random forest. We find that prevalence is impacted by seasonality, but the trend of the seasonal signature is not noticeable in urban and peri-urban areas. While urban districts have a slightly lower prevalence, there are still pockets with higher rates within these regions. These areas of high prevalence are linked to proximity to water bodies and waterways, but the rise in these same variables is not associated with the increase of prevalence in peri-urban areas. The increase in nightlight reflectance in rural areas is associated with an increased prevalence. We conclude that urbanization is not the main factor driving the decline in malaria. However, the data indicate that understanding and managing malaria prevalence in urbanization will necessitate a focus on these contextual factors. Finally, we design an interactive tool, ’malDecision’ that allows data-supported decision-making.

60 APPLIED LIFE SCIENCES

Forward variable selection enables fast and accurate dynamic system identification with Karhunen-Loève decomposed Gaussian processes

A promising approach for scalable Gaussian processes (GPs) is the Karhunen-Loève (KL) decomposition, in which the GP kernel is represented by a set of basis functions which are the eigenfunctions of the kernel operator. Such decomposed kernels have the potential to be very fast, and do not depend on the selection of a reduced set of inducing points. However KL decompositions lead to high dimensionality, and variable selection thus becomes paramount. This paper reports a new method of forward variable selection, enabled by the ordered nature of the basis functions in the KL expansion of the Bayesian Smoothing Spline ANOVA kernel (BSS-ANOVA), coupled with fast Gibbs sampling in a fully Bayesian approach. It quickly and effectively limits the number of terms, yielding a method with competitive accuracies, training and inference times for tabular datasets of low feature set dimensionality. Theoretical computational complexities are O ( N P 2 ) in training and O ( P ) per point in inference, where N is the number of instances and P the number of expansion terms. The inference speed and accuracy makes the method especially useful for dynamic systems identification, by modeling the dynamics in the tangent space as a static problem, then integrating the learned dynamics using a high-order scheme. The methods are demonstrated on two dynamic datasets: a ‘Susceptible, Infected, Recovered’ (SIR) toy problem, along with the experimental ‘Cascaded Tanks’ benchmark dataset. Comparisons on the static prediction of time derivatives are made with a random forest (RF), a residual neural network (ResNet), and the Orthogonal Additive Kernel (OAK) inducing points scalable GP, while for the timeseries prediction comparisons are made with LSTM and GRU recurrent neural networks (RNNs) along with the SINDy package.

Hayes, Kyle

FIRM image analysis: A machine learning workflow for quantifying extracellular matrix components from electron microscopy images

The extracellular matrix (ECM) is a complex network of biomolecules that plays an integral role in the structure, processes, and signaling mechanisms of cells and tissues. Identifying and quantifying changes in these matrix components provides insight into the mechanisms behind specific tissue remodeling processes; however, quantifying these changes is challenging due to difficult imaging conditions, complexity of the ECM, and the subtlety of these changes. Current imaging techniques allow us to visualize these critical remodeling events and developments in image analysis have employed a combination of analysis software and machine learning techniques to improve the efficiency and accuracy with which features are measured. Although image analysis has seen much improvement in recent years, there has been no technique developed to address ambiguity in feature edges in electron microscopy images. Presented here is a new machine learning-based workflow for the analysis of microscopy images named FIRM (Feature Identification from Raw Microscopy) that uses a random forest classifier to identify ECM features of interest and generate binary segmentation masks for quantification with ImageJ-FIJI. FIRM performed with an F1 score of 0.794 and greater than 80% accuracy for number and size of features detected. FIRM had similar deviation from the ground truth in the number of identified fibrils, fibril size, and size distributions when compared to human analyses. The results suggest that FIRM performs as well as manual analysis and requires a fraction of the time. This analysis technique is more efficient, eliminates user bias, and can be easily optimized to identify a variety of features, making it useful for any discipline requiring image analysis.

Science & Technology - Other Topics

Machine learning of factors for improving oyster hatchery production

Oyster aquaculture and restoration in the Chesapeake Bay are vital, yet hatcheries frequently struggle with inconsistent larval growth and sudden mass mortality events. Unpredictable disruptions in larval production cause large economic losses, represent a perceived risk to growers, and impede industry expansion. To better understand associations between production yield and its potential predictors, we applied machine learning (random forest, and neural network) and statistical (generalized additive model) models to a comprehensive dataset of environmental, water quality, and operational parameters from a Maryland oyster hatchery, aiming to identify key yield predictors and develop a robust forecasting tool. We used recursive Boruta algorithm for variable selection, pinpointing critical predictors, and employed cross-validation to fine-tune model settings. Shapley value analysis offered crucial insights into model interpretations, highlighting week number, Normalized Difference Vegetation Index, salinity, turbidity, and fecundity as primary drivers of yield variability. For low-yield cases, salinity-related variables were particularly important. Our findings provide an early warning system for potential production downturns, empowering hatchery operators to make data-driven decisions for optimizing water conditions, feeding schedules, and broodstock management. By boosting predictability and efficiency, this research directly supports economic stability of the oyster industry and ecological health of the Chesapeake Bay.

Vishwakarma, Srishti [Oak Ridge National Laborator