Search NASA⌕ Search

SEARCH · Search NASA

Results for “random forest regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Characterization and prediction of the electromechanical wear of contact tips during wire arc additive manufacturing of 316L stainless steel

Here, this study seeks to better understand the degradation of the contact tip with respect to WAAM for a 316L wire electrode as well as explore methods of monitoring the contact tip state from process data. The contact tip, a consumable component, positions the wire and serves as the electrical contact surface between the wire electrode and the welding power supply. The wear of the contact tip was characterized in terms of material loss and material contamination for a set of tips worn to discrete levels as measured by the amount of wire fed or arc time. Geometrical characterization found a 49% increase in the bore exit area at 180 meters of wire fed. Machine learning models were developed to predict the relative bore exit area of the contact tip from arc-based process data and a random forest classifier exhibited favorable performance with a cross-validated f1-score of 0.84. The regression architecture implemented a multi-layer perceptron with the ability to predict the relative exit area with an $R^2$ score of 0.75. Key features used in the prediction include the standard deviation of the voltage and the time between shorts.

Contact tip wear↗

A framework to evaluate machine learning crystal stability predictions

The rapid adoption of machine learning in various scientific domains calls for the development of best practices and community agreed-upon benchmarking tasks and metrics. We present Matbench Discovery as an example evaluation framework for machine learning energy models, here applied as pre-filters to first-principles computed data in a high-throughput search for stable inorganic crystals. We address the disconnect between (1) thermodynamic stability and formation energy and (2) retrospective and prospective benchmarking for materials discovery. Alongside this paper, we publish a Python package to aid with future model submissions and a growing online leaderboard with adaptive user-defined weighting of various performance metrics allowing researchers to prioritize the metrics they value most. To answer the question of which machine learning methodology performs best at materials discovery, our initial release includes random forests, graph neural networks, one-shot predictors, iterative Bayesian optimizers and universal interatomic potentials. We highlight a misalignment between commonly used regression metrics and more task-relevant classification metrics for materials discovery. Accurate regressors are susceptible to unexpectedly high false-positive rates if those accurate predictions lie close to the decision boundary at 0 eV per atom above the convex hull. The benchmark results demonstrate that universal interatomic potentials have advanced sufficiently to effectively and cheaply pre-screen thermodynamic stable hypothetical materials in future expansions of high-throughput materials databases.

Riebesell, Janosh↗

Machine learning enhanced predictions of ICRF heating: Overcoming numerical limitations via data curation

In this work, we present the development of robust surrogate models for Ion Cyclotron Range of Frequencies (ICRF) and High-Harmonic Fast Wave (HHFW) heating predictions in fusion plasmas. Building upon our previous efforts to achieve real-time capable models, we identify the cause of the outliers found using TORIC in certain HHFW heating scenarios. The outliers are observed to be spurious ion Bernstein wave (IBW)-like modes caused by a wavelength control algorithm designed to address challenging scenarios with high perpendicular wavenumbers. The effect arises from the modulation in the perpendicular susceptibility, which can induce sign reversal and IBW-like propagation for scenarios featuring normalized ion Larmor radius λ i ≫ 1. We use TORIC with this algorithm disabled to generate a novel HHFW-NSTX database that is free of outliers. Surrogate models trained on this database, including Random Forest Regressor (RFR), Multi-Layer Perceptrons, and Gaussian Process Regressors (GPR), demonstrate the ability to accurately predict HHFW heating profiles, with regression scores of R 2 ∈[0.93−0.99]. Additionally we demonstrate that it is possible to generalize predictions beyond training data by the use of both RFR and GPR models, enabling the prediction of scenarios previously limited to the original model. GPR models also provide uncertainty quantification, offering insights into model confidence. This work introduces a comprehensive Verification, Validation, and Uncertainty Quantification methodology for surrogate modeling, applicable not only to ICRF heating but also to other RF heating challenges and fusion physics problems. Beyond accelerated inference, these models show effective extrapolation capabilities, providing an alternative for addressing numerical challenges.

Artificial neural networks↗

Risk-Aware Framework Development for Disruption Prediction: Alcator C-Mod and DIII-D Survival Analysis

Abstract Survival regression models can achieve longer warning times at similar receiver operating characteristic performance than previously investigated models. Survival regression models are also shown to predict the time until a disruption will occur with lower error than other predictors. Time-to-event predictions from time-series data can be obtained with a survival analysis statistical framework, and there have been many tools developed for this task which we aim to apply to disruption prediction. Using the open-source Auton-Survival package we have implemented disruption predictors with the survival regression models Cox Proportional Hazards, Deep Cox Proportional Hazards, and Deep Survival Machines. To compare with previous work, we also include predictors using a Random Forest binary classifier, and a conditional Kaplan-Meier formalism. We benchmarked the performance of these five predictors using experimental data from the Alcator C-Mod and DIII-D tokamaks by simulating alarms on each individual shot. We find that developing machine-relevant metrics to evaluate models is an important area for future work. While this study finds cases where disruptive conditions are not predicted, there are instances where the desired outcome is produced. Giving the plasma control system the expected time-to-disruption will allow it to determine the optimal actuator response in real time to minimize risk of damage to the device.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Verbal Learning and Memory Deficits across Neurological and Neuropsychiatric Disorders: Insights from an ENIGMA Mega Analysis

Deficits in memory performance have been linked to a wide range of neurological and neuropsychiatric conditions. While many studies have assessed the memory impacts of individual conditions, this study considers a broader perspective by evaluating how memory recall is differentially associated with nine common neuropsychiatric conditions using data drawn from 55 international studies, aggregating 15,883 unique participants aged 15–90. The effects of dementia, mild cognitive impairment, Parkinson’s disease, traumatic brain injury, stroke, depression, attention-deficit/hyperactivity disorder (ADHD), schizophrenia, and bipolar disorder on immediate, short-, and long-delay verbal learning and memory (VLM) scores were estimated relative to matched healthy individuals. Random forest models identified age, years of education, and site as important VLM covariates. A Bayesian harmonization approach was used to isolate and remove site effects. Regression estimated the adjusted association of each clinical group with VLM scores. Memory deficits were strongly associated with dementia and schizophrenia (p < 0.001), while neither depression nor ADHD showed consistent associations with VLM scores (p > 0.05). Differences associated with clinical conditions were larger for longer delayed recall duration items. By comparing VLM across clinical conditions, this study provides a foundation for enhanced diagnostic precision and offers new insights into disease management of comorbid disorders.

Neurosciences & Neurology↗

Southern Rockies Western Slope Agriculture: Identifying Drivers of Rangeland Production for Drought Planning on the Western Slope of the Southern Rockies

Over the last decade, the southern Rocky Mountains of the United States experienced severe and variable drought. Local ranchers and landowners have reported strain on their operations, citing decreasing forage for their cattle and a need to adjust their business models. This study identified Major Land Resource Area-48 (MLRA-48) and northwestern Colorado as the key region for analysis. NASA DEVELOP partnered with the BLM Colorado River Field Office, Colorado State University Extension, USDA Forest Service, and the National Drought Mitigation Center to address concerns regarding the efficacy of remotely sensed rangeland production platforms and identify early warning climatic indicators of drought. The study identified two key platforms, The Rangeland Productivity Monitoring Service (RPMS) and Rangeland Analysis Platform (RAP), which use NASA Landsat 5 TM, Landsat 7 ETM+, Landsat 8 OLI, and Landsat 9 OLI-2 to estimate rangeland biomass. We regressed these with in situ biomass data to validate their efficacy and found that RAP was more effective than RPMS in estimating rangeland biomass, though it presents a tendency to overestimate. Our study performed a random forest analysis, comparing monthly RAP biomass estimates to a variety of climate variables, including mean precipitation, temperature, Palmer Drought Severity Index, snow water equivalent, snow persistence from Terra MODIS, wind speed and direction, and vapor pressure deficit. We determined that vapor pressure deficit and precipitation are key indicators in predicting forage production in MLRA-48. Our climate analysis provided our partners with greater understanding of the influence of various climate variables in determining rangeland production and allows them to assist land managers in drought mitigation.

remote sensing↗

Implementation of stacked ensemble machine learning for the detection of surrogate plutonium contamination in soil via LIBS

Supervised machine learning methods have demonstrated increased utility for the quantification of lanthanide and actinide elements in atomic spectroscopy applications. This study implements laser-induced breakdown spectroscopy (LIBS) for the identification of plutonium surrogate material (CeO 2 ) in soil matrices by training supervised machine learning methods on the recorded spectral data. A bagged ensemble using Random Forest yields the highest sensitivity predictions with a detection limit of 0.015 wt.% CeO 2 . However, high precision in Ce content prediction required the use of a stacked ensemble regression, which provided the superlative Ce quantification model with an error of 0.107% and a detection limit of 0.022 wt.%. Furthermore, the high performance of the stacked ensemble demonstrates its potential to enhance the accuracy and sensitivity of nuclear contaminant detection using field-deployable spectroscopic analyzers in real-world scenarios.

47 OTHER INSTRUMENTATION↗

Mapping tree canopy cover and canopy height with L-band SAR using LiDAR data and Random Forests

Light detection and ranging (LiDAR) data can provide direct measurements of vegetation structures but are limited by the sparse spatial coverage. Polarimetric synthetic aperture radar (SAR) can perform large-scale high-resolution mapping without weather constraints but the information about vegetation and ground subsurface are mixed in the backscatter data. In this paper, we adopted the Random Forests algorithm to train an upscaling function using tree canopy cover (TCC) and canopy height model (CHM) derived from Goddard’s LiDAR, Hyperspectral and Thermal Imager (G-LiHT) data. The regression model is then applied to the L-band Uninhabited Aerial Vehicle Synthetic Aperture Radar (UAVSAR) data acquired during the 2017 Arctic-Boreal Vulnerability Experiment (ABoVE) airborne campaign to map the TCC and CHM over the Delta Junction area in interior Alaska.

Moghaddam, Mahta↗

Northern Rockies Ecological Conservation: Leveraging Earth Observations to Monitor and Predict Populations of Federally Threatened Whitebark Pine (Pinus albicaulis) across the Intermountain West

Whitebark pine (WBP; Pinus albicaulis) is an ecologically important species in North America. As a federally listed threatened species, an understanding of WBP habitat, distribution, and health is important for the natural resource managers of the National Park Service, United States Forest Service, Bureau of Land Management, Fish and Wildlife Service, and non-profit organizations such as the Whitebark Pine Ecosystem Foundation. Previous attempts to develop models of WBP habitat suitability and distribution lack confidence in their validity and integrity for these organizations. The updated models of habitat suitability and distribution developed by this study would provide managers with a capability to be employed in the conservation and future research direction for WBP. Thus, we developed a habitat suitability model of WBP at a high spatial resolution (Landsat 9 Operational Land Image-2, National Land Cover Database, NASA Shuttle Radar Topography Mission; 30m pixels) using a generalized logistic regression with an area under the curve value of 0.754. We extracted spectral reflectance signatures from overlapped ground sample points and Sentinel-2 Multispectral Instrument. The spectral signature analysis indicates WBP is separable from other tree species. We also utilized a visual validation approach and random forest (RF) modeling to separate WBP from limber pine. Through visual validation the RF classifier successfully identified 8out of 10 WBP trees gathered through ground truth points. Additionally, we achieved an overall accuracy of 91%in our confusion matrix for the distribution model using a dependent validation approach. The derived products from this study allow project partners to assess current suitable habitat and apparent health status in areas of identified WBP occurrence, providing data to aid future research regarding WBP health.

Sentinel-2↗

Geographical Insights into Suicide Mortality Through Spatial Machine Learning

Suicide mortality is a leading cause of death in the United States, with an upward trend that emphasizes its significance as a public health issue. Previous research has employed global models like ordinary least squares (OLS) regression and local models such as geographically weighted regression (GWR). While local models are useful for analyzing spatial variations in suicide mortality, they share limitations with traditional global models, particularly about their inability to handle multi-collinearity and non-linear relationships. Machine learning approaches, like random forests (RF), can address some of these limitations but often fail to account for spatial variability. This gap highlights the need for spatial ML models specifically designed to tackle suicide mortality. This research seeks to fill this void by using a geographically weighted random forest model (GWRF) to examine the associations between county-level suicide mortality in the U.S. from 2010 to 2020 and various social and environmental determinants of health. A key aspect of our methodology is disciplined feature selection, which reduces the pool of explanatory variables by about 90%. This refinement enhances the explanatory power of both global (R2 improved from 0.59 to 0.67) and local (R2 improved from 0.64 to 0.67) RF models while reducing their run times. An analysis of the importance scores for these selected features reveals that the drivers of suicide mortality vary by context. Thus, to effectively address regional disparities and inform targeted public health interventions, a holistic approach that incorporates multiple county-level characteristics is essential.

Lebakula, Viswadeep [ORNL] (ORCID:0000000152935914↗

Learning epistatic polygenic phenotypes with Boolean interactions

Detecting epistatic drivers of human phenotypes is a considerable challenge. Traditional approaches use regression to sequentially test multiplicative interaction terms involving pairs of genetic variants. For higher-order interactions and genome-wide large-scale data, this strategy is computationally intractable. Moreover, multiplicative terms used in regression modeling may not capture the form of biological interactions. Building on the Predictability, Computability, Stability (PCS) framework, we introduce the epiTree pipeline to extract higher-order interactions from genomic data using tree-based models. The epiTree pipeline first selects a set of variants derived from tissue-specific estimates of gene expression. Next, it uses iterative random forests (iRF) to search training data for candidate Boolean interactions (pairwise and higher-order). We derive significance tests for interactions, based on a stabilized likelihood ratio test, by simulating Boolean tree-structured null (no epistasis) and alternative (epistasis) distributions on hold-out test data. Finally, our pipeline computes PCS epistasis p-values that probabilisticly quantify improvement in prediction accuracy via bootstrap sampling on the test set. We validate the epiTree pipeline in two case studies using data from the UK Biobank: predicting red hair and multiple sclerosis (MS). In the case of predicting red hair, epiTree recovers known epistatic interactions surrounding MC1R and novel interactions, representing non-linearities not captured by logistic regression models. In the case of predicting MS, a more complex phenotype than red hair, epiTree rankings prioritize novel interactions surrounding HLA-DRB1 , a variant previously associated with MS in several populations. Taken together, these results highlight the potential for epiTree rankings to help reduce the design space for follow up experiments.

59 BASIC BIOLOGICAL SCIENCES↗

Southern Colorado Disasters: Using NASA Observations to Map Aspen Extent and Recovery Due to Wildfire

Quaking aspen (Populus tremuloides) is an important species for wildlife, watershed health, and ecosystem resilience across its range. Heavy ungulate browsing and factors influenced by a changing climate including seasonal temperature changes and moisture deficit have led to reduced post-fire aspen regeneration rates in southern Colorado. This project partnered with Trinchera Ranch and the Colorado State Forest Service to estimate aspen recovery after the Spring Creek Fire, which ignited in June of 2018. The Southern Colorado Disasters team utilized field measurements and satellite imagery from Landsat Operational Land Imager (OLI), Sentinel-2 MultiSpectral Instrument (MSI), and the Shuttle Radar Topography Mission (SRTM) to train and run several random forest models that detect pre- and post-fire aspen extent. Ocular sampling of over 500 points on high-resolution pre-fire and post-fire images identified percentage aspen cover in 30 x 30-meter grid cells. This process provided training data for regression models, which were able to detect aspen across the landscape for both time periods using multiple remote sensing vegetation health indices. In addition, landscape suitability for aspen regeneration was modeled to provide a guide for managers on where to monitor for aspen regeneration post-fire.

DEVELOP Project Summary↗

Southern Colorado Disasters: Using NASA Earth Observations to Map Aspen Extent and Recovery Due to Wildfire

Quaking aspen (Populus tremuloides) is an important species for wildlife, watershed health, and ecosystem resilience across its range. Heavy ungulate browsing and factors influenced by a changing climate including seasonal temperature changes and moisture deficit have led to reduced post-fire aspen regeneration rates in southern Colorado. This project partnered with Trinchera Ranch and the Colorado State Forest Service to estimate aspen recovery after the Spring Creek Fire, which ignited in June of 2018. The Southern Colorado Disasters team utilized field measurements and satellite imagery from Landsat Operational Land Imager (OLI), Sentinel-2 MultiSpectral Instrument (MSI), and the Shuttle Radar Topography Mission (SRTM) to train and run several random forest models that detect pre- and post-fire aspen extent. Ocular sampling of over 500 points on high-resolution pre-fire and post-fire images identified percentage aspen cover in 30 x 30-meter grid cells. This process provided training data for regression models, which were able to detect aspen across the landscape for both time periods using multiple remote sensing vegetation health indices. In addition, landscape suitability for aspen regeneration was modeled to provide a guide for managers on where to monitor for aspen regeneration post-fire.

DEVELOP Technical Paper↗

Forest Biomass Mapping From Lidar and Radar Synergies

The use of lidar and radar instruments to measure forest structure attributes such as height and biomass at global scales is being considered for a future Earth Observation satellite mission, DESDynI (Deformation, Ecosystem Structure, and Dynamics of Ice). Large footprint lidar makes a direct measurement of the heights of scatterers in the illuminated footprint and can yield accurate information about the vertical profile of the canopy within lidar footprint samples. Synthetic Aperture Radar (SAR) is known to sense the canopy volume, especially at longer wavelengths and provides image data. Methods for biomass mapping by a combination of lidar sampling and radar mapping need to be developed. In this study, several issues in this respect were investigated using aircraft borne lidar and SAR data in Howland, Maine, USA. The stepwise regression selected the height indices rh50 and rh75 of the Laser Vegetation Imaging Sensor (LVIS) data for predicting field measured biomass with a R(exp 2) of 0.71 and RMSE of 31.33 Mg/ha. The above-ground biomass map generated from this regression model was considered to represent the true biomass of the area and used as a reference map since no better biomass map exists for the area. Random samples were taken from the biomass map and the correlation between the sampled biomass and co-located SAR signature was studied. The best models were used to extend the biomass from lidar samples into all forested areas in the study area, which mimics a procedure that could be used for the future DESDYnI Mission. It was found that depending on the data types used (quad-pol or dual-pol) the SAR data can predict the lidar biomass samples with R2 of 0.63-0.71, RMSE of 32.0-28.2 Mg/ha up to biomass levels of 200-250 Mg/ha. The mean biomass of the study area calculated from the biomass maps generated by lidar- SAR synergy 63 was within 10% of the reference biomass map derived from LVIS data. The results from this study are preliminary, but do show the potential of the combined use of lidar samples and radar imagery for forest biomass mapping. Various issues regarding lidar/radar data synergies for biomass mapping are discussed in the paper.

Sun, Guoqing↗

Data driven investigation to understand the influence of total solids on biological biogas upgrading

In situ biogas upgrading achieves CO 2 conversion to CH 4 via hydrogenotrophic methanogenesis; however, gas-liquid mass transfer constraints limit the upgrading performance. Recognizing that optimization studies often underrepresent the effects of total solids (TS) and organic loading rate (OLR), this study undertook a holistic, statistics driven assessment of operating conditions for in situ H 2 assisted biogas upgrading, centering the analysis on TS and OLR. A dataset of 31 studies was compiled and comprised 99 observations. A rigorous analytical framework was employed, combining data standardization, fixed- and random-effects (REML) weighted regressions with cluster-robust errors, stratified analyses, and machine learning. Mixed-effects meta regression indicated that TS was the main factor explaining differences of methane fraction (CH 4 %) when considering the between studies heterogeneity. Focusing on a near-stoichiometric subset (H 2 /CO 2 ≈ 4:1), TS remained significant. Stratified results showed a stronger negative relationship between TS and CH 4 % in UASB reactors than in CSTRs, with a negative effect under mesophilic conditions and no significant effect under thermophilic conditions. A Random Forest model corroborated the statistical findings, consistently ranking H 2 /CO 2 ratio, OLR, TS, and hydrogen injection rate (HIR) as the most influential predictors. These findings delineate trends across increasing TS levels, particularly between 1% and 10%, and provide preliminary insights for TS above 15% in in situ biogas upgrading. They further provide insights for the influence of TS by reactor type and temperature, thereby advancing the evidence base for implementing biological CO 2 conversion to CH 4 in practice.

In situ biogas upgrading↗

Addressing bias in bagging and boosting regression models

As artificial intelligence (AI) becomes widespread, there is increasing attention on investigating bias in machine learning (ML) models. Previous research concentrated on classification problems, with little emphasis on regression models. This paper presents an easy-to-apply and effective methodology for mitigating bias in bagging and boosting regression models, that is also applicable to any model trained through minimizing a differentiable loss function. Our methodology measures bias rigorously and extends the ML model's loss function with a regularization term to penalize high correlations between model errors and protected attributes. We applied our approach to three popular tree-based ensemble models: a random forest model (RF), a gradient-boosted model (GBT), and an extreme gradient boosting model (XGBoost). We implemented our methodology on a case study for predicting road-level traffic volume, where RF, GBT, and XGBoost models were shown to have high accuracy. Despite high accuracy, the ML models were shown to perform poorly on roads in minority-populated areas. Our bias mitigation approach reduced minority-related bias by over 50%.

97 MATHEMATICS AND COMPUTING↗

Detection of Diversion in a Realistic Heat Pipe Microreactor Using Supervised Machine Learning

Microreactors (MRs) pose new challenges for international safeguards. Here, their small size and mass reproducibility make them ideal for deployment in greater numbers and in remote locations, making the job of safeguards inspectors more challenging. Machine learning (ML) is currently being applied to many fields to augment human performance and increase automation; in particular, ML could be used to provide insight for international inspectors to help detect the diversion of nuclear fuel from MR cores. Four ML model types (k-nearest neighbors, decision tree, random forest, and histogram-based gradient boosted ensemble) were trained on integrated flux and critical control drum angle data generated with Serpent 2 for a realistic heat pipe MR design, achieving nearly 100% binary classification accuracy of nominal and diversion core configurations by the end of 1 full power year for three of the four model types. Regression model variants were also trained, using the same input data, for predicting the number of fuel pins diverted. Root-mean-square errors below 5% of the total number of fuel pins were achieved by the 1 full power year mark for all models.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Nearest-Neighbor Machine Learning Feature Selection for Interpretation of Microbial Molecular Signatures from Isotope Ratio Mass Spectrometry Data

Mass spectrometry (MS) promises to be a powerful tool for potential biosignature detection during astrobiological missions on ocean worlds in our solar system. Accurate and generalizable machine learning methods could enhance science return on investment by predicting seawater chemistry and classifying isotopic biosignatures, either as a signature consistent with microbial life (biotic) or as a novelty (unclassified/unique). However, machine learning models are likely to be complex and involve interactions between MS features, making biosignatures difficult to interpret. Feature selection methods provide biological and chemical context that help interpret the mechanisms of machine learning models, but these methods also need the ability to detect complex interactions. Previously, we developed a machine learning feature selection algorithm called nearest-neighbor projected distance regression (NPDR) that has the ability to identify important model features that involve complex interactions and automatically reduce correlation and the dimensionality in a high-dimensional variable space. The standard distance metrics used in NPDR – Manhattan and Euclidean – assume the multivariate data are isotropic, which is often violated in real data due to differences in the covariance between variables. Thus, we extend NPDR to include a random forest distance, and other anisotropic distance metrics, for computing nearest neighbors. We also augment the isotope-ratio MS data with time-series features from the raw MS signal to improve biotic classification. We test NPDR on our novel experimental ocean world seawater analog MS data. We measure isotope fractionations of volatile CO 2 that could be measured in exospheres or plumes. Samples include baseline abiotic conditions using a range of possible seawater chemistry consistent with Europa and Enceladus, and biotic samples that include microbes in these seawaters. We use penalized NPDR with random forest proximity to identify interpretable microbial molecular signatures. We compare features with random forest importance, and we train a classifier that discriminates between biotic and abiotic samples with high accuracy. These ML-trained ocean-world analog MS data could be used to assist in identifying biosignatures during future missions.

geochemistry↗