Search NASA⌕ Search

SEARCH · Search NASA

Results for “Random Forest”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Integrating Crack Detection and Pipe Shape Optimization for Enhanced Sewage System Durability

Crack detection in underground reinforced concrete pipes has been essential in determining the state of stormwater infrastructure. Detection models have been implemented for detecting cracks and other defects in pipes using CCTV footage for stormwater drainage systems. In addition, Finite element models have been used to determine optimum shapes and pipe thickness for different boundary conditions such as header pipes in power plants. The concept of shape optimization emerges as a crucial factor in power plant design and operation, with the potential to maximize performance while minimizing the use of materials. Shape optimization not only enhances efficiency but also contributes to reducing the environmental footprint. This paper discusses the integration of both topics by using the cracks detected in underground pipes as boundary conditions for shape optimization of the pipes. A machine learning model has been developed which uses limited data for training and outlines the location of detected cracks. A shape optimization methodology is proposed in which ANSYS modules are used to analyze fluid flow and then optimize the shape of the pipe. The crack detection model developed has been applied to a crack detected in lab setting and machine learning model used has an accuracy of 98% using a random forest algorithm.

20 FOSSIL-FUELED POWER PLANTS↗

Mapping Rare Earths and Toxics in E-Waste via Hyperspectral Imaging and Machine Learning

Electronic waste (e-waste) presents a mounting challenge to environmental sustainability due to its complex composition, which includes high-value rare earth elements, hazardous organic compounds, and non-recyclable plastics. Accurate and scalable material classification is essential for enabling efficient resource recovery and safe recycling practices. This study introduces a confidence-aware classification pipeline that combines mid-infrared hyperspectral imaging (HSI), spectral angle mapping (SAM), and iterative machine learning to perform pixel-level material identification across e-waste devices. A curated spectral library encompassing artificial materials (e.g., plastic iron oxide, galvanized metals), minerals (e.g., allanite, hematite), and organic compounds (e.g., benzanthracene, toluene) was used to generate pseudo-labels, each assigned a confidence score based on SAM-derived spectral similarity. High-confidence samples from seven consumer electronics—digital cameras, keyboards, laptop fans, modems, motherboards, TV remotes, and speakers—were iteratively expanded and classified using models such as Support Vector Machine (SVM), Random Forest, Gradient Boosting Classifier, Partial Least Squares Discriminant Analysis (PLSDA) and Logistic Regression. The best-performing classifiers achieved macro F1 scores approaching 1.0. Results revealed widespread plastic content (dominated by plastic iron oxide), the presence of rare earth-bearing minerals like cerium-containing allanite, and pervasive detection of hazardous organics such as benzanthracene. Principal Component Analysis (PCA) visualizations and confusion matrices confirmed high separability and robust classification performance. This methodology enables precise, non-destructive, and scalable classification of heterogeneous e-waste streams. It supports automated, hazard-aware sorting in recycling workflows, facilitating selective recovery of critical materials and compliance with circular economy goals. The confidence-aware framework provides a foundation for real-time deployment in industrial settings, offering significant implications for smart e-recycling infrastructure and policy-driven material stewardship.

Circular economy↗

Machine Learning for Predicting Multipactor Susceptibility in Planar RF Structures

Multipactor discharge is a persistent challenge in high-power microwave (HPM) and accelerator systems, where secondary electron avalanches can cause heating, vacuum degradation, and failure. This work presents the first supervised machine learning (ML) framework for multipactor prediction, trained on high-fidelity 3D Particle-in-Cell (PIC) simulation data in planar geometries. The model maps operational, geometric, and material-dependent secondary electron yield (SEY) parameters to the time-averaged electron growth rate, enabling rapid reconstruction of susceptibility charts. Among the models evaluated, tree-based ensemble methods such as Random Forest and Extra Trees demonstrate superior generalization to unseen materials compared to neural networks such as multilayer perceptron (MLP). Performance metrics, including Intersection over Union (IoU), Structural Similarity Index Measure (SSIM), and Pearson correlation, show close agreement with simulation benchmarks. Principal Component Analysis attributes generalization limits to material feature-space disjointedness.

43 PARTICLE ACCELERATORS↗

Virtual Growth of SRF Materials: A Machine Learning Approach to Predict the Crystalline Structural Ordering in Nb Surface Oxides

Niobium's native surface oxide affects SRF cavity and superconducting qubit performance, motivating interest in controlling its crystalline structure. We combine a literature-derived machine-learning analysis with temperature-dependent XRD to study crystalline ordering in Nb2O5. Random Forest models, trained on 74 processing conditions from 17 papers and validated by leave-one-group-out cross-validation, predicted broad crystallinity outcomes well (balanced accuracy 0.809), but struggled with specific polymorph identity (0.577). Annealing temperature was the dominant predictor across all targets; oxygen partial pressure showed negligible importance, reflecting narrow literature coverage rather than physical irrelevance. Temperature-dependent XRD on anodized and H2O2-treated Niobium showed structural evolution consistent with the machine learning predictions. Our model and overall approach provide a data-driven framework for identifying and optimizing conditions that promote crystallization in initially amorphous oxides. This framework can guide the selection of growth and post-annealing conditions for Nb surfaces by narrowing the experimental parameter space, thereby reducing trial-and-error efforts in developing oxide structures relevant to SRF applications.

Tilkin, Anthony [Unlisted, US, IL; Fermilab]↗

Transitioning from Simulation to Reality: Applying Chatter Detection Models to Real-World Machining Data

Chatter, a self-excited vibration phenomenon, is a critical challenge in high-speed machining operations, affecting tool life, product surface quality, and overall process efficiency. While machine learning models trained on simulated data have shown promise in detecting chatter, their real-world applicability remains uncertain due to discrepancies between simulated and actual machining environments. The primary goal of this study is to bridge the gap between simulation-based machine learning models and real-world applications by developing and validating a Random Forest-based chatter detection system. This research focuses on improving manufacturing efficiency through reliable chatter detection by integrating Operational Modal Analysis (OMA), Receptance Coupling Substructure Analysis (RCSA), and Transfer Learning (TL). The study applies a Random Forest classification model trained on over 140,000 simulated machining datasets, incorporating techniques like Operational Modal Analysis (OMA), Receptance Coupling Substructure Analysis (RCSA), and Transfer Learning (TL) to adapt the model for real-world operational data. The model is validated against 1600 real-world machining datasets, achieving an accuracy of 86.1%, with strong precision and recall scores. The results demonstrate the model’s robustness and potential for practical implementation in industrial settings, highlighting challenges such as sensor noise and variability in machining conditions. This work advances the use of predictive analytics in machining processes, offering a data-driven solution to improve manufacturing efficiency through more reliable chatter detection.

42 ENGINEERING↗

Global Declines in Human-Driven Mangrove Loss

Global mangrove loss has been attributed primarily to human activity. Anthropogenic loss hotspots across Southeast Asia and around the world have characterized the ecosystem as highly threatened, though natural processes such as erosion can also play a significant role in forest vulnerability. However, the extent of human and natural threats has not been fully quantified at the global scale. Here, using a Random Forest-based analysis of over one million Landsat images, we present the first 30-meter resolution global maps of the drivers of mangrove loss from 2000-2016, capturing both human-driven and natural stressors. We estimate that 62% of global losses between 2000-2016 resulted from land-use change, primarily through conversion to aquaculture and agriculture. Up to 80% of these human-driven losses occurred within six Southeast Asian nations, reflecting the regional emphasis on enhancing aquaculture for export to support economic development. Both anthropogenic and natural losses declined between 2000-2016, though slower declines in natural loss caused an increase in their relative contribution to total global loss area. We attribute the decline in anthropogenic losses to the regionally-dependent combination of increased emphasis on conservation efforts and a lack of remaining mangroves viable for conversion. While efforts to restore and protect mangroves appear to be effective over decadal time scales, the emergence of natural drivers of loss presents an immediate challenge for coastal adaptation. We anticipate that our results will inform decision making within conservation and restoration initiatives by providing a locally-relevant understanding of the causes of mangrove loss.

Coastal wetlands↗

Estimating Forest Vertical Structure from Multialtitude, Fixed-Baseline Radar Interferometric and Polarimetric Data

Parameters describing the vertical structure of forests, for example tree height, height-to-base-of-live-crown, underlying topography, and leaf area density, bear on land-surface, biogeochemical, and climate modeling efforts. Single, fixed-baseline interferometric synthetic aperture radar (INSAR) normalized cross-correlations constitute two observations from which to estimate forest vertical structure parameters: Cross-correlation amplitude and phase. Multialtitude INSAR observations increase the effective number of baselines potentially enabling the estimation of a larger set of vertical-structure parameters. Polarimetry and polarimetric interferometry can further extend the observation set. This paper describes the first acquisition of multialtitude INSAR for the purpose of estimating the parameters describing a vegetated land surface. These data were collected over ponderosa pine in central Oregon near longitude and latitude -121 37 25 and 44 29 56. The JPL interferometric TOPSAR system was flown at the standard 8-km altitude, and also at 4-km and 2-km altitudes, in a race track. A reference line including the above coordinates was maintained at 35 deg for both the north-east heading and the return southwest heading, at all altitudes. In addition to the three altitudes for interferometry, one line was flown with full zero-baseline polarimetry at the 8-km altitude. A preliminary analysis of part of the data collected suggests that they are consistent with one of two physical models describing the vegetation: 1) a single-layer, randomly oriented forest volume with a very strong ground return or 2) a multilayered randomly oriented volume; a homogeneous, single-layer model with no ground return cannot account for the multialtitude correlation amplitudes. Below the inconsistency of the data with a single-layer model is followed by analysis scenarios which include either the ground or a layered structure. The ground returns suggested by this preliminary analysis seem too strong to be plausible, but parameters describing a two-layer compare reasonably well to a field-measured probability distribution of tree heights in the area.

Treuhaft, Robert N.↗

Front Range Wildland Fires: Evaluating the Efficacy of Remote Sensing Imagery in Monitoring Forest Fuels Treatment Methods

Over the last several decades, wildfire frequency and severity in forested areas along Colorado’s Front Range have increased due to a buildup of fuels. This has led to an increase in forest treatments, as well as an increased need to evaluate the success of these treatments. Remote sensing products offer an efficient and cost-effective way to monitor forest treatments; however, not all remote sensing products and analysis techniques have been explored by Coloradan land managers. Specifically, project partners at the Colorado State Forest Service (CSFS) and the Colorado Forest Restoration Institute (CFRI) were interested in using an effective and streamlined method of mapping canopy cover to better monitor forest treatment success. To support their needs, the NASA DEVELOP Front Range Wildland Fires team explored National Agricultural Imagery Program (NAIP) imagery at different spatial resolutions and numbers of training points with NASA’s Shuttle Radar Topography Mission (SRTM) Data Elevation Model (DEM) as a predictor in addition to NAIP imagery spectral predictors. From this analysis, we created classified canopy cover rasters, and compared accuracy metrics across model iterations. We also determined that the best performing model, with an overall accuracy of 0.900 uses 2021 NAIP imagery at 2-meter resolution, 800 training points, 200 testing points, does not use topographic predictors, and reclassifies shadow pixels via a pre-selected NDVI threshold.

Remote Sensing↗

Comparison of machine learning and electrical resistivity arrays to inverse modeling for locating and characterizing subsurface targets

Here, this study evaluates the performance of multiple machine learning (ML) algorithms and electrical resistivity (ER) arrays for inversion with comparison to a conventional Gauss-Newton numerical inversion method. Four different ML models and four arrays were used for the estimation of only six variables for locating and characterizing hypothetical subsurface targets. The combination of dipole-dipole with Multilayer Perceptron Neural Network (MLP-NN) had the highest accuracy. Evaluation showed that both MLP-NN and Gauss-Newton methods performed well for estimating the matrix resistivity while target resistivity accuracy was lower, and MLP-NN produced sharper contrast at target boundaries for the field and hypothetical data. Both methods exhibited comparable target characterization performance, whereas MLP-NN had increased accuracy compared to Gauss-Newton in prediction of target width and height, which was attributed to numerical smoothing present in the Gauss-Newton approach. MLP-NN was also applied to a field dataset acquired at U.S. DOE Hanford site.

54 ENVIRONMENTAL SCIENCES↗

The impact of curation errors in the PDBBind Database on machine learning predictions of protein–protein binding affinity

The PDBBind database has been widely utilized for the computational prediction of protein–protein binding affinities. While the accuracy of the PDBBind-curated equilibrium dissociation constants (K D ) has been reported for the protein–ligand subset of the PDBBind database, the curation accuracy has not been reported for the protein–protein subset. Here, we present a detailed manual analysis for the subset of PDBBind records with PubMed Central Open Access primary publications and find that ~19% of these records had K D values that were not supported by their primary publications. The impact of these putative curation errors on the machine learning-based prediction of K D from experimental protein–protein 3D structures was evaluated and correcting the curation errors improved the Pearson correlation coefficient between measured and random forest-predicted log 10 (K D ) values by ~8 percentage points. This finding underscores the importance of dataset accuracy for computational modelling and highlights the need for more stringent curation processes when extracting information from the scientific literature.

59 BASIC BIOLOGICAL SCIENCES↗

TCAD-Machine Learning Enabled TID Compact Model Development for Commercial SiC MOSFET

We propose a TCAD (Technology Computer Aided Design)-machine learning coupled approach that combines a TCAD tool (Charon), optimization/uncertainty quantification tool (Dakota), surrogate models, and Bayesian learning capabilities. The coupling approach is used for accurate modeling and calibration of total ionizing dose (TID) induced threshold voltage (V th ) shifts in Commercial-Off-The-Shelf (COTS) semiconductor devices and to develop physics-informed TID compact models. This versatile approach is applied to model the TID effect in an exemplar COTS 3.3 kV SiC power MOSFET (Metal-Oxide-Semiconductor Field-Effect Transistor). With the Charon-Dakota coupling, we can determine key device geometry and doping values based on device physics, which are difficult to obtain or not available for COTS devices but important for TCAD simulation; additionally, we can efficiently generate thousands of simulation results in a large parameter space, which makes it possible to develop data-driven surrogate models and perform Bayesian calibration. Utilizing the full tool-coupling approach, we achieve calibrated TCAD simulation models that accurately capture the average TID-induced V th shifts behavior with total doses and V th shifts saturation at high doses as observed in experimental data. More importantly, the calibrated TCAD simulations are obtained with determined TID model parameters (e.g., hole trap density and capture cross section) values that contain well quantified uncertainties. Furthermore, we can isolate and quantify the noises that are not captured by the TCAD models but exist in the measured data due to measurements and devices variabilities. Lastly, the calibrated surrogate models are used to develop physics-informed TID compact models. The method is generalizable to other devices and/or radiation conditions with few modifications and can provide well-determined uncertainties.

COTS↗

Machine-learning based approach to examine ecological processes influencing the diversity of riverine dissolved organic matter composition

Dissolved organic matter (DOM) assemblages in freshwater rivers are formed from mixtures of simple to complex compounds that are highly variable across time and space. These mixtures largely form due to the environmental heterogeneity of river networks and the contribution of diverse allochthonous and autochthonous DOM sources. Most studies are, however, confined to local and regional scales, which precludes an understanding of how these mixtures arise at large, e.g., continental, spatial scales. The processes contributing to these mixtures are also difficult to study because of the complex interactions between various environmental factors and DOM. Here we propose the use of machine learning (ML) approaches to identify ecological processes contributing toward mixtures of DOM at a continental-scale. We related a dataset that characterized the molecular composition of DOM from river water and sediment with Fourier-transform ion cyclotron resonance mass spectrometry to explanatory physicochemical variables such as nutrient concentrations and stable water isotopes ( 2 H and 18 O). Using unsupervised ML, distinctive clusters for sediment and water samples were identified, with unique molecular compositions influenced by environmental factors like terrestrial input and microbial activity. Sediment clusters showed a higher proportion of protein-like and unclassified compounds than water clusters, while water clusters exhibited a more diversified chemical composition. We then applied a supervised ML approach, involving a two-stage use of SHapley Additive exPlanations (SHAP) values. In the first stage, SHAP values were obtained and used to identify key physicochemical variables. These parameters were employed to train models using both the default and subsequently tuned hyperparameters of the Histogram-based Gradient Boosting (HGB) algorithm. The supervised ML approach, using HGB and SHAP values, highlighted complex relationships between environmental factors and DOM diversity, in particular the existence of dams upstream, precipitation events, and other watershed characteristics were important in predicting higher chemical diversity in DOM. Our data-driven approach can now be used more generally to reveal the interplay between physical, chemical, and biological factors in determining the diversity of DOM in other ecosystems.

54 ENVIRONMENTAL SCIENCES↗

Chronic obstructive pulmonary disease among former United States Department of Energy workers: comorbidities and lung function changes

Chronic Obstructive Pulmonary Disease (COPD) is a major cause of morbidity and mortality in the United States and is frequently associated with multiple comorbidities which lead to poor COPD outcomes in the general population. However, little is known regarding COPD comorbidities in occupational cohorts whose exposure experiences could result in differences in comorbidities compared to the general population. These differences may also be important for assessing COPD outcomes such as lung function changes or decline. Therefore, the objectives of this study were to: (1) identify and describe clusters of COPD comorbidities among Department of Energy (DOE) former workers; (2) assess if the attributes of the identified clusters differ from those identified among the general population based on the published literature, and (3) identify predictors of lung function changes and decline among DOE former workers.

60 APPLIED LIFE SCIENCES↗

A Photometric Machine-Learning Method to Infer Stellar Metallicity

Following its formation, a star's metal content is one of the few factors that can significantly alter its evolution. Measurements of stellar metallicity ([Fe/H]) typically require a spectrum, but spectroscopic surveys are limited to a few x 10(exp 6) targets; photometric surveys, on the other hand, have detected > 10(exp 9) stars. I present a new machine-learning method to predict [Fe/H] from photometric colors measured by the Sloan Digital Sky Survey (SDSS). The training set consists of approx. 120,000 stars with SDSS photometry and reliable [Fe/H] measurements from the SEGUE Stellar Parameters Pipeline (SSPP). For bright stars (g' < or = 18 mag), with 4500 K < or = Teff < or = 7000 K, corresponding to those with the most reliable SSPP estimates, I find that the model predicts [Fe/H] values with a root-mean-squared-error (RMSE) of approx.0.27 dex. The RMSE from this machine-learning method is similar to the scatter in [Fe/H] measurements from low-resolution spectra..

machine learning↗

Porosity Area Fraction Analysis of Bonded Borosilicate Specimens From in-Situ Optical Imagery

Image segmentation methods are routinely used to analyze data derived in nondestructive testing scenarios. A typical problem in this area is to estimate porosity from X-ray Computed Tomography (CT) inspections. This investigation aims to understand the effectiveness of open-source image segmentation methods for porosity estimation from optical images in the context of strength analysis of bonded materials and provides a benchmark of performance against a standard commercial software package. The rank ordering of the specimens by porosity area fraction is consistent across both methods.

Porosity area fraction↗

Porosity Area Fraction Analysis of Bonded Borosilicate Specimens From in-Situ Optical Imagery

Image segmentation methods are routinely used to analyze data derived in nondestructive testing scenarios. A typical problem in this area is to estimate porosity from X-ray Computed Tomography (CT) inspections. This investigation aims to understand the effectiveness of open-source image segmentation methods for porosity estimation from optical images in the context of strength analysis of bonded materials and provides a benchmark of performance against a standard commercial software package. The rank ordering of the specimens by porosity area fraction is consistent across both methods.

Porosity area fraction↗