Search NASASearch

SEARCH · Search NASA

Results for “Random Forest”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Southern Rockies Western Slope Agriculture: Identifying Drivers of Rangeland Production for Drought Planning on the Western Slope of the Southern Rockies

Over the last decade, the southern Rocky Mountains of the United States experienced severe and variable drought. Local ranchers and landowners have reported strain on their operations, citing decreasing forage for their cattle and a need to adjust their business models. This study identified Major Land Resource Area-48 (MLRA-48) and northwestern Colorado as the key region for analysis. NASA DEVELOP partnered with the BLM Colorado River Field Office, Colorado State University Extension, USDA Forest Service, and the National Drought Mitigation Center to address concerns regarding the efficacy of remotely sensed rangeland production platforms and identify early warning climatic indicators of drought. The study identified two key platforms, The Rangeland Productivity Monitoring Service (RPMS) and Rangeland Analysis Platform (RAP), which use NASA Landsat 5 TM, Landsat 7 ETM+, Landsat 8 OLI, and Landsat 9 OLI-2 to estimate rangeland biomass. We regressed these with in situ biomass data to validate their efficacy and found that RAP was more effective than RPMS in estimating rangeland biomass, though it presents a tendency to overestimate. Our study performed a random forest analysis, comparing monthly RAP biomass estimates to a variety of climate variables, including mean precipitation, temperature, Palmer Drought Severity Index, snow water equivalent, snow persistence from Terra MODIS, wind speed and direction, and vapor pressure deficit. We determined that vapor pressure deficit and precipitation are key indicators in predicting forage production in MLRA-48. Our climate analysis provided our partners with greater understanding of the influence of various climate variables in determining rangeland production and allows them to assist land managers in drought mitigation.

remote sensing

California & Oregon Ecological Forecasting: Detecting and Forecasting Fog Occurrence, Frequency, and Change to Support Coast Redwood (Sequoia sempervirens) Habitat Assessments

Fog and low clouds play an important role in providing moisture to coastal ecosystems. Coast redwood (Sequoia sempervirens) forests are currently distributed along a narrow strip of coastline in California and Oregon and rely on the presence of marine fog for moisture availability during the dry season (June-October). Recent time series analyses presented an uncertain future of fog frequency; however, a decline in fog presence may have adverse effects on the coast redwood habitat. To support Save the Redwoods League, a non-profit organization dedicated to coast redwood forest management, the team analyzed hourly fog data from the Geospatial Operational Environmental Satellite 17 (GOES-17) Advanced Baseline Imager (ABI) and daily cloud cover data from the Moderate Resolution Imaging Spectroradiometer (MODIS) aboard the Terra satellite. To explore present day fog longevity, GOES-17 was utilized to map the number of fog hours per day for the 2019 and 2020 dry seasons. The MODIS cloud flag was used to map the presence or absence of daily fog, which was summarized to create a monthly fog frequency dataset and identify trends in fog presence between 2000-2020. Both datasets were used as inputs into the random forest machine learning algorithm to identify climatic drivers of fog presence and longevity over the landscape. The present-day models suggested that daily temperature difference is a driving force behind fog presence and longevity. Trends in fog presence from 2000-2020 indicated great interannual variability. Finally, fog presence was modeled under a 2080 climate projection to shed light on the future of fog presence under a projected warmer climate. Model results projected an overall decline in fog presence during the dry season in the 2080s. Decreased fog presence as a result of increased temperature difference under a warmer climate remains to be a topic of investigation as to the impact on future redwood habitat suitability.

DEVELOP Project Summary

California & Oregon Ecological Forecasting: Detecting and Forecasting Fog Occurrence, Frequency, and Change to Inform Coast Redwood (Sequoia sempervirens) Habitat Assessments

Fog and low clouds play an important role in providing moisture to coastal ecosystems. Coast redwood (Sequoia sempervirens) forests are currently distributed along a narrow strip of coastline in California and Oregon and rely on the presence of marine fog for moisture availability during the dry season (June-October). Recent time series analyses presented an uncertain future of fog frequency; however, a decline in fog presence may have adverse effects on the coast redwood habitat. To complement ongoing work by Save the Redwoods League, a non-profit organization dedicated to coast redwood forest management, the team analyzed hourly fog data from the Geospatial Operational Environmental Satellite 17 (GOES-17) Advanced Baseline Imager (ABI) and daily cloud cover data from the Moderate Resolution Imaging Spectroradiometer (MODIS) aboard the Terra satellite. To explore present day fog longevity, GOES-17 was utilized to map the number of fog hours per day for the 2019 and 2020 dry seasons. The MODIS cloud flag was used to map the presence or absence of daily fog, which was summarized to create a monthly fog frequency dataset and identify trends in fog presence between 2000-2020. Both datasets were used as inputs into the random forest machine learning algorithm to identify climatic drivers of fog presence and longevity over the landscape. The present-day models suggested that daily temperature difference is a driving force behind fog presence and longevity. Trends in fog presence from 2000-2020 indicated great interannual variability. Finally, fog presence was modeled under a 2080 climate projection to shed light on the future of fog presence under a projected warmer climate. Model results projected an overall decline in fog presence during the dry season in the 2080s. Decreased fog presence as a result of increased temperature difference under a warmer climate remains to be a topic of investigation as to the impact on future redwood habitat suitability.

DEVELOP Technical Paper

Mapping tree height in complex terrain of northern China using ultra-high-resolution images

Tree height is a key parameter for estimating forest biomass and carbon sequestration. In recent years, notable progress has been made in mapping tree height using satellite imagery. However, existing tree height products show low accuracy in mountainous and complex terrains, and few studies typically addressed tree height estimations in mountain areas. This study examines the Mentougou district of Beijing, China, characterized by complex terrain and mountainous landscapes. We analyzed two methods for estimating tree height: one using only spectral features and another combining spectral features with topographic factors (elevation, slope, aspect). We used 3-m resolution PlanetScope 8-band multispectral imagery, with 710 field-measured individual tree heights averaged to obtain 471 pixel-level tree height values as ground-truth, to develop tree height prediction models using eXtreme Gradient Boosting (XGBoost), Random Forest (RF), and Gradient Boosting Machine (GBM) models. The results show that the XGBoost model consistently presented the highest accuracy for both methods evaluated. Specifically, the XGBoost model that combined spectral data with elevation and slope variables with an R² of 0.75 and an RMSE of 2.69 m. Using the XGBoost model, we generated the tree height map for the Mentougou area at 3 m resolution, showing tree heights ranging from 0.5 to 30.4 m, and the model’s prediction error standard deviations ranged from 2.50 to 4.71 m, indicating reliable performance across varied terrain. Additionally, we compared and evaluated the global tree height products, identifying limitations in the accuracy within complex terrains. This study demonstrates the potential for accurately predicting tree heights by combining high-resolution multispectral satellites with a terrain factor modeling approach.

Complex terrain

New York Ecological Forecasting: Utilizing NASA Earth Observations to Map Ash Distribution and Inform Emerald Ash Borer Control

Since their first sightings in the U.S. in 2002, emerald ash borer beetles (Agrilus planipennis; EAB) have killed millions of native ash (Fraxinus spp.) trees across 35 states. Infected ash stands frequently exhibit complete mortality, with the predicted result being the functional extinction of native ash in U.S. forests. In August of 2020, EAB was discovered in the 6.1-million-acre Adirondack Park. The team’s partners at the Adirondack Park Invasive Plant Program (APIPP) desired ash tree distribution and EAB susceptibility information to help improve EAB bio-control efficiency and apply the methodology to future invasive programs. To assist, the team mapped ash tree distribution using NASA Earth observations from Landsat 7 Enhanced Thematic Mapper Plus (ETM+) and Shuttle Radar Topography Mission (SRTM), along with hyperspectral imagery from the Airborne Visible/Infrared Imaging Spectrometer (AVIRIS). Field data from the Monitoring and Managing Ash (MaMA) project, iMapInvasives and iNaturalist databases, and the New York State Department of Environmental Conservation (NYSDEC) provided ground truthing for mapping and modeling. Results indicate that for ash detection, the team’s Spectral Angle Mapping (SAM) hyperspectral classification is slightly more sensitive but less accurate than multispectral Random Forest (RF) classification, though neither method was above a ~20% detection rate. End products include maps of ash extent derived from both imagery types, a model forecasting future spread scenarios based on current EAB presence, and outreach materials. These products inform APIPP’s management decisions and facilitate public awareness of EAB’s threat to communities within the region.

Liam Megraw

Northern Rockies Ecological Conservation: Leveraging Earth Observations to Monitor and Predict Populations of Federally Threatened Whitebark Pine (Pinus albicaulis) across the Intermountain West

Whitebark pine (WBP; Pinus albicaulis) is an ecologically important species in North America. As a federally listed threatened species, an understanding of WBP habitat, distribution, and health is important for the natural resource managers of the National Park Service, United States Forest Service, Bureau of Land Management, Fish and Wildlife Service, and non-profit organizations such as the Whitebark Pine Ecosystem Foundation. Previous attempts to develop models of WBP habitat suitability and distribution lack confidence in their validity and integrity for these organizations. The updated models of habitat suitability and distribution developed by this study would provide managers with a capability to be employed in the conservation and future research direction for WBP. Thus, we developed a habitat suitability model of WBP at a high spatial resolution (Landsat 9 Operational Land Image-2, National Land Cover Database, NASA Shuttle Radar Topography Mission; 30m pixels) using a generalized logistic regression with an area under the curve value of 0.754. We extracted spectral reflectance signatures from overlapped ground sample points and Sentinel-2 Multispectral Instrument. The spectral signature analysis indicates WBP is separable from other tree species. We also utilized a visual validation approach and random forest (RF) modeling to separate WBP from limber pine. Through visual validation the RF classifier successfully identified 8out of 10 WBP trees gathered through ground truth points. Additionally, we achieved an overall accuracy of 91%in our confusion matrix for the distribution model using a dependent validation approach. The derived products from this study allow project partners to assess current suitable habitat and apparent health status in areas of identified WBP occurrence, providing data to aid future research regarding WBP health.

Sentinel-2

Assessing Alaskan boreal forest landcover affected by climate-wildfire interactions from ground truth surveys and NASA airborne remote sensing

Alaska’s boreal forest is facing unprecedented challenges under rapid climate warming (increasingly severe fires, droughts, pest/disease outbreaks) that may destabilize its function as a global carbon sink. Forests near Fairbanks may be especially vulnerable, impacting air quality and ecosystem services. We combined GT (ground truthing) with Airborne Visible InfraRed Imaging Spectrometer (AVIRIS-NG) images collected by the NASA Arctic-Boreal Vulnerability Experiment (ABoVE) program (2017-2019) to assess landcover change at five recently burned sites (2001-2019) of different fire severities and moisture regimes within 30 miles of Fairbanks. GT included tree seedling counts, understory % cover and >50% leaf canopy color assessment. 36 circular plots (1/30 ha radius) including 6 moderate to severely burned plots were selected across sites. 31 additional sites including 12 burned sites were geotagged in photos. AVIRIS images were processed from 29 spectral bands selected to identify changes in chlorophyll and water content. Images were segmented into natural boundaries (polygons) using ENVI 5.5 software. A spectral library of 8 AVIRIS bands with high between-class/low within-class variation was used in two random forest models to predict vegetation classes (model 1: 12 classes, model 2: 14 classes) in each AVIRIS scene, using 20% of the data as training data. Model 2 classified 20% more polygons overall, but only 42% of GT/geotagged polygons were correctly classified by both models. More forest sites were correctly classified (63%) than open vegetation (32%) or post-fire sites (46%). 50% of aspen forest and post-fire polygons were misclassified as shrubland. GT revealed that post-fire plots supported 134,000 (± 48,000) tree seedlings and saplings ha-1 (0.2 - 4 m height, 64% deciduous) versus 2500 (± 2100) shrubs ha-1 (1-6 m height). > 50% canopy browning was observed in conifer forest (8 plots) with no signs of insect infestation. Canopy herbivory > 50% (leaf miner, leaf beetle) and moose herbivory of tree bark was seen across aspen sites. Our study suggests: 1) low canopy vegetation presents challenges for improved landcover classification, and 2) aspen forest should be differentiated in vegetation maps which would aid in tracking herbivory.

Alaska

Water Across Synthetic Aperture Radar Data (WASARD): SAR Water Body Classification for the Open Data Cube

The detection of inland water bodies from Synthetic Aperture Radar (SAR) data provides a great advantage over water detection with optical data, since SAR imaging is not impeded by cloud cover. Traditional methods of detecting water from SAR data involves using thresholding methods that can be labor intensive and imprecise. This paper describes Water Across Synthetic Aperture Radar Data (WASARD): a method of water detection from SAR data which automates and simplifies the thresholding process using machine learning on training data created from Geoscience Australia’s WOFS algorithm. Of the machine learning models tested, the Linear Support Vector Machine was determined to be optimal, with the option of training using solely the VH polarization or a combination of the VH and VV polarizations. WASARD was able to identify water in the target area with a correlation of 97% with WOFS. Sentinel-1, Open Data Cube, Earth Observations, Machine Learning, Water Detection 1. INTRODUCTION Water classification is an important function of Earth imaging satellites, as accurate remote classification of land and water can assist in land use analysis, flood prediction, climate change research, as well as a variety of agricultural applications [2]. The ability to identify bodies of water remotely via satellite is immensely cheaper than contracting surveys of the areas in question, meaning that an application that can accurately use satellite data towards this function can make valuable information available to nations which would not be able to afford it otherwise. Highly reliable applications for the remote detection of water currently exist for use with optical satellite data such as that provided by LANDSAT. One such application, Geoscience Australia’s Water Observations from Space (WOFS) has already been ported for use with the Open Data Cube [6]. However, water detection using optical data from Landsat is constrained by its relatively long revisit cycle of 16 days [5], and water detection using any optical data is constrained in that it lacks the ability to make accurate classifications through cloud cover [2]. The alternative solution which solves these problems is water detection using SAR data, which images the Earth using cloud-penetrating microwaves. Because of its advantages over optical data, much research has been done into water detection using SAR data. Traditionally, this has been done using the thresholding method, which involves picking a polarization band and labeling all pixels for which this band’s value is below a certain threshold as containing water. The thresholding method works since water tends to return a much lower backscatter value to the satellite than land [1]. However, this method can be flawed since estimating the proper threshold is often imprecise, complicated, and labor intensive for the end user. Thresholding also tends to use data from only one SAR polarization, when a combination of polarizations can provide insight into whether water is present. [2] In order to alleviate these problems, this paper presents an application for the Open Data Cube to detect water from SAR data using support vector machine (SVM) classification. 2. PLATFORM WASARD is an application for the Open Data Cube, a mechanism which provides a simple yet efficient means of ingesting, storing, and retrieving remote sensing data. Data can be ingested and made analysis ready according to whatever specifications the researcher chooses, and easily resampled to artificially alter a scene’s resolution. Currently WASARD supports water detection on scenes from ESA’s Sentinel-1 and JAXA’s ALOS. When testing WASARD, Sentinel-1 was most commonly used due to its relatively high spatial resolution and its rapid 6 day revisit cycle [5]. With minor alterations to the application's code, however, it could support data from other satellites. 3. METHODOLOGY Using supervised classification, WASARD compares SAR data to a dataset pre-classified by WOFS in order to train an SVM classifier. This classifier is then used to detect water in other SAR scenes outside the training set. Accuracy was measured according to the following metrics:  Precision: a measure of what percentage of the points WASARD labels as water are truly water  Recall: a measure of what percentage of the total water cover WASARD was able to identify.  F1 Score: a harmonic average of the precision and recall scores Both precision and recall are calculated at the end of the training phase, when the trained classifier is compared to a testing dataset. Because the WOFS algorithm’s classifications are used as the truth values when training a WASARD classifier, when precision and recall are mentioned in this paper, they are always with respect to the values produced by WOFS on a similar scene of Landsat data, which themselves have a classification accuracy of 97% [6]. Visual representations of water identified by WASARD in this paper were produced using the function wasard_plot(), which is included in WASARD. 3.1 Algorithm Selection The machine learning model used by WASARD is the Linear Support Vector Machine (SVM). This model uses a supervised learning algorithm to develop a classifier, meaning it creates a vector which can be multiplied by the vector formed by the relevant data bands to determine whether a pixel in a SAR scene contains water. This classifier is trained by comparing data points from selected bands in a SAR scene to their respective labels, which in this case are “water” or “not water” as given by the WOFS algorithm. The SVM was selected over the Random Forest model, which outperformed the SVM in training speed, but had a greater classification time and lower accuracy, and the Multilayer Perceptron Artificial Neural Network, which had a slightly higher average accuracy than the SVM, but much greater training and classification times. Figure 1: Visual representation of the SVM Classifier. Each white point represents a pixel in a SAR scene. In Figure 1, the diagonal line separating pixels determined to be water from those determined not to be water represents the actual classification vector produced by the SVM. It is worth noting that once the model has been trained, classification of pixels is done in a similar manner as in the thresholding method. This is especially true if only one band was used to train the model. 3.1 Feature Selection Sentinel-1 collects data from two bands: the Vertical/Vertical polarization (VV) and the Vertical/Horizontal polarization (VH). When 100 SVM classifiers were created for each polarization individually, and for the combination of the two, the following results were achieved: Figure 2: Accuracy of classifiers trained using different polarization bands. Precision and Recall were measured with respect to the values produced by WOFS. Figure 2 demonstrates that using both the VV and VH bands trades slightly lower recall for significantly greater precision when compared with the VH band alone, and that using the VV band alone is inferior in both metrics. WASARD therefore defaults to using both the VV and VH bands, and includes the option to use solely the VH band. The VV polarization’s lower precision compared to the VH polarization is in contrast to results from previous research and may merit further analysis [4]. 3.2 Training a Classifier The steps in training a classifier with WASARD are 1. Selecting two scenes (one SAR, one optical) with the same spatial extents, and acquired close to each other in time, with a preference that the scenes are taken on the same day. 2. Using the WOFS algorithm to produce an array of the detected water in the scene of optical data, to be used as the labels during supervised learning 3. Data points from the selected bands from the SAR acquisition are bundled together into an array with the corresponding labels gathered from WOFS. A random sample with an equal number of points labeled “Water” and “Not Water” is selected to be partitioned into a training and a testing dataset 4. Using Scikit-Learn’s LinearSVC object, the training dataset is used to produce a classifier, which is then tested against the testing dataset to determine its precision and recall The result is a wasard_classifier object, which has the following attributes: 1. f1, recall, and precision: 3 metrics used to determine the classifier’s accuracy 2. Coefficient: Vector which the SVM uses to make its predictions. The classifier detects water when the dot product of the coefficient and the vector formed by the SAR bands is positive 3. Save(): allows a user to save a classifier to the disk in order to use it without retraining 4. wasard_classify(): Classifies an entire xarray of SAR data using the SVM classifier All of the above steps are performed automatically when the user creates a wasard_classifier object. 3.3 Classifying a Dataset Once the classifier has been created, it can be used to detect water in an xarray of SAR data using wasard_classify(). By taking the dot product of the classifier’s coefficients and the vector formed by the selected bands of SAR data, an array of predictions is constructed. A classifier can effectively be used on the same spatial extents as the ones where it was trained, or on any area with a similar landscape. While

Kreiser, Zachary

Performance Prediction of High‐Entropy Perovskites La 0.8 Sr 0.2 Mn x Co y Fe z O 3 with Automated High‐Throughput Characterization of Combinatorial Libraries and Machine Learning

Perovskite oxides form a large family of materials with applications across various fields, owing to their structural and chemical flexibility. Efficient exploration of this extensive compositional space is now achievable through automated high-throughput experimentation combined with machine learning. In this study, we investigate the composition–structure–performance relationships of high-entropy La 0.8 Sr 0.2 Mn x Co y Fe z O 3±𝞭 perovskite oxides (0 < x, y, z <1; x+y+z≈1) for application as oxygen electrodes in Solid Oxide Cells. Following the deposition of a continuous compositional map using thin-film combinatorial pulsed laser deposition, compositional, structural, and performance properties are characterized using six different techniques with mapping capabilities. Random forests effectively model electrochemical performance, consistently identifying Fe-rich oxides as optimal compounds with the lowest area-specific resistance values for oxygen electrodes at 700 °C. Additionally, the models identify a statistical correlation between oxygen sublattice distortion—derived from spectral analysis of Raman-active modes—and enhanced performance.

high entropy oxides

Data‐Driven Insights into Rare Earth Mineralization: Machine Learning Applications Using Functional Material Synthesis Data

Understanding rare‐earth element (REE) mineralization mechanisms is essential for developing efficient separation strategies. Although the geochemical pathways that generate REE deposits are qualitatively known, quantitative links between specific conditions and mineralization outcomes remain limited. Herein, the repurpose laboratory REE hydrothermal synthesis data—originally collected for functional‐materials fabrication—as a surrogate for studying mineralization with data‐driven methods. The compiled 1,200+ hydrothermal reaction records and trained three machine‐learning models—K‐nearest neighbors (KNN), random forest (RF), and extreme gradient boosting (XGB)—to predict product elements and phases from precursors, additives, reaction conditions, and engineered features. Validation shows XGB achieves the highest accuracy. Feature importance indicates thermodynamic properties of cations and anions dominate model decisions. Correlations reveal positive relationships among precursor concentration, reaction time, pH, and temperature, consistent with classical crystallization behavior. XGB‐based regressors are built to predict crystallization temperature and pH from precursor/product attributes. Performance is strongest when similar training examples exist, while accuracy declines for underrepresented reactions, notably REE carbonates and heavy‐REE systems. Overall, the study shows that functional‐materials datasets can illuminate REE mineralization and provide priors for exploration and processing. Expanding datasets with less‐studied chemistries and conditions will improve generality and support deposit discovery and more efficient REE recovery.

feature importance analysis

ytopt: Autotuning Scientific Applications for Energy Efficiency at Large Scales

As we enter the exascale computing era, efficiently utilizing power and optimizing the performance of scientific applications under power and energy constraints has become critical and challenging. We propose a low-overhead autotuning framework to autotune performance and energy for various hybrid MPI/OpenMP scientific applications at large scales and to explore the tradeoffs between application runtime and power/energy for energy efficient application execution, then use this framework to autotune four ECP proxy applications—XSBench, AMG, SWFFT, and SW4lite. Our approach uses Bayesian optimization with a Random Forest surrogate model to effectively search parameter spaces with up to 6 million different configurations on two large-scale HPC production systems, Theta at Argonne National Laboratory and Summit at Oak Ridge National Laboratory. The experimental results show that our autotuning framework at large scales has low overhead and achieves good scalability. Using the proposed autotuning framework to identify the best configurations, we achieve up to 91.59% performance improvement, up to 21.2% energy savings, and up to 37.84% EDP (energy delay product) improvement on up to 4096 nodes.

Autotuning

Predictive links between microbial communities and biological oxygen utilization in the Arctic Ocean

Microbial metabolism influences rates of net community production (NCP), exerting a direct biological control on marine oxygen and carbon fluxes. In the Arctic, it is increasingly important to understand and quantify this process, as ecological and oceanographic conditions shift due to changing climate. Here, we describe potential ecological links between pelagic microbial diversity and an NCP precursor, biological oxygen utilization, using machine learning and paired observations of community structure and metabolic activity from a seasonally and spatially variable transect of the Arctic Ocean (2019–2020 MOSAiC Expedition). Community structure was determined using 16S (prokaryotic) and 18S (eukaryotic) rRNA gene amplicon sequencing, and metabolic activity was derived from ΔO 2 /Ar. Using self-organizing maps, we identified clear successional patterns in observed microbial community structure that were seasonally driven in the upper ocean and vertically stratified with depth. Metabolic activity was also stratified, with a primarily net heterotrophic water column (median −1.5% biological oxygen saturation), excepting periodic oxygen supersaturation (maximum: 13.6%) within the mixed layer. Using DNA sequences as predictor variables, we then constructed a random forest regression model that reliably reconstructed biological oxygen concentrations (root mean squared error = 4.14 μmol kg −1 ). Top predictors from this model were from heterotrophic (bacteria) or potentially mixotrophic (dinoflagellate) taxa. These analyses highlight biologically driven diagnostic tools that can be used to expand biogeochemical datasets and improve the microbial perspectives and metabolisms represented in ecological models of net productivity and carbon flux in a changing Arctic Ocean.

Chamberlain, Emelia J. [Univ. of San Diego, San Di

A data-driven framework for predicting machining stability: employing simulated data, operational modal analysis, and enhanced transfer learning

Chatter, a self-excited vibration phenomenon, presents a significant challenge in machining operations, particularly in high-speed milling, where it can degrade tool life, reduce material removal efficiency, and compromise workpiece quality. Addressing this challenge requires a reliable predictive model that can accommodate the complex dynamics of various machining scenarios. This study introduces a novel, data-driven approach to predicting machining stability, leveraging over 140,000 simulated datasets and employing advanced techniques such as operational modal analysis (OMA), enhanced transfer learning (TL), and receptance coupling substructure analysis (RCSA). By integrating these methodologies, the framework effectively classifies and predicts chatter across diverse operational modes, achieving robust and accurate outcomes. Our model utilizes a Random Forest (RF) classifier trained with the comprehensive dataset, which demonstrates substantial improvements in both predictive accuracy and robustness. Specifically, the RF model achieved an accuracy rate of 85%, an area under the curve (AUC) of 0.90, and an F1 score of 0.88, underscoring its capability to adapt to varying machining configurations. These results highlight the framework’s potential to enhance operational efficiency and machining quality by providing reliable chatter predictions across a broad range of machining parameters. In conclusion, this research thus offers a significant advancement in predictive maintenance for machining processes, enabling more stable and efficient manufacturing operations.

42 ENGINEERING

Evaluating ecosystem water use efficiency under drought stress: a case study of the Helan Mountain region, northwest China

Context Water use efficiency (WUE) is a fundamental ecological indicator links carbon assimilation and water loss in terrestrial ecosystems. Understanding its responses to drought stress is essential for adaptive ecosystem management, particularly in climate-sensitive mountain landscapes. Objectives This study aimed to investigate drought-driven variations in WUE across major vegetation types in the Helan Mountain region of Northwest China. Specifically, we sought to identify dominant ecological drivers of WUE variability and to disentangle their relative importance and causal pathways. Methods We quantified WUE using the Moderate Resolution Imaging Spectroradiometer (MODIS) products and the Drought Severity Index (DSI) data from 2001 to 2020. To examine WUE – drought relationships across contrasting vegetation types, we employed a spatially explicit analytical framework integrating Random Forest (RF) modeling, partial correlation analysis, and structural equation modeling (SEM). Results Regional WUE exhibited relatively stable interannual dynamics, yet pronounced spatial heterogeneity that was strongly modulated by drought conditions. Vegetation properties, particularly Leaf Area Index (LAI) and Normalized Difference Vegetation Index (NDVI), emerged as the dominant determinants of WUE, with NDVI alone explaining over 20% of its spatial variance in forest and grassland during non-drought periods. SEM analyses revealed that climate forcing influenced WUE mainly through indirect pathways mediated by soil moisture availability and vegetation structural dynamics, rather than through direct climatic controls. Among all regulating factors, LAI acted as the central control node governing ecosystem carbon–water coupling. In contrast, short-term climatic stress, especially atmospheric demand and drought duration, exerted weak or negative direct effects on WUE. Ecosystem-specific responses were observed, with croplands mainly regulated by soil water availability, whereas forests and grasslands showed more sensitive to atmospheric drought stress. Together, these results reveal a hierarchical control framework where soil–vegetation interactions mediate climate impacts on WUE, driving strong spatial heterogeneity in drought responses across mountain landscapes. Conclusions Our findings highlight the pivotal role of indirect drought effects mediated by vegetation and soil processes in shaping ecosystem WUE. The identified soil–vegetation–climate regulatory hierarchy provides mechanistic insight into landscape–scale drought sensitivity and supports integrated modeling approaches for evaluating ecosystem resilience and sustainable management in arid mountain regions.

China

Evaluation of normalization strategies for mass spectrometry-based multi-omics datasets

Introduction Data normalization is crucial for multi-omics integration, reducing systematic errors and maximizing the likelihood of discovering true biological variation. Most studies assess normalization for a single omics type or use datasets from separate experiments. Few address time-course data, where normalization might bias temporal differentiation. In this study, we compared common normalization methods and a machine learning approach, Systematical Error Removal using Random Forest (SERRF), using multi-omics datasets generated from the same experiment—even from the same cell lysate. Objectives To develop a straightforward process to assess normalization effects and identify the most robust methods across multi-omics datasets. Methods We analyzed metabolomics, lipidomics, and proteomics datasets from primary human cardiomyocytes and motor neurons exposed to acetylcholine-active compounds over time. Normalization effectiveness was evaluated based on improvement in QC features consistency and observing the change in treatment and time-related variance. Results Probabilistic Quotient Normalization (PQN) and Locally Estimated Scatterplot Smoothing (LOESS) QC were identified as optimal for metabolomics and lipidomics, while PQN, Median, and LOESS normalization excelled for proteomics. These methods consistently enhanced QC feature consistency in metabolomics and lipidomics, and preserved time-related variance or treatment-related variance in proteomics, demonstrating their effectiveness and robustness. SERRF normalization, applied only to metabolomics in this study, outperformed other methods in some datasets but inadvertently masked treatment-related variance in others. Conclusion Our evaluation identified PQN and LoessQC as the top methods for metabolomics and lipidomics, and PQN, Median, and Loess normalization for proteomics, in multi-omics integration in a temporal study.

60 APPLIED LIFE SCIENCES

Macroscopic Traffic Modeling Using Probe Vehicle Data: A Machine Learning Approach

Abstract The macroscopic fundamental diagram (MFD) captures an orderly relationship among traffic flow, density, and speed at the network level. It is a simple yet powerful tool for modeling traffic dynamics in large urban networks with broad application in traffic control and management. However, empirically derived MFDs in urban regions require high-resolution traffic data from the network. Having the network flow and vehicular density estimated at the (granular) census tract level using vehicle probe data, we apply machine learning methods to predict the MFDs across U.S. urban areas and capture the impacts of location-specific input features on the network flow–density relationships at a large scale. The results show that, among the four tested machine learning approaches (Random Forest, XGBoost, Support Vector Machine, and Neural Network), XGBoost delivers the best performance in predicting network traffic flow based on vehicular density and location attributes. Using interaction Shapley Additive explanation (SHAP) values and partial correlation analysis, we examine the factors influencing MFD shapes across different locations. Our empirical findings reveal that across U.S. urban areas, network topology, transportation infrastructure, and land use are primary factors shaping MFD curves, while demand and trip-related factors play a lesser role. Specifically, higher ranking roads, centrality, and development levels correlate positively with network capacity and critical density, whereas negative associations are observed for network connectivity, mixed-use development, and road roughness levels.

Jin, Ling

Analyzing the impact of design factors on solar module thermomechanical durability using interpretable machine learning techniques

Solar modules in utility-scale systems are expected to maintain decades of lifetime to rival conventional energy sources. However, cyclic thermomechanical loading often degrades their long-term performance, highlighting the importance of effective design to mitigate thermal expansion mismatches between module materials. Given the complex composition of solar modules, isolating the impact of individual components on overall durability remains a challenging task. In this work, we analyze a comprehensive data set that comprises bill-of-materials (BOM) and thermal cycling power loss from 251 distinct module designs to identify the predominant design factors and their impacts on the thermomechanical durability of modules. The methodology of our analysis combines machine learning modeling (random forest) and Shapley additive explanation (SHAP) to correlate design factors with power loss and interpret the model’s decision-making. The interpretation reveals that silicon type (monocrystalline or polycrystalline), encapsulant thickness, busbar numbers, and wafer thickness predominantly influence the degradation. With lower power loss of around 0.6% on average in the SHAP analysis, monocrystalline cells present better durability than polycrystalline cells. This finding is further substantiated by statistical testing on our raw data set. The SHAP analysis also demonstrates that while thicker encapsulants lead to reduced power loss, further increasing their thickness over around 0.6 to 0.7 mm does not yield additional benefits, particularly for the front side one. In addition, other important BOM features such as the number of busbars are analyzed. This study provides a blueprint for utilizing explainable machine learning techniques in a complex material system and can potentially guide future research on optimizing the design of solar modules.

14 SOLAR ENERGY

Constraining Galaxy-Halo connection using machine learning

We investigate the potential of machine learning (ML) methods to model small-scale galaxy clustering for constraining Halo Occupation Distribution (HOD) parameters. Our analysis reveals that while many ML algorithms report good statistical fits, they often yield likelihood contours that are significantly biased in both mean values and variances relative to the true model parameters. This highlights the importance of careful data processing and algorithm selection in ML applications for galaxy clustering, as even seemingly robust methods can lead to biased results if not applied correctly. ML tools offer a promising approach to exploring the HOD parameter space with significantly reduced computational costs compared to traditional brute-force methods if their robustness is established. Using our ANN-based pipeline, we successfully recreate some standard results from recent literature. Properly restricting the HOD parameter space, transforming the training data, and carefully selecting ML algorithms are essential for achieving unbiased and robust predictions. Among the methods tested, artificial neural networks (ANNs) outperform random forests (RF) and ridge regression in predicting clustering statistics, when the HOD prior space is appropriately restricted. We demonstrate these findings using the projected two-point correlation function (w p (r p )), angular multipoles of the correlation function (ξ ℓ (r)), and the void probability function (VPF) of Luminous Red Galaxies from Dark Energy Spectroscopic Instrument mocks. Our results show that while combining w p (r p ) and VPF improves parameter constraints, adding the multipoles ξ 0 , ξ 2 , and ξ 4 to w p (r p ) does not significantly improve the constraints.

cosmology