Search NASA⌕ Search

SEARCH · Search NASA

Results for “Random Forest”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Automated classification of scientific publications linked to GES DISC datasets

The data collections archived and distributedby the GES DISC NASA data center arewidely utilized for various Earth Science studies.As these collections are created, many researchworks are published regarding the collections, algorithms,validations and applications. SinceGES DISC collects these publications and providestheir citations for the users, it is helpful tocategorize them based on how they relate to the datasetsthey are associated with. Specifically,whether the publication that is linked to GES DISCdataset is using it for applicational research,or if it describes the algorithm for dataset creation,or the validation of the dataset, or providesthe general overview of the data collection. Currently,this process requires simple manuallabelling, and as such, may be possible to solve viaautomation. To approach this problem, wedeveloped machine learning classifiers to predictthe category a publication belongs to. We usedmanually labeled publications as training data forsupervised machine learning algorithms:Random Forest and Naive Bayes. We achieved classificationaccuracy that is substantially betterthan the baseline accuracy, thus greatly improvingthe efficiency of the publication internalanalysis.

Rohan Dayal↗

Classifying Agnostic Biosignatures using Raman, VNIR, and Elemental Data

How can we use our current wealth of terrestrial data, encompassing biogenic and abiogenic systems, to determine the distinguishing properties of life? SCOBI (Statistical Classification of Biosignature Information) uses machine learning techniques to algorithmically identify combinations of measurements that are “indicative of life”. A set of ~1000 observations, comprising elemental abundance, isotopic fractionation, VNIR reflectance, and (in progress) Raman spectra, have been assembled from existing literature and databases. The observations cover systems classified as “indicative alive” (e.g., cells, vegetation), “indicative non-alive” (e.g., fossils, teeth), “mixed indicative” (e.g., soil, pond water), or “non-indicative” (e.g., rocks, meteorites). VNIR data was preprocessed by linear interpolation from 400-2100 nm and smoothed with a Savitzky-Golay filter. To limit the amount of Earth-biochemistry-specific (non-agnostic) information included, the first five spectral features extracted were number of peaks, number of troughs, mean reflectance, mean peak width, and broadest peak width. To help further emphasize agnostic biosignatures, Earth-specific features such as chlorophylls have been manually flagged so that feature importance with and without them can be compared. Classifiers including k-nearest neighbors (KNN), Gaussian Naïve Bayes (GNB), logistic regression (LR), random forest (RF), and support vector machine (SVM) were implemented, as was a combination voting classifier. Performance metrics included false positive rates, false negative rates, and AUC with 50-50 test/train splits (Monte Carlo simulations). Key takeaways from this stage, prior to the inclusion of Raman spectra, are (1) the overall success rate of 0.933 AUC was most heavily influenced by the elemental abundance data; and (2) VNIR reflectance had the lowest classification performance with 0.52 AUC (58% of objects correctly classified). The next steps are to complete integration of Raman spectral data and to improve the approach to pre-processing and feature extraction for both types of spectral data, such as automated baseline removal, whole spectrum matching, and dimensionality reduction.

Biosignatures↗

Machine Learning Applications to Metal-Silicate Equilibria and their Insights into Core Formation

An extensive number of studies have experimentally investigated how elements distribute between metal and silicate phases, to better constrain core-mantle chemical equilibrium. Here, we present a new database compiling all (to our knowledge) experimental data on liquid metal-silicate partitioning from 118 peer-reviewed publications. We applied various machine learning techniques to gain further insights into these partitioning equilibria and their dependencies. We performed a network analysis to investigate the relationship between experiments and partition coefficients, which enables visualizing gaps in the experimental dataset and biases related to varying experimental conditions and analytical setup. In addition, semi-empirical thermodynamic models are commonly used to extrapolate these chemical reactions to the wide range of pressure, temperature and compositional conditions of planetary differentiation. These models are based on linear regressions that assume continuous relationship between partition coefficients and experimental variables. Here, we considered random forest regressions, which are algorithms based on ensembles of decision trees and does not consider continuous effects of each variable. The application of this regression significantly improves the prediction of metal-silicate partitioning for several elements including Ni, Si and Cr. We will show how this new approach improves our understanding of elemental exchange between metal and silicate and their implications for the Earth’s core formation.

siderophile element↗

Spatial Variation of Fine Particulate Matter Levels in Nairobi Before and During the COVID-19 Curfew: Implications for Environmental Justice

The temporary decrease of fine particulate matter (PM(sub 2.5)) concentrations in many parts of the world due to the COVID-19 lockdown spurred discussions on urban air pollution and health. However there has been little focus on sub-Saharan Africa, as few African cities have air quality monitors and if they do, these data are often not publicly available. Spatial differentials of changes in PM(sub 2.5) concentrations as a result of COVID also remain largely unstudied. To address this gap, we use a serendipitous mobile air quality monitoring deployment of eight Sensirion SPS 30 sensors on motorbikes in the city of Nairobi starting on 16 March 2020, before a COVID-19 curfew was imposed on 25 March and continuing until 5 May 2020. We developed a random-forest model to estimate PM(sub 2.5) surfaces for the entire city of Nairobi before and during the COVID-19 curfew. The highest PM(sub 2.5) concentrations during both periods were observed in the poor neighborhoods of Kariobangi, Mathare, Umoja, and Dandora, located to the east of the city center. Changes in PM(sub 2.5) were heterogeneous over space. PM(sub 2.5) concentrations increased during the curfew in rapidly urbanizing, the lower-middle-class neighborhoods of Kahawa, Kasarani, and Ruaraka, likely because residents switched from LPG to biomass fuels due to loss of income. Our results indicate that COVID-19 and policies to address it may have exacerbated existing air pollution inequalities in the city of Nairobi. The quantitative results are preliminary, due to sampling limitations and measurement uncertainties, as the available data came exclusively from low-cost sensors. This research serves to highlight that spatial data that is essential for understanding structural inequalities reflected in uneven air pollution burdens and differential impacts of events like the COVID pandemic. With the help of carefully deployed low cost sensors with improved spatial sampling and at least one reference-quality monitor for calibration, we can collect data that is critical for developing targeted interventions that address environmental injustice in the African context.

Environmental justice↗

Algorithmic Classification of Raman Spectra Biosignatures: Improving Life Detection Confidence

“Agnostic” biosignatures – indicators of life (or the absence of life), independent of a particular biochemistry – are increasingly considered a high standard for life detection. The Ladder of Life Detection (2018) called for investigating how combinations of independent and different potential biosignatures affect confidence. To address this gap, statistical classification of elemental abundances, isotopic fractionation, and reflectance spectroscopy (VNIR) has been implemented. Raman spectroscopy, highly desirable due to its wide availability, has the potential to improve this predictive power. This work implemented biosignature classification algorithms on Raman data alone, in preparation for combination with the other data types. Raman spectroscopy data was collected from published databases and papers as part of a manually curated dataset of “indicative” and “non-indicative of life” samples. These currently include 61 non-indicative samples (meteorites, magnetite); 3 indicative living samples (bacteria); 20 indicative non-living samples (chalk, bone); and 12 indicative mixed (with non-indicative material) samples (soil, microbial mats). Laboratory work is ongoing to characterize additional samples, particularly a greater breadth of mixed systems. Spectra were interpolated, filtered with the Savitzsky-Golay filter, and de-noised. For a preliminary examination, agnostic features were manually extracted including mean intensity, number of peaks, and mean peak width. Different peak prominences and filtering polynomials were used to refine features. Classification algorithms were implemented: k-nearest neighbors (KNN), logistic regression (LR), linear support vector machines (SVM), random forest (RF), Gaussian naïve bayes (GNB). Lastly, Monte Carlo simulations on 1,000 50%-train-test-splits were used to validate classification performance and feature significance. The preliminary feature set achieved its highest AUC of 0.52 with LR, with no strongly discriminatory features. Work to improve feature extraction, such as through deep learning with back propagation, is planned. In future work, the Raman data will be combined with the other data types, and potentially new data types such as enantiomeric excess. This project was partially supported through the NASA Ames Project EXcellence (APEX) incubator program.

Astrobiology↗

Remote Sensing of Lineage Functional Types for Modeling and Monitoring Biodiversity

Hyperspectral remote sensing has the potential to continuously scale plant function and plant diversity information from landscape to global extents. Numerous studies have indicated that VSWIR (400-2500 nm) reflectance properties of vegetation capture evolutionarily conserved biochemical, structural, and other functional attributes of plant species. Spectral properties conserved in plants provide the opportunity to both 1) aggregate species into lineages with improved classification accuracy and 2) link those lineages directly to plant traits. Full realization of this goal will enable parameterization of Land Surface Models (LSMs) with remotely sensed information, e.g., canopy nitrogen, and better representations of biodiversity and functional diversity in biogeographic studies. In this study, we use hyperspectral AVIRIS data from the 2013 HyspIRI campaign over the Southern Sierra Nevada, California flight box to investigate the potential for incorporating evolutionary thinking into landcover classification. We link the airborne hyperspectral data with vegetation plot data from roughly 1372 surveys and a phylogeny representing 1361 species. We aggregate species into lineages ranging from species level groups down to similar number of Plant Functional Types as often used in LSMs. We assessed the ability of Random Forest and Partial Least Squares Discriminant Analysis to discriminate across these different phylogenetic scales and determine the optimal number of lineages to classify. Although there are some temporal and spatial differences in our training data, our best approaches achieved moderate classification accuracy (Kappa > 0.65). Given an optimal number of lineages, we explored approaches to improve classifications including machine learning and unmixing approaches. This work suggests that lineage-based methods may be a promising way to leverage the huge amounts of data that will come from high resolution and high return interval hyperspectral data planned for the Surface Biology and Geology mission with sparsely sampled existing ground-based ecological data.

Hyperspectral↗

Powder River Basin Water Resources: Mapping Russian Olive in the Powder River Basin to Inform Invasive Species Management

Since its introduction in the late 1800s, Elaeagnus augustifolia (Russian olive) has become a widespread invasive shrub that poses a threat to native riparian species in the United States by competing with native riparian plants for space and resources. To date, limited information on the distribution of Russian olive in the Powder River Basin of Montana and Wyoming have hampered management efforts and decision making. Here, we detect and model the distribution of Russian Olive using field surveys, ocular sampling, and variables from Landsat 8 Operational Land Imager (OLI), Sentinel-2 MultiSpectral Instrument (MSI), and Shuttle Radar Topography Mission (SRTM) using the Random Forest algorithm. We derived topographic, spectral, and hydrological variables from Landsat 8 OLI, Sentinel-2 MSI, and SRTM to utilize as model inputs. The team was able to successfully create a spectral Russian olive detection map for the Powder River Basin (RMSE =15.44%, R2 = 0.6482). The team also examined change in stream channel geomorphology from 1984-2020 in a time-series analysis using Landsat visible imagery and the RivMap MATLAB package and found little change. Our results will help our partners at the Powder River County Weed Board, Gay Ranch, United States Geological Survey, and University of Northern Colorado to locate and prioritize areas for riparian habitat restoration and to understand the region’s hydrology and geomorphology.

Catherine Buczek↗

Dynamic Ensemble Prediction of Cognitive Performance in Space

Astronauts are exposed to a unique set of stressors in spaceflight. Microgravity, isolation, confinement, and environmental and operational hazards: all of these can impact sleep, vigilant attention, and alertness, which are critical to mission success. In this paper, we seek to understand the most important predictors of alertness over the course of a space mission, using self-reported, cognitive, and environmental data collected from 24 astronauts on 6-month missions to the International Space Station (ISS). Alertness was repeatedly and objectively assessed on the ISS with a brief 3-minute Psychomotor Vigilance Test (PVT) that is highly sensitive to sleep deprivation. To relate PVT performance to time-varying and sparsely-measured environmental, operational, and psychological covariates, we propose a n ensemble prediction model comprising of linear mixed effects regression, random forest, and functional concurrent regression models. An extensive cross-validation procedure reveals that this ensemble outperforms any one of its components alone. We also discover that a participant’s past performance, reported fatigue and stress, and temperature and radiation exposure were among the most important variables associated with alertness. This method is broadly applicable to environmental studies where the main goal is accurate, individualized prediction involving a mixture of person-level traits and irregularly measured time series.

Danni Tu↗

Supervised Machine Learning Approach for Classifying Earth Science Publications

The data collections archived and distributed by the GES DISC NASA data center are widely utilized for various Earth Science studies. As these collections are created, many research works are published regarding these collections' algorithms, their validation, and their applications. As NASA data centers collect these publications for public use, it is helpful to categorize them based on how they relate to their associated datasets. Specifically, whether the publication linked to the GES DISC dataset is using it for applicational research, describing the algorithm used for the dataset creation, validating the dataset, or providing a general overview of the data collection. Currently, this process requires simple manual labeling, and as such, it may be possible to solve via automation. To approach this problem, machine learning classifiers were developed to predict a publication's category. Manually labeled publications were used as the training data for the supervised machine learning algorithms, specifically Random Forest and Multinomial Naïve Bayes. After balancing the dataset and implementing the Multinomial Naïve Bayes algorithm, the classification accuracy achieved was substantially higher than the baseline accuracy, thus significantly improving the efficiency of publication labeling.

Rohan Dayal↗

Variance Decomposition of MEDLI2 Reconstructed Heating Using Neural Networks

The Mars Entry, Descent, and Landing Instrumentation (MEDLI2) sensor suite collected data during entry of the Mars 2020 Perseverance rover into Mars’ atmosphere. This suite included a network of MEDLI2 Instrumented Sensor Plugs (MISPs). Each MISP was comprised of a cylinder made of Thermal Protection System (TPS) material with 1-3 embedded thermocouples (TCs), and it was flush mounted into the heatshield or backshell. Data from these in-depth TCs were used to reconstruct the aeroheating environment of the vehicle throughout entry. Surface heating was posed as an inverse problem, with the goal of estimating the surface heating by minimizing an objective function of the difference between MISP temperature measurements during flight and the temperature predictions derived from the Fully Implicit Ablation and Thermal response (FIAT) program. Given an aerothermal environment, FIAT calculates the material response and provides in-depth temperatures throughout the TPS material. To achieve the reverse, an internal tool called FIAT_Opt runs through multiple different environments until the output temperature at the TC depth closely matches the flight data. 95% confidence intervals on the reconstructed surface heating were obtained using Monte Carlo analysis, in which uncertainties in the thermocouple depth and the TPS material properties (e.g., density, thermal conductivity, heat capacity, emissivity) based on flight-lot material testing were included. A variance decomposition method using Sobol indices was employed to assess the sensitivity of the reconstructed peak heating to the TC placement and material property uncertainties. Variance decomposition was found to require tens of thousands of FIAT_Opt runs in order for the Sobol indices to converge. With a single FIAT_Opt run taking on the order of 40 minutes, the required number of computations would take months to complete, even if using multiple CPUs. To mitigate this problem, three machine learning models (ridge regression with cross-validation, random forest regression, and a deep neural network) were trained and tested using the 2000 Monte Carlo runs that were already completed. A subset of 1600 runs were used to train the model (i.e., training set), while the remaining 400 runs were used as the test set. The predictions from the deep neural network (DNN) on the test set showed nearly perfect agreement to the actual values computed with FIAT_Opt (R2 > 0.99). Using the DNN as a surrogate model, the variance decomposition using 50,000 runs was completed within minutes. The resulting Sobol indices showed that the reconstructed peak surface heating was most sensitive to the uncertainties in the thermal conductivity (ST = 0.37) and heat capacity (ST = 0.26). This method can be leveraged to provide requirements for material property measurements needed to improve the accuracy of surface heating prediction and ultimately lead to the reduction of design margins in the future. This presentation will include background on the MEDLI2 suite; the method used for inverse heating estimation; the way that material property uncertainties were accounted for using Monte Carlo analysis; a brief background on variance decomposition; the motivation for using machine learning in this context; how a neural network was trained on the data to enable variance decomposition in a fraction of the time; and the variance decomposition results for one of the MISPs.

Hannah Alpert↗

A Regional L-band High Biomass Estimation Framework Leveraging Spaceborne Lidar and Interferometric Data to Overcome Backscatter Saturation

We propose a framework to estimate high above ground biomass (AGB) from L-band SAR imagery leveraging spaceborne lidars such as GEDI or ICESat-2 and repeat-pass coherence. Our results indicate we are able to overcome model saturation typically associated with purely backscatter methodologies. We validate our approach using lidar-derived AGB maps from the AfriSAR datasets at Mondah, Ogooue, and Lope. We apply our framework to UAVSAR and ALOS- 2 imagery to obtain 50 meter resolution biomass maps. We obtain < 60% nRMSE (in some cases much better) with negligible relative bias using a multiscale random forest model. We illustrate that the inclusion of coherence can significantly improve high AGB estimation particularly at the coastal site Mondah.

Liao, Tien-hao↗

Global benefits of non-continuous flooding to reduce greenhouse gases and irrigation water use without rice yield penalty

Non-continuous flooding is an effective practice for reducing greenhouse gas emissions (GHGs) and irrigation water use (IRR) in rice fields. However, advancing global implementation is hampered by the lack of comprehensive understanding of GHGs and IRR reduction benefits without compromising rice yield. Here, we present the largest observational data set for such effects as of yet. By using Random Forest regression models based on 636 field trials at 105 globally georeferenced sites, we identified the key drivers of effects of non-continuous flooding practices and mapped maximum GHGs or IRR reduction benefits under optimal non-continuous flooding strategies. The results show that variation in effects of non-continuous flooding practices are primarily explained by the UnFlooded days Ratio (UFR, that is the ratio of the number of days without standing water in the field to total days of the growing period). Non-continuous flooding practices could be feasible to be adopted in 76% of global rice harvested areas. This would reduce the global warming potential (GWP) of CH4 and N2O combined from rice production by 47% or the total GWP by 7% and alleviate irrigation water use by 25%, while maintaining yield levels. The identified UFR targets far exceed currently observed levels particularly in South and Southeast Asia, suggesting large opportunities for climate mitigation and water use conservation, associated with the rigorous implementation of non-continuous flooding practices in global rice cultivation.

climate change mitigation↗

Grand Canyon Ecological Forecasting: Using NASA Earth Observations to Monitor and Model Juniper Woodland Mortality in Grand Canyon National Park

Significant die-offs of the drought tolerant species Utah juniper (Juniperus osteosperma) and one-seeded juniper (Juniperus monosperma) have been observed throughout central and northern Arizona, including Grand Canyon National Park (GCNP). As climate models project rising temperatures and continuous drought, land managers are concerned for the future of juniper in and around GCNP. This project incorporated data from Landsat 8 Operational Land Imager (OLI), the Shuttle Radar Topography Mission (SRTM), and ocular samples of the National Agriculture Imagery Program (NAIP) in a random forest model to identify patterns between characteristics of the landscape and locations of juniper woodland mortality and to model areas subject to vulnerability. This study found no significant correlation between ocularly sampled juniper tree mortality and remotely sensed environmental variables used thus, accurately modeling mortality vulnerability in the future was not feasible. The ocular sampling, however, allows the partners at the GCNP’s Science & Resource Management Division to better understand areas of juniper tree woodland mortality and the relative amount of mortality in the park. Additionally, the areas of juniper tree mortality found in this project provide the partners with guidance for future field sampling.

Scarlet Jackson↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), has typically limited machine learning (ML) in space studies and further study of radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNAseq) data from 6 mouse liver GeneLab datasets (GLDS) with a total of 113 spaceflight and ground-control samples to determine top features relevant to spaceflight including the effect of radiation exposure. Data was normalized within each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. The top MRMR features were used to predict spaceflight vs. ground-control samples using a Random Forest (RF) classifier with 5-fold cross validation (CV). The ML-based gene sets were further compared against differential gene expression results from individual GLDS. CV training using the top 100 MRMR genes show averages of 86% accuracy and 0.95 AUC value on the validation set over 5 folds (Figure 1A). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 811 or 68 DEGs overlapping between at least 2 or 3 studies, respectively (Figure 1B). Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism. Set analysis between the MRMR features and the DEGs showed 60 or 8 genes overlapping with at least 1 or 2 studies, respectively. MRMR feature selection and ensemble ML methods (e.g. RF) improve performance relative to a Naïve Bayes classifier when NGS data sets are analyzed. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise ratio. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from RNASeq analysis. Non-intersecting sets introduce opportunity to explore spaceflight relevant genes and implementing ML methods across existing NGS datasets may overcome sample size limitations. ML coupled with existing analytical methods enhances understanding of disease by revealing common underlying pathways across datasets.

Machine Learning↗

Yellowstone Ecological Forecasting: Assessing Change in Aspen Extent in Northern Yellowstone National Park

The removal and reintroduction of the gray wolf (Canis lupus) in Yellowstone National Park have played an important role in shaping the ecological composition of this distinct landscape, and it is a textbook example of multi-trophic dynamics. With particular importance to conservation science, the inter-trophic cascades between wolves and species such as the elk (Cervus canadensis) and the quaking aspen (Populus tremuloides) have been extensively studied. In conjunction with the National Park Service, Yellowstone National Park, Utah State University, and the University of Wisconsin–Stevens Point, this project utilized satellite remote sensing to investigate the long-term trends in aspen extent. Through random forest modeling and phenological approaches, Sentinel-2 Multispectral Instrument (MSI; years 2017–2019) and Landsat 5 Thematic Mapper (TM; years 1987–2011) datasets were used to derive an Enhanced Vegetation Index (EVI), a Normalized Difference Vegetation Index (NDVI), Tasseled Cap Indices (Brightness, Greenness, Wetness), and RGB true color composites. The International Space System Global Ecosystem Dynamics Investigation (ISS GEDI) was used to analyze canopy height. Results were consolidated into maps and time-series that provide an in-depth and intricate depiction of aspen stand extent. The end products will assist the National Park Service in its management practices and inform wildlife restoration and rewilding decisions within and beyond the contexts of Yellowstone National Park.

Kyle Steen↗

Advancements in Blowing Dust Detection at Night via Machine Learning

This presentation introduces operational users to a machine-learning based Dust Probability product developed by the NASA SPoRT program for the application of detecting and monitoring blowing dust plumes at night. Advances in earth observing satellites has improved monitoring and detection of dust both day and night through derived imagery such as the Dust RGB. However, limitations of the RGB at night result in less contrast between dust and land surface features, as seen by the user. A Machine Learning (ML) model has been developed and applied to GOES-16 ABI to overcome this limitation and improve nighttime dust detection. The ML capability is a subset of Artificial Intelligence methods. In this case the Dust ML model was developed using a simple Random Forest (RF) model, typically used to solve classification challenges (or to provide regression type output). The goal was to leverage the strengths of the RF model to learn how to identify blowing dust, and hence, overcome the limitation of a user trying to detect blowing dust within the satellite imagery by eye alone. A brief description of the ML model development will be provided. However, the focus of the presentation will be on the initial user feedback from the assessment of this tool for the 2022 blowing dust events of March through April. During this time several users across the U.S. Southwest collaborated to apply this Dust ML product at night as a complement to the existing Dust RGB in order to determine if it provided greater operational efficiency and value.

Machine Learning↗

Yellowstone Ecological Forecasting: Assessing Change in Aspen Extent in Northern Yellowstone National Park

The removal and reintroduction of the gray wolf (Canis lupus) in Yellowstone National Park have shaped the ecological composition of this distinct landscape, representing a textbook example of trophic dynamics. With particular importance to conservation science, researchers have studied the trophic cascades between wolves and species such as elk (Cervus canadensis) and quaking aspen (Populus tremuloides). In conjunction with the National Park Service, Yellowstone National Park, Utah State University, and the University of Wisconsin–Stevens Point, this project utilized satellite remote sensing to investigate the long-term trends in aspen extent. Through random forest modeling and phenological approaches, Landsat 5 Thematic Mapper (TM; years 1986–2011) and Sentinel-2 Multispectral Instrument (MSI; years 2017–2019) datasets were used to derive color composites, Normalized Difference Vegetation Index (NDVI), Enhanced Vegetation Index (EVI), and Tasseled Cap Indices (Brightness, Greenness, Wetness). The International Space System (ISS) Global Ecosystem Dynamics Investigation (GEDI) provided canopy height data. The team consolidated results into maps and time-series which provide an in-depth depiction of aspen stand extent. The National Park Service will use these end products to assist in its management practices and inform wildlife restoration decisions within and beyond Yellowstone National Park.

Kyle Steen↗

What’s That Supposed to Mean? Capturing Micro-Behaviors in Teams

Future long-duration space exploration (LDSE) crews will require extensive coordination, cooperation, and team functioning as they face a myriad of challenges rooted in both taskwork and teamwork (Bell et al., 2015; Landon et al., 2018). While exposed to extreme conditions, crew members must navigate living and working together in prolonged confinement. Moreover, astronaut teams are becoming increasingly diverse, introducing significant variability in team composition. This increasing diversity, alongside traditional constraints of LDSE, introduces additional challenges into effective team functioning. To date, most methods for capturing team functioning rely on self-report measures. Such measures are prone to several limitations, including but not limited to social desirability bias, halo effect, and leniency effects (Trull & Ebner-Priemer, 2013), which skew data and limit nuanced understandings of phenomena at play. Self-report measures broadly capture team functioning, lending the nature of such methods to identifying underlying “macro”-behaviors (i.e., behaviors that are long-standing and last over time). However, team functioning is far more complex than a series of macro-behaviors, rendering reliance on self-report data deficient for accurate measurement. Recent research demonstrates the potential of alternative methods for capturing team functioning, such as speech and physiological data (Chaffin et al., 2017; Murray & Oertel, 2018). Consequently, these methods are more suitable for capturing micro-behaviors: brief, often unconscious expressions that affect the extent to which an individual feels included by others around them (Paletz et al., 2013). Micro-behaviors can be further classified into microaggressions (i.e., subtle, negative exchanges; Keller & Galgay, 2010) or micro-affirmations (i.e., subtle, positive exchanges; Kyte et al. 2020), both of which influence team functioning. Due to the subtle nature of micro-behaviors, contextual factors have a significant impact when determining if it is aggressive or affirmative. Additionally, several iterations of microbehaviors can have lingering effects on team interactions. For example, the use of “mm-hmm” by a crew member can function as both a micro-affirmation and micro-aggression. Specifically, it can be indication of active listening (i.e., micro-affirmation) or as an expression of annoyance (i.e., aggression) depending on the context in which it occurs. Auditory features (e.g., tone, frequency) can help delineate between the two forms; however, the contextual factors (e.g., previous interactions between team members, crew demographics) add a layer of complexity that render auditory features alone as insufficient to capture micro-behaviors. Consequently, this paper seeks to provide a novel approach in which multi-modal data (i.e., auditory features and contextual features) are used in a random-forest model to better identify distinguishing characteristics between micro-affirmations and micro-aggressions. In turn, detected micro-behaviors are used to predict team performance, thereby demonstrating the value of capturing micro-behaviors as supplemental data to macro-behaviors.

Sydney R. Begerowski↗