Search NASASearch

Engineering topics

Alaa Khader

Publications and source records attributed to Alaa Khader.

A Preliminary Study on the Feasibility of Large Language Models for Detecting Micro-Behaviors Among Team Members in Space Missions

Large-language models (LLMs) have been recently used for spoken language understanding (SLU) to infer meaning and semantics from speech in tasks such as speaker intent and sentiment classification. Due to being trained on large amounts of data, and their ability to understand context and relationships between words, LLMs are competent, enabling them to generalize across tasks without requiring many task-specific training samples. This research examines the feasibility of few-shot learning in LLMs for detecting subtle, brief, and possibly unconscious interactions between team members, called ``micro-behaviors," and provides insights into the appropriate design of LLMs for this task. Our data came from 5 teams participating in a 45-day mission at the US National Aeronautics and Space Administration’s (NASA) Human Exploration Research Analog (HERA). More specifically we used data collected from team interaction battery (TIB) tasks teams performed five times in-mission which comprise an average 1.5 hours of conversation data per day. Micro-behaviors were coded according to an adapted version of Smith & Griffins (2022) theoretical framework in terms of Violation (i.e., presence of valenced behavior, uplifting/positive or discouraging/negative), Intensity (i.e., force of behavior in terms of how uplifting or discouraging is the behavior), and Intent (i.e., motive of the behavior in terms of whether it was deliberate or unintentional). We explore the ability of LLMs to detect the presence and intensity of micro-behaviors. We examine employing and fine-tuning readily available LLMs (i.e., RoBERTa, DistilBERT), as well as prompting state-of-the-art sequence classification models (i.e., Llama-2, Llama-3). In a total of 13,058 conversational turns (17.8% uplifting, 3.3% discouraging, 75.76% neutral, 3.14% nulls), we compute the macro F1-score of the 3-way micro-behavior classification task (i.e., classifying among uplifting, discouraging, and neutral; 33% chance). Results indicate that the RoBERTa model achieves a F1-score of 36.2% (uplift: 43.3% precision (P), 15.1% recall (R); discourage: 20% P, 0.5% R). These results significantly improve when we augment the data via paraphrasing in the RoBERTa model, reaching a 41.2% macro F1-score (uplift: 37.7% P, 86.3% R; discourage: 3.5% P, 1.8% R). Finally, the Llama-2 model with 3-shot prompting yields 38% macro F1-score (uplift: 28.7% P, 20% R; discourage: 7.2% P, 18% R), which is slightly better compared to the RoBERTa model without data augmentation, highlighting the effectiveness of sequence classification models in detecting minority classes with a small sample size. Findings indicate that LLMs hold potential to detect subtle behaviors in conversations, which could be valuable in assessing team behavior in space exploration missions. Future studies will evaluate the performance of different LLM prompting strategies or fine-tuning methods.

Ankush Raut

Navigating Team Dynamics: Automated Detection of Micro-Behaviors Between Team Members Through Longitudinal Interaction Data

The success in future long term space exploration missions will depend on the cooperation, coordination, and mutual understanding among the crew members. Micro-behaviors are momentary, subtle linguistic and paralinguistic indicators of thinking and feeling toward another member of the team (Cortina et al., 2001; Smith & Griffiths, 2022) that can significantly impact team dynamics and influence the overall team performance (Paromita & Chaspari, 2024). Due to their interactive nature, micro-behaviors have a sender (i.e., the team member expressing the micro-behavior) and a target (the team member impacted by the micro-behavior). Detection of these behaviors can assist in avoiding possible conflict among crew members and promoting the overall team success. Our prior research focused on an initial proof of concept of machine learning (ML) models and natural language processing (NLP) techniques that were used for automatically detect micro-behaviors among crew members of the US National Aeronautics and Space Administration’s (NASA) Human Exploration Research Analog (HERA) Campaigns 4 and 5 missions (Paromita et al., 2023). Results underscored the importance of incorporating contextual information in the ML models in the form of sentiment analysis, type of task, and dyadic interaction among team members. Here, we expand the scope of our prior work in two ways. First, we assess ML/NLP methods on new behavioral annotations coded using an adapted version of Smith & Griffins (2022) theoretical framework in terms of Violation (i.e., presence of valenced behavior, uplifting/positive or discouraging/negative), Intensity (i.e., force of behavior in terms of how uplifting or discouraging is the behavior), and Intent (i.e., motive of the behavior in terms of whether it was deliberate or unintentional). Second, we expand the design of the ML model to preserve information about the role of each team member within the occurrence of the micro-behavior (in contrast to the previous model that only considered the sender and the target without determining the team member role). This allows to consider all team members' contributions in the conversation and model long-term dependencies in the dialogue. Our experiments for this study are conducted on data from 5 teams of the NASA HERA C4 (NASA grant NNX16AQ48G (PI: Bell)). Conversations were extracted from the 1.5 hour Team Interaction Battery (TIB) task that occurred 5 times in-mission per crew. This resulted in a total of 13,058 conversational turns (i.e., 17.8% uplifting, 3.3% discouraging, 75.76% neutral, 3.14% nulls). Our findings with the revised behavioral coding and ML/NLP models indicate a 43.66% macro F1-score (i.e., 38.29% precision (P), 50.8% recall (R)) for a dialog state-tracking model that includes information from the sender only, and a 40.9% F1-score (i.e., 38.7% P, 43.36% R) for the same model that includes information from both the sender and the target of the micro-behavior. These are significantly higher compared to simple random forest models that classify behaviors strictly based on speech content and do not consider iterative team dynamics, achieving a 36.07% F1-score (i.e., 39.04% R, 33.53% P). Our findings demonstrate potential ways to leverage large conversational datasets to better capture complex team dynamics. We will discuss future directions including proposed models that can incorporate additional mission days and tasks beyond the TIB for objectively quantifying team behavior at high temporal resolution in space exploration missions.

Projna Paromita

What’s That Supposed to Mean? Capturing Micro-Behaviors in Teams

Future long-duration space exploration (LDSE) crews will require extensive coordination, cooperation, and team functioning as they face a myriad of challenges rooted in both taskwork and teamwork (Bell et al., 2015; Landon et al., 2018). While exposed to extreme conditions, crew members must navigate living and working together in prolonged confinement. Moreover, astronaut teams are becoming increasingly diverse, introducing significant variability in team composition. This increasing diversity, alongside traditional constraints of LDSE, introduces additional challenges into effective team functioning. To date, most methods for capturing team functioning rely on self-report measures. Such measures are prone to several limitations, including but not limited to social desirability bias, halo effect, and leniency effects (Trull & Ebner-Priemer, 2013), which skew data and limit nuanced understandings of phenomena at play. Self-report measures broadly capture team functioning, lending the nature of such methods to identifying underlying “macro”-behaviors (i.e., behaviors that are long-standing and last over time). However, team functioning is far more complex than a series of macro-behaviors, rendering reliance on self-report data deficient for accurate measurement. Recent research demonstrates the potential of alternative methods for capturing team functioning, such as speech and physiological data (Chaffin et al., 2017; Murray & Oertel, 2018). Consequently, these methods are more suitable for capturing micro-behaviors: brief, often unconscious expressions that affect the extent to which an individual feels included by others around them (Paletz et al., 2013). Micro-behaviors can be further classified into micro-aggressions (i.e., subtle, negative exchanges; Keller & Galgay, 2010) or micro-affirmations (i.e., subtle, positive exchanges; Kyte et al. 2020), both of which influence team functioning. Due to the subtle nature of micro-behaviors, contextual factors have a significant impact when determining if it is aggressive or affirmative. Additionally, several iterations of micro-behaviors can have lingering effects on team interactions. For example, the use of “mm-hmm” by a crew member can function as both a micro-affirmation and micro-aggression. Specifically, it can be indication of active listening (i.e., micro-affirmation) or as an expression of annoyance (i.e., aggression) depending on the context in which it occurs. Auditory features (e.g., tone, frequency) can help delineate between the two forms; however, the contextual factors (e.g., previous interactions between team members, crew demographics) add a layer of complexity that render auditory features alone as insufficient to capture micro-behaviors. Consequently, this paper seeks to provide a novel approach in which multi-modal data (i.e., auditory features and contextual features) are used in a random-forest model to better identify distinguishing characteristics between micro-affirmations and micro-aggressions. In turn, detected micro-behaviors are used to predict team performance, thereby demonstrating the value of capturing micro-behaviors as supplemental data to macro-behaviors.

Sydney Begerowski

Meaningful Work as a Resilience Countermeasure in Extreme Environments

Evaluate risk factors, biomarkers, and countermeasures for adaptation and resilience in ICE environments. Identify how meaningful work influences the relationship between risk factors, the valence and social process domains, and operational and performance outcomes. Develop an operationally acceptable and valid measure of meaningful work in ICE. Examine meaningful work as a countermeasure.

Lauren Blackwell Landon