Search NASA⌕ Search

Engineering topics

Bethany P. Theiling

Publications and source records attributed to Bethany P. Theiling.

Interpretable Machine Learning for Molecular Biosignatures: a Novel Single-Sample Feature Importance Method That Is Sensitive To Statistical Interactions

Isotope ratio mass spectrometry (IRMS) of volatiles (e.g., CO 2 ) promises to be a powerful tool for potential biosignature detection for future missions to ocean worlds (OW) such as Europa and Enceladus. Machine learning (ML) methods for IRMS data could enable science autonomy by onboard prediction of seawater chemistry and biosignature presence. However, ML models are likely to be complex and involve statistical interactions between features (variables), which can make predictions seem opaque and enigmatic. For ML predictions as significant as extraterrestrial biosignatures, we must place extraordinary confidence in models. It is therefore essential that these models make interpretable predictions (i.e., human-understandable) and include false-prediction diagnostics. We achieve high accuracy and interpretability in ML biosignature and seawater chemistry models for OW through a nearest-neighbors feature selection tool that detects statistical interactions between predictors, constructs interaction networks for visualization of selected features working together to make a prediction, and reports single-sample feature importance scores for false-detection diagnostics. Here we develop a novel single-sample nearest-neighbors projected distance regression(ssNPDR) feature selection method that improves upon existing single-sample algorithms through the inclusion of statistical interactions while providing false-prediction diagnostics for ML models.

geochemistry↗

Science Autonomy for Ocean Worlds Astrobiology: A Perspective

Astrobiology missions to ocean worlds in our solar system must overcome both scientific and technological challenges due to extreme temperature and radiation conditions, long communication times, and limited bandwidth. While such tools could not replace ground-based analysis by science and engineering teams, machine learning algorithms could enhance the science return of these missions through development of autonomous science capabilities. Examples of science autonomy include onboard data analysis and subsequent instrument optimization, data prioritization (for transmission), and real-time decision-making based on data analysis. Similar advances could be made to develop streamlined data processing software for rapid ground-based analyses. Here we discuss several ways machine learning and autonomy could be used for astrobiology missions, including landing site selection, prioritization and targeting of samples, classification of “features” (e.g., proposed biosignatures) and novelties (uncharacterized, “new” features, which may be of most interest to agnostic astrobiological investigations), and data transmission.

ocean worlds↗

NASA’s Goddard Space Flight Center’s Distributed Systems Missions Architecture

Space and Earth Science are being transformed by applying a distributed approach to missions, where the fusion of data from components, systems, instruments, models, and observation locations works in concert with timely responses and feedback mechanisms to multiply the knowledge obtained. Additionally, a disaggregated approach allows for a distributed cost and schedule that can be shared across multiple organizations to enable the greater mission. With the advances in reduced size, weight, and power for space-worthy components leading to the revolution in smaller spacecraft, the cost and timeliness proposition for launching multiple space assets has also greatly improved. Thus, the aerospace industry is undergoing a paradigm shift toward a proliferation of small satellites as a networked approach to meet mission objectives. This paper will describe the impetus, goal, and path to provide an openly available framework as a unifying catalyst for broad-ranging Distributed Systems Missions (DSMs) contributors.

Distributed Systems↗

Squeezing Every Last 'Bit' of Information from Enceladus Mass Spectrometry

Potential opportunities to return to Enceladus in Discovery and Flagship class missions inspire development of next-generation instruments and creative approaches to sample collection, sample analysis, and data analysis and transmission strategies. Mass spectrometers (MS) are ideally suited to future Enceladus missions due to their analytical power in identifying a range of molecular and ionic compositions – including complex organics – and potentially astrobiologically-important features such as isotope ratios, chirality, and enantiomeric excess. However, long communication delays from Enceladus and limited bandwidth limits the data transmission from these higher-data-volume instruments, likely delaying mission-related response to new data. We explore the utility of data science and machine learning (ML) on isotope ratio (IR)MS data collected from laboratory analogs of Enceladus to: 1) process data quickly for rapid ground-based analyses, 2) understand if compositional and biosignature information could be extracted from IRMS data, and 3) evaluate whether onboard ML techniques could improve sample analysis, cadence, and transmission prioritization. Laboratory analogs analyzed isotopes of volatile CO2 that interacted with seawaters of varying composition, and include both abiotic and biotic (microbially-influenced) experiments. Enceladus’s alkaline oceans promote speciation of carbon into multiple forms (e.g., H2CO3 / CO2, HCO3-, and CO32-), each of which could be isotopically fractionated by abiotic or biotic reactions. Large (>2‰) changes in carbon isotopes (δ13C) are observed from some biotic experiments inoculated with complex microbial ecosystems relative to the abiotic seawaters. ML training and classification suggests that microbial samples can be distinguished from abiotic samples, yet that a broad range of microbial experiments are necessary to train ML models to cover a range of complexities including disequilibria, and isotopic and compositional fractionation.

geochemistry↗

Extraterrestrial Molecular Indicators of Life Investigation (EMILI)

Future missions to Enceladus, Europa, Mars, and beyond may seek the molecular signs of extraterrestrial life through chemical analysis of acquired samples. Particularly on ocean worlds such as Enceladus and Europa, samples may contain trace ocean-borne molecular biosignatures of extant life that may or may not share similarities to those of terrestrial life. In situ analyses must be prepared to detect and characterize a wide range of possible molecular species, structures, and patterns, typically with exquisite sensitivity and within a complex, poorly-characterized planetary environment. The Extraterrestrial Molecular Indicators of Life Investigation (EMILI) is designed to meet or exceed the requirements of such missions for organic molecular analysis through a powerful combination of dual chemical separation and both optical and mass spectrometry detection techniques, realized in an integrated, compact instrument package fully compatible with anticipated flight resources and conditions. The full EMILI instrument combines two sample analysis subsystems to provide wide-ranging and complementary detection of organic compounds and inorganic salts. The Gas Analysis Processing System (GAPS) uses a chemical derivatization protocol with gas chromatography (GC) separation prior to detection in an ion trap mass spectrometer (ITMS) to enable full characterization of lower-polarity, volatile and semi-volatile molecules such as fatty acids and hydrocarbons. The Organic Capillary Electrophoresis ANalysis System (OCEANS) uses a liquid-based extraction protocol with CE separation to enable precise analysis of more water-soluble/polar compounds. OCEANS features a laser-induced fluorescence detection mode to perform ultra-sensitive quantitative analysis of chiral amino acids. In EMILI, OCEANS is additionally coupled to the same ITMS through a novel electrospray ionization interface. The common ITMS allows EMILI to identify and cross-correlate molecular species and patterns, detected through either or both protocols, of molecular weights to over 1000 u, potentially even revealing complex biosignatures such as alien oligopeptides and informational polymers.

Europa↗

Interpretable Machine Learning Models for Autonomous Characterization of Analogue Ocean World Seawater Chemistry and Biosignature Potential Using Isotope Ratio Data

Background: Future missions to ocean worlds, such as Enceladus and Europa, will attempt to characterize the subsurface seawater chemistry and assess the potential for life. Such missions will be equipped with capabilities to precisely measure volatile isotopes in plumes, atmospheres, and exospheres. Motivation: While large isotopic fractionations can indicate a biological source, there are signatures resulting from abiotic geochemical processes that mimic isotopic biosignatures. While machine learning (ML) has the potential to disentangle competing effects and biotic mimicry, high-dimensional isotope ratio mass spectrometry (IRMS) data is likely to contain noise/irrelevant features and involve complex statistical interactions that make human inference and interpretation difficult. Further, ML predictions with as far-reaching implications as an extraterrestrial biosignature on an ocean world requires the use of interpretable models (i.e., not “black box” models) with physically and mathematically meaningful feature spaces along with false positive diagnostics. Methods: We use volatile CO2 IRMS data of analogue ocean world seawaters to validate an ML approach to provide biogeochemical context for biosignature detection. We employ a feature selection method called nearest-neighbor projected distance regression (NPDR) that detects statistical interactions and helps elucidate the mechanisms of the Random Forest classification models. Results: We train and validate predictive ML models on volatile CO2 IRMS data of analogue ocean world seawaters to predict major salt components (e.g., MgSO4, NaHCO3), pH, ionic strength, and the presence of biosignatures. Features derived from IRMS measurements are augmented with extracted time-series features. Our results show high test accuracy and interpretability, which is increased by interaction network visualization, sample-wise variable importance scores, and single-sample class probability estimates. We demonstrate an ML mission software solution that triggers autonomous data transmission and biogeochemical sample prediction.

geochemistry↗

A Science-Focused Artificial Intelligence (AI) Responding in Real-Time to New Information: Capability Demonstration for Ocean World Missions

Introduction: Artificial intelligence (AI) has long been considered a potential mechanism to explore increasingly challenging environments, including those with extreme temperatures and pressures, limited communication capabilities, or those with demanding terrain. We posit that missions in extreme environments could deploy an onboard AI focused on science observations and goals in order to augment a traditional concept(s) of operations (ConOps). An onboard AI capability could perform functions such as data analysis in order to make high-level decisions, including prioritized data transmission for analysis by ground-based teams or autonomously-guided follow-on analyses that maximize science return. Such a capability would empower missions to respond to scientific data of interest in real-time; a mission could make observations and perform a preliminary analysis to alert ground-based scientists to an observation of interest, enabling an informed, rapid response from Earth-based teams. Enceladus Case Study for Onboard AI: We are developing an onboard AI capability for real-time telemetry response that formulates and carries-out informed decisions in service to established mission goals, enabling increased science return of a mission. We focus our AI development for use on a constellation of SmallSats orbiting Enceladus. Our Enceladus case study tests autonomous decision-making capabilities in scenarios with complex orbital dynamics, plume ejecta, extreme cold environments, power restrictions, and a requirement to maximize science return for a potential positive detection of life, while critically evaluating the potential for false positives. Telemetry includes simulated scientific data, spacecraft onboard operational data (e.g., position, velocity, and rotation), and engineering hardware performance data. Enceladus SmallSat Constellation. Our constellation includes eight SmallSat spacecraft in an 8:35 resonant orbit-based formation, leveraging Saturn’s gravitational forces to maintain stable orbits with global coverage around Enceladus. To our knowledge, we simulate the first stable configuration of multiple spacecraft in closed orbits around Enceladus, using a full ephemeris force model (Russell and Lara, 2009). Each spacecraft’s orbit will precess, causing an eastward ground track shift (from an orbiter’s perspective) of each spacecraft for each orbit. However, all spacecraft return to their original positions relative to Enceladus after eight Enceladus revolutions around Saturn. We model communication pathways between SmallSats to understand how information would need to be transmitted across the constellation to enable AI-driven decision-making and resource allocation across the fleet. Capability Demonstration. Our simulated capability demonstration inputs position, velocity, and rotation telemetry from our Enceladus-focused constellation simulations, and mass spectrometry data collected from abiotic and biotic laboratory-analog ocean world experiments (Theiling et al., 2018; Theiling, 2021; Da Poian et al., 2023). Data from these experiments are used to simulate MS measurements and different scenarios of science observations for onboard analysis performed on each of the eight spacecraft. For these demonstrations, we integrate 24 machine learning (ML) algorithms into an onboard intelligence as a ‘knowledge base’, including algorithms evaluating data quality and those predicting (with % confidence) gas composition, ocean aqueous chemistry, and whether the sample was influenced by microbial life. The onboard AI capability is designed to use the knowledge base to come to a consensus-based decision in the interpretation of the observed data in order to request additional action outside of a pre-defined ConOps. Requested actions could include e.g., prioritized downlink to Earth (for analysis by ground-based teams) or follow-on analyses performed across the constellation. The spacecraft’s intelligent onboard planner must then determine whether sufficient resources (e.g., time, power, etc.) are available and weigh the request with mission priorities. In our simulation, the constellation is able to identify potential biosignatures using onboard ML algorithms, evaluate the confidence of that prediction, and perform follow-on analyses across the fleet to confirm the detection, in order to best prepare a transmission of these data to Earth-based teams.

astrobiology↗