Search NASA⌕ Search

SEARCH · Search NASA

Results for “Machine Learning Models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35

Parametric Analysis of a Hover Test Vehicle using Advanced Test Generation and Data Analysis

Large complex aerospace systems are generally validated in regions local to anticipated operating points rather than through characterization of the entire feasible operational envelope of the system. This is due to the large parameter space, and complex, highly coupled nonlinear nature of the different systems that contribute to the performance of the aerospace system. We have addressed the factors deterring such an analysis by applying a combination of technologies to the area of flight envelop assessment. We utilize n-factor (2,3) combinatorial parameter variations to limit the number of cases, but still explore important interactions in the parameter space in a systematic fashion. The data generated is automatically analyzed through a combination of unsupervised learning using a Bayesian multivariate clustering technique (AutoBayes) and supervised learning of critical parameter ranges using the machine-learning tool TAR3, a treatment learner. Covariance analysis with scatter plots and likelihood contours are used to visualize correlations between simulation parameters and simulation results, a task that requires tool support, especially for large and complex models. We present results of simulation experiments for a cold-gas-powered hover test vehicle.

Gundy-Burlet, Karen↗

Using Machine Learning to Identify Novel Hydroclimate States

Anthropogenic climate change is expected to alter drought risk in the future. However, droughts are not uncommon or unprecedented, as documented in tree-ring-based reconstructions of the summer average Palmer drought severity index (PDSI). Using an unsupervised machine-learning method trained on these reconstructions of pre-industrial climate, we identify outliers: years in which the spatial pattern of PDSI is unusual relative to ‘normal' variability. We show that in many regions, outliers are more frequently identified in the twentieth and twenty-first centuries. This trend is more pronounced when the regional drought atlases are combined into a single global dataset. By definition, outlier patterns at the 10% level are expected to occur once per decade, but from 1950 to 2000 more than 6 years per decade are identified as outliers in the global drought atlas (GDA). Extending the GDA through 2020 using an observational dataset suggests that anomalous global drought conditions are present in 80% of years in the twenty-first century. Our results indicate, without recourse to climate models, that the world is more frequently experiencing drought conditions that are highly unusual in the context of past natural climate variability.

Drought risk↗

Squeezing Every Last 'Bit' of Information from Enceladus Mass Spectrometry

Potential opportunities to return to Enceladus in Discovery and Flagship class missions inspire development of next-generation instruments and creative approaches to sample collection, sample analysis, and data analysis and transmission strategies. Mass spectrometers (MS) are ideally suited to future Enceladus missions due to their analytical power in identifying a range of molecular and ionic compositions – including complex organics – and potentially astrobiologically-important features such as isotope ratios, chirality, and enantiomeric excess. However, long communication delays from Enceladus and limited bandwidth limits the data transmission from these higher-data-volume instruments, likely delaying mission-related response to new data. We explore the utility of data science and machine learning (ML) on isotope ratio (IR)MS data collected from laboratory analogs of Enceladus to: 1) process data quickly for rapid ground-based analyses, 2) understand if compositional and biosignature information could be extracted from IRMS data, and 3) evaluate whether onboard ML techniques could improve sample analysis, cadence, and transmission prioritization. Laboratory analogs analyzed isotopes of volatile CO2 that interacted with seawaters of varying composition, and include both abiotic and biotic (microbially-influenced) experiments. Enceladus’s alkaline oceans promote speciation of carbon into multiple forms (e.g., H2CO3 / CO2, HCO3-, and CO32-), each of which could be isotopically fractionated by abiotic or biotic reactions. Large (>2‰) changes in carbon isotopes (δ13C) are observed from some biotic experiments inoculated with complex microbial ecosystems relative to the abiotic seawaters. ML training and classification suggests that microbial samples can be distinguished from abiotic samples, yet that a broad range of microbial experiments are necessary to train ML models to cover a range of complexities including disequilibria, and isotopic and compositional fractionation.

geochemistry↗

GeoAI advances in specific landform mapping

Landform mapping (also referred to as geomorphology or geomorphometry) can be divided into two domains: general and specific (Evans 2012). Whereas general landform mapping categorizes all elements of the study area into landform classes, such as ridges, valleys, peaks, and depressions, the mapping of specific landforms requires the delineation (even if fuzzy) of individual landforms. The former is mainly driven by physical properties such as elevation, slope, and curvature. The latter, however, must consider the cognitive (human) reasoning that discriminates individual landforms in addition to these physical properties (Arundel and Sinha 2018). Both mapping forms are important. General geomorphometry is needed to understand geological and ecological processes and as boundary layer input to climate and environmental models. Specific geomorphometry supports such activities as disaster management and recovery, emergency response, transportation, and navigation. In the United States, individual landforms of interest are named in the U.S. Geological Survey (USGS) Geographic Names Information System, a point dataset captured specifically to digitize geographic names from the USGS Historical Topographic Map Collection (HTMC). Named landform extent is represented only by the name placement in the HTMC. Recent work has investigated CNN-based deep learning methods to capture these extents in machine-readable form. These studies first relied on physical properties (Arundel et al. 2020) and then included the HTMC as a band in RGB images in limited testing (Arundel et al. 2023). Results from the HTMC dataset surpassed those using just physical properties and using the HTMC alone performed best due to the hillshading and elevation (contour) data incorporated into the topographic maps. However, results fell short of an operational capacity to map all named landforms in the United States. Thus, our current work expands upon past research by focusing on the HTMC and physical information as inputs and the named landform label extents. Specifically, we propose to leverage pre-trained foundation models for segmentation and optical character recognition (OCR) models to jointly map landforms in the United States. Our approach aims to bridge the disparities among the independent information sources to facilitate informed decision-making. The modeling pipeline performs (1) segmentation using the physical information and (2) information extraction using OCR, in parallel. Then a computer vision approach merges the two branches into a labeled segmentation. References: Arundel, Samantha T., Wenwen Li, and Sizhe Wang. 2020. “GeoNat v1.0: A Dataset for Natural Feature Mapping with Artificial Intelligence and Supervised Learning.” Transactions in GIS 24 (3): 556–72. https://doi.org/10.1111/tgis.12633. Arundel, Samantha T, and Gaurav Sinha. 2018. “Validating GEOBIA Based Terrain Segmentation and Classification for Automated Delineation of Cognitively Salient Landforms BT - Proceedings of Workshops and Posters at the 13th International Conference on Spatial Information Theory (COSIT 2017).” In Proceedings of Workshops and Posters at the 13th International Conference on Spatial Information Theory (COSIT 2017), Lecture Notes in Geoinformation and Cartography, edited by Paolo Fogliaroni, Andrea Ballatore, and Eliseo Clementini, 9–14. Cham: Springer International Publishing. Arundel, Samantha T., Gaurav Sinha, Wenwen Li, David P. Martin, Kevin G. McKeehan, and Philip T. Thiem. 2023. “Historical Maps Inform Landform Cognition in Machine Learning.” Abstracts of the ICA 6 (August): 1–2. https://doi.org/10.5194/ica-abs-6-10-2023. Evans, Ian S. 2012. “Geomorphometry and Landform Mapping: What Is a Landform?” Geomorphology 137 (1): 94–106. https://doi.org/10.1016/j.geomorph.2010.09.029.

machine learning↗

Introduction to Fuzzy Set Theory

An introduction to fuzzy set theory is described. Topics covered include: neural networks and fuzzy systems; the dynamical systems approach to machine intelligence; intelligent behavior as adaptive model-free estimation; fuzziness versus probability; fuzzy sets; the entropy-subsethood theorem; adaptive fuzzy systems for backing up a truck-and-trailer; product-space clustering with differential competitive learning; and adaptive fuzzy system for target tracking.

Kosko, Bart↗

Machine Vision based Sample-Tube Localization for Mars Sample Return

A potential Mars Sample Return (MSR) architecture is being jointly studied by NASA and ESA. As currently envisioned, the MSR campaign consists of a series of 3 missions: sample cache, fetch and return to Earth. In this paper, we focus on the fetch part of the MSR, and more specifically the problem of autonomously detecting and localizing sample tubes deposited on the Martian surface. Towards this end, we study two machine-vision based approaches: First, a geometrydriven approach based on template matching that uses hardcoded filters and a 3D shape model of the tube; and second, a data-driven approach based on convolutional neural networks (CNNs) and learned features. Furthermore, we present a large benchmark dataset of sample-tube images, collected in representative outdoor environments and annotated with ground truth segmentation masks and locations. The dataset was acquired systematically across different terrain, illumination conditions and dust-coverage; and benchmarking was performed to study the feasibility of each approach, their relative strengths and weaknesses, and robustness in the presence of adverse environmental conditions.

Detry, R.↗

Enhancing Air Traffic Control Planning with Automatic Speech Recognition

The decisions made during the Federal Aviation Administration Air Traffic Control System Command Center's planning teleconferences hold significant sway over the National Airspace System. Held every two hours, these teleconferences convene air traffic managers and stakeholders from across the nation to discuss airspace conditions, weather, and constraints, leading to the formulation and adjustment of traffic management initiatives. Given the critical nature of these decisions, the need for accurate and efficient record-keeping is paramount. In recent years, the application of automatic speech recognition has gained popularity across diverse industries, including aviation. While traditional applications focus on transcribing air traffic control communication, this paper explores a unique application of automatic speech recognition by converting the audio from planning teleconferences into text transcriptions. This innovative approach addresses key challenges in the field, presenting potential benefits for quality assurance, real-time participation, and downstream natural language processing tasks. A notable breakthrough in the machine learning community, namely the transformer neural network architecture, forms the backbone of the proposed solution in this paper. The transformer architecture's role in this research represents a paradigm shift in the efficiency of automatic speech recognition models. By reducing the amount of in-domain training data required, this architecture allows for the fine-tuning of such models like Whisper, originally pretrained on vast English speech datasets. The adaptability of the transformer architecture proves invaluable in capturing the nuances of aviation terminology and specific language used in planning teleconferences. Leveraging the Whisper model as a baseline, our research details the fine-tuning and validation using a dataset comprising 20 hours of meticulously transcribed planning teleconferences. Notably, the baseline pretrained Whisper model exhibited a word error rate of 18.77%. Through the fine-tuning process, the model achieved a substantial improvement, demonstrating an impressive performance with a reduced word error rate of 6.82%. This substantial decrease in WER not only highlights the effectiveness of the transformer architecture but also emphasizes the practical advancements achieved through the application of automatic speech recognition in this specific domain. The utilization of automatic speech recognition in planning teleconferences in this work introduces several novelties. Firstly, the creation of text transcriptions offers a valuable tool for quality assurance and facilitates the efficient review of teleconferences. This is an important aspect of the proposed solution, given the time-sensitive and high-stakes nature of decisions made during these meetings. Furthermore, text-searchable transcriptions provide a streamlined approach for locating and validating critical information, potentially saving hours of manual effort in searching through audio recordings. Moreover, our research identifies a key use case for external facilities and stakeholders. In situations where attendance at the planning teleconference is not feasible, having access to text transcriptions in real-time or shortly after the teleconference ends, proves to be a time-saving and informative resource. This feature enhances collaboration and ensures that stakeholders can stay abreast of important discussions and decisions even in their absence. Despite the efficiency gains facilitated by the transformer architecture in automatic speech recognition technology, it is essential to acknowledge the human factors in data creation. Subject matter experts play a crucial role in accurately transcribing planning teleconferences due to the specificity and complexity of the information discussed. The research dataset, consisting of 20 hours of transcribed planning teleconferences, forms the foundation for fine-tuning and validating the Whisper model. The achieved word error rate of 6.82% demonstrates promising advancements, particularly in recognizing essential aviation terminology within the teleconferences. In conclusion, this paper presents a comprehensive exploration of the application of automatic speech recognition in Air Traffic Control System Command Center planning teleconferences, leveraging the transformer architecture for enhanced efficiency. The novel contributions lie in the improved accessibility of decision-making records, real-time participation opportunities for external stakeholders, and the potential for downstream natural language processing advancements. As the aviation industry continues to evolve, the integration of automatic speech recognition technologies holds the promise of revolutionizing decision-making processes and contributing to the overall safety and efficiency of air traffic management.

ATM↗

Automatic Detection and Classification of Aurora in THEMIS All‐Sky Images

We report a novel machine-learning algorithm for automatically detecting and classifying aurora in all–sky images (ASI) that is largely trained without requiring ground–truth labels. By including a small number of labeled images, we are able to automatically label all of the approximately 700 million images in the Time History of Events and Macroscale Interactions during Substorms (THEMIS) ASI data set from 2008 to 2022. We use a two–stage approach. In the first stage, we adapt the Simple framework for Contrastive Learning of Representations (SimCLR) algorithm to learn latent representations of THEMIS all–sky images. We then finetune a classifier network on the latent representations our model learns of the manually labeled Oslo aurora THEMIS (OATH) data set. We demonstrate that this two–stage approach achieves excellent classification results on data for which there is no current ML classification benchmark. The outcome of this work will facilitate efficient information retrieval for researchers interested in specific categories of aurora and will enable large scale statistical studies and machine learning analyses of THEMIS all–sky images that have not previously been possible. To demonstrate possible ways to utilize this database, we performed a statistical analysis of the occurrence rates of auroral labels with respect to solar wind parameters, interplanetary magnetic field vector, and geomagnetic indices. We further investigate the occurrence rates of auroral phenomena in the annotated data set and their geoeffectiveness by utilizing the co–located THEMIS ground magnetometer data set.

Jeremiah W Johnson↗

Aurora Detection From Nighttime Lights for Earth and Space Science Applications

This research leverages data from the Day/Night Band (DNB) of the Visible Infrared Imaging Radiometer (VIIRS) instrument onboard the Suomi National Polar-orbiting Partnership (S-NPP) satellite. We demonstrate the value of mining the VIIRS DNB for aurora and describe our use of unsupervised machine learning to create a binary mask for aurora occurrence. This mask can be used to flag aurora-contaminated observations for NASA's nighttime lights products for Earth science applications. The identification of auroral regions can also be used for Space Weather applications, for example, for comparison with aurora forecast model and with other satellite- or ground-based aurora observations. The DNB is a broadband channel that is sensitive to wavelengths from 500 to 900 nm, which covers most of the visible light spectrum, and as the name implies, captures light even at night with a sensitivity at the nanowatt level. This band is suitable for aurora observations since the light emitted by the aurora tends to be dominated by emissions from atomic oxygen, resulting in a greenish glow at a wavelength of 557.7 nm, especially at an altitude of 110 km. This study compares the global nighttime derived aurora regions for 17 and 18 March with the NOAA Space Weather Prediction Center's (SWPC) probability product for the St. Patrick's Day geomagnetic storm in 2015. VIIRS sensors are slated to be added to the next generation of polar-orbiting operational satellites. Our novel automated approach to aurora identification opens up an efficient way to leverage this unique data source.

Aurora↗

Detection of Chlorophyll and Leaf Area Index Dynamics from Sub-weekly Hyperspectral Imagery

Temporally rich hyperspectral time-series can provide unique time critical information on within-field variations in vegetation health and distribution needed by farmers to effectively optimize crop production. In this study, a dense time series of images were acquired from the Earth Observing-1 (EO-1) Hyperion sensor over an intensive farming area in the center of Saudi Arabia. After correction for atmospheric effects, optimal links between carefully selected explanatory hyperspectral vegetation indices and target vegetation characteristics were established using a machine learning approach. A dataset of in-situ measured leaf chlorophyll (Chll) and leaf area index (LAI), collected during five intensive field campaigns over a variety of crop types, were used to train the rule-based predictive models. The ability of the narrow-band hyperspectral reflectance information to robustly assess and discriminate dynamics in foliar biochemistry and biomass through empirical relationships were investigated. This also involved evaluations of the generalization and reproducibility of the predictions beyond the conditions of the training dataset. The very high temporal resolution of the satellite retrievals constituted a specifically intriguing feature that facilitated detection of total canopy Chl and LAI dynamics down to sub-weekly intervals. The study advocates the benefits associated with the availability of optimum spectral and temporal resolution spaceborne observations for agricultural management purposes.

Houborg, Rasmus↗

Evaluating Meteorological Dust Events and Machine-Learning Based Dust Identification in Geostationary Satellite Imagery

NASA scientists in the Short-term Prediction Research and Transition Center (SPoRT) developed a physically-based machine learning approach to identify dust in satellite imagery with a focus on night-time dust detection (Berndt et al. 201; DustTracker-AI). NASA/NOAA Geostationary Environmental Operational Satellite-16 (GOES-16) imagery was used for training and model inputs. The training, testing and validation data set consists of 28 events in the Southwest United States, capturing dust and null events in the region from 2018-2020.With 83 distinct images and millions of pixels a random forest model was trained and validated, correctly labeling 85% of dust pixels.For the first time, the model was run in near-real time production during the spring of 2022 and dust probability visualizations were made available to NOAA National Weather Service (NWS) forecasters to assess its utility for dust forecasting. Results indicated the model helped increase the confidence in the presence of dust and enabled dust tracking for a longer period of time into the night-time hours. Forecaster assessment and running the model in near real-time allowed for the team to determine the types of events missed, captured, and false alarms. To gain additional context on model performance,the SPoRT team sought to gather more detailed information on the training database(e.g., meteorological characteristics and drivers). The goal of this project was to identify the meteorological drivers for the dust events and create a database which synthesized information from observations, forecaster discussions, and analyses pertaining to the dust events to understand the types of events currently used to train the model. A more detailed meteorological synopsis was created for each dust event in the training, testing, and validation datasets. Following the completion of the database and documentation, the classification details revealed that 88% of the dust events were synoptically driven while mesoscale events were less prevalent in model datasets. Meteorological conditions found such as mixing layer depth and wind velocity had mean values of 645mb and 21kt respectively.With conditions of deep mixed layers and moderate to strong surface winds a mesoscale thunderstorm outflow event was considered and subsequently added to the model training data set to test the impact of additional mesoscale training data. The model was retrained and then qualitatively tested on a sample thunderstorm outflow case that the original model was unable to identify. Preliminary results showed potential that the addition of more mesoscale events included in the training data could help to better identify indistinct and localized dust events.

Connor Welch↗

Parameterization of Vertical Cloud Distribution from C3M and MERRA Data Using ML Method

Clouds play a key role in regulating the hydrological cycle and the Earth's radiative energy budget. However, global climate models (GCMs) with a horizontal grid spacing on the order of 100 km have limitations in representing sub-grid cloud dynamics with spatial scales on the order of 1 km, leading to potential uncertainties in cloud radiative feedback on the global scale. In our research, we will leverage the capabilities of Deep Machine Learning (DML) methods to construct parameterizations of sub-grid volumetric cloud fraction (VCF), which is the frequency of occurrence on a grid volume accumulated in the horizontal and vertical directions. Our investigation delves into the intricate relationship between VCF obtained from the NASA CALIPSO-CloudSat-CERES-MODIS (CCCM) satellite observation data and 3-D MERRA-2 reanalysis meteorological profiling data (e.g., wind, relative humidity, temperature). Through a comprehensive one-year data training utilizing the Sequence to Sequence DML method, we have successfully disentangled the complicated cloud formation dynamics across diverse meteorological conditions through a day-to-day analysis framework. Preliminary findings reveal promising statistical agreements in geographical and vertical distributions and seasonal variations of volumetric cloud fraction between ML prediction and satellite measurements. These results underscore the aptitude of our DML model to discern underlying cloud physical processes and accurately represent sub-grid cloud formation dynamics. Additionally, we have also employed trained neural network to analyze uncertainties arising from errors in meteorological data, further enhancing the robustness of our VCF parameterization.

Shan Zeng↗

Using Machine Learning to Infer Pre-Entry Properties for Asteroid Threat Analysis

Accurately assessing asteroid threats relies on knowledge of the asteroid’s pre-entry properties such as size, velocity, and mass. Directly measuring these properties can be infeasible due to the sparsity of events and the accuracy and fidelity of various sensors. Current analysis of an asteroid’s pre-entry properties involves modeling the asteroid’s entry into the Earth’s atmosphere. This process can be time consuming and can require manual adjustment of uncertain modeling specific parameters. NASA Ames has developed a genetic algorithm that can help automate asteroid modeling using the Fragment-Cloud Model (FCM). The algorithm generates realistic energy deposition curves based on actual energy deposition curves from real, observed asteroids. By using these synthetic, labeled energy deposition curves, we developed a one-dimensional convolutional neural network that can predict an asteroid’s pre-entry parameters.

ATAP↗

Lessons Learned in the Application of Machine Learning Techniques to Air Traffic Management

There is an increasing interest in applying methods based on Machine Learning Techniques (MLT) to problems in Air Traffic Management (ATM). The current interest is based on developments in Cloud Computing, the availability of open software and the success of MLT in automation, consumer behavior and finance involving large databases. This paper reviews the current-state-of-the art in applying MLT to aviation operations, its promises and challenges. Historically aviation operations have been analyzed using physics-based models and provide information for making operational decisions. Aviation operations involving many decision makers, multiple objectives, poor or unavailable physics-based models and a rich historical database are prime candidates for analysis using data-driven methods. The promises and challenges in applying MLT to ATM is traced through three examples based on the authors’ experience, each separated by a decade, to show the influence of data and feature selection in the successful application of MLT to ATM. As always, the best approach depends on the task, the physical understanding of the problem and the quality and quantity of the available data.

Machine Learning Techniques↗

Document Classification Techniques for Aviation Letters of Agreement

Often when working with technical documents, it is helpful to classify them into specific categories. In this paper, we conduct a thorough review of natural language processing techniques to perform this classification task on Letters of Agreement (LOAs), technical aviation documents outlining rules for utilizing US airspace. We evaluate multiple techniques, including Transfer Learning, for representing the text in the documents as embeddings: unigram and bigram Term Frequency Inverse Document Frequency (TFIDF), Word2Vec, Doc2Vec, GloVe and RoBERTa. We investigate a wide range of classification models: K-Nearest Neighbors, Random Forest, Support Vector Machines (SVM), Logistic Regression, Naive Bayes, Feed-Forward Neural Network, Convolutional Neural Networks (CNNs) and Long-Short Term Memory (LSTM). By comparing the different methods, we found the best overall approach for our task was to use unigram TFIDF representations with SVM while also gaining insight into how the other methodologies performed on a small technical datasets.

Aayushi Batra↗

Development of an Airspace Simulation and Modeling Tool for Enhanced Spectrum Management

The emergence of new aerial vehicles into the National Airspace System creates an increased demand for aeronautical communications to support aviation operations. However, the issue of spectrum scarcity remains an ever-present concern, and the growing demand cannot be supported using existing spectrum allocation strategies. As a result, a new spectrum management approach is required, and the National Aeronautics and Space Administration (NASA) is investigating advanced concepts to modernize the management and use of aviation spectrum by leveraging the latest advancements in wireless communications, big data and machine learning. This research proposes an autonomous spectrum allocation concept, which allocates communications resources, such as spectrum and power, based on the predicted communications and air traffic demands throughout the airspace, as opposed to the use of fixed allocations as is done today. This approach will result in improved spectrum utilization efficiency and enhanced airspace capacity. The autonomous spectrum allocation concept decomposes into three research areas: demand prediction, resource allocation, and use case evaluation. As part of the use case evaluation effort, a modeling and simulation capability is currently under development. This simulation capability includes the implementation of various features, including visualization of both live or virtually-generated airspace traffic, simulation scenario development, simulation management with data collection, and flight plan creation with corresponding trajectory generation. This modeling and simulation capability will continue to evolve as new and advanced airspace applications are introduced into existing and emerging operational environments.

Eric J. Knoblock↗

Biological Research and Space Health Enabled by Machine Learning to Support Deep Space Missions

A key science goal of the NASA “Moon to Mars” campaign is to understand how biology responds to the Lunar, Martian, and deep space environments in order to advance fundamental knowledge, reduce risk, and support safe, productive human space missions. Through the powerful emerging computer science approaches of artificial intelligence (AI) and machine learning (ML), a paradigm shift has begun in biomedical science and engineered astronaut health systems, to enable Earth-independence and autonomy of mission operations. We present a decadal view of AI/ML architecture to support deep space mission goals, developed in concert with leaders in the field. We describe current AI/ML methods to support 1) fundamental biology, 2) in situ analytics, 3) high performance computing hardware, 4) automated science, 5) self-driving labs, 6) remote data management, 7) integrated real-time mission biomonitoring, and 8) a Precision Space Health system. Cutting-edge AI/ML approaches that can be integrated to support these domains include active learning, explainable AI, adaptive learning, causal inference, knowledge graphs, federated learning, transfer learning, and large language models. Finally, we present results from several current ML projects that are underway in the field to address key challenges of small sample n, high feature count, heterogeneity, and sparse data. These include 1) connecting omics data to phenotypic data using an ensemble model to infer causality of spaceflight rodent liver health disruption, 2) usage of explainable ML to interrogate the muscular underpinnings of spaceflight muscle atrophy, 3) ML models analyzing and determining directed acyclic graphs of human space health risk leveraging rodent bone datasets, 4) usage of large pre-trained models connecting biomedical knowledgebases with small spaceflight datasets to understand gene-to-gene interaction networks, and 5) a suite of benchmarked open science datasets (spaceflight mouse liver; radiation DNA damage) enabling programmers to identify the best ML algorithms to answer space biological science questions.

space biology↗

SIM_EXPLORE: Software for Directed Exploration of Complex Systems

Physics-based numerical simulation codes are widely used in science and engineering to model complex systems that would be infeasible to study otherwise. While such codes may provide the highest- fidelity representation of system behavior, they are often so slow to run that insight into the system is limited. Trying to understand the effects of inputs on outputs by conducting an exhaustive grid-based sweep over the input parameter space is simply too time-consuming. An alternative approach called "directed exploration" has been developed to harvest information from numerical simulators more efficiently. The basic idea is to employ active learning and supervised machine learning to choose cleverly at each step which simulation trials to run next based on the results of previous trials. SIM_EXPLORE is a new computer program that uses directed exploration to explore efficiently complex systems represented by numerical simulations. The software sequentially identifies and runs simulation trials that it believes will be most informative given the results of previous trials. The results of new trials are incorporated into the software's model of the system behavior. The updated model is then used to pick the next round of new trials. This process, implemented as a closed-loop system wrapped around existing simulation code, provides a means to improve the speed and efficiency with which a set of simulations can yield scientifically useful results. The software focuses on the case in which the feedback from the simulation trials is binary-valued, i.e., the learner is only informed of the success or failure of the simulation trial to produce a desired output. The software offers a number of choices for the supervised learning algorithm (the method used to model the system behavior given the results so far) and a number of choices for the active learning strategy (the method used to choose which new simulation trials to run given the current behavior model). The software also makes use of the LEGION distributed computing framework to leverage the power of a set of compute nodes. The approach has been demonstrated on a planetary science application in which numerical simulations are used to study the formation of asteroid families.

Burl, Michael↗