Search NASASearch

SEARCH · Search NASA

Results for “open datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Creating and Using Sensors That Tell Us About Precipitation

Precipitation is deceptively simple to measure - just put a container in the back yard - and this was the only technology available until radar was discovered to be capable of sensing precipitation in the 1940's,leading to quantitative estimates by the 1970's.Meanwhile, satellite-based sensors started advancing.The first precipitation estimates from space re purposed the existing geosynchronous satellite infrared (GEO-IR) data, but purpose-built passive microwave (PMW) sensors soon became a reality. In 1987, the launch of the first Special Sensor Microwave/Imager on the Defense Meteorological Satellite Program F08 by the U.S. Department of Defense, and their decision to open the dataset top public use, created a boom in precipitation algorithms that continues to this day. Experimental work to create a global multi-satellite product by Global Precipitation Climatology Project, and then a "virtual constellation"of PMW sensors from satellite agencies around the globe by the NASA/JAXA Tropical Rainfall Measuring Mission (TRMM) and by NOAA/NWS Climate Prediction Center, resulted in a new generation of quasi-global multi-satellite precipitation estimates at increasingly fine time and space scales. TRMM and the NASA/JAXA Global Precipitation Measurement mission have hosted precipitation radars in space, providing critical new quasi-global information about 3-D precipitation structures and enabling improved calibration of the PMW constellation's estimates.

Huffman, George J.

NASA SpaceCube Edge TPU SmallSat Card for Autonomous Operations and Onboard Science-Data Analysis

Using state-of-the-art artificial intelligence (AI)frameworks onboard spacecraft is challenging because common spacecraft processors cannot provide comparable performance to datacenters with server-grade CPUs and GPUs available for terrestrial applications and advanced deep-learning networks. This limitation makes small, lo w-p o we r AI microchip architectures, such as the Google Coral Edge Tensor Processing Unit (TPU), attractive for space missions where the application-specific design enables both high-performance and power-efficient computing for AI applications. To address these challenging considerations for space deployment, this research introduces the design and capabilities of a CubeSat-sized Edge TPU-based co-processor card, known as the SpaceCube Low-power Ed g e Artificial Intelligence Resilient Node (SC-LEARN). This design conforms to NASA’s CubeSat Card Specification (CS2) for integration into next-generation SmallSat and CubeSat systems. This paper describes the overarching architecture and design of the SC-LEARN, as well as, the supporting test card designed for rapid prototyping and evaluation. The SC-LEARN was developed with three operational modes: (1) a high-performance parallel-processing mode,(2)a fault-tolerant mode for onboard resilience, and (3) a power-saving mode with cold spares. Importantly, this research also elaborates on both training and quantization of Tensor Flow models for the SC-LEARN for use onboard with representative, open-source datasets. Lastly, we describe future research plans, including radiation-beam testing and flight demonstration.

Advanced avionics

Mapping Global Forest Age from Forest Inventories, Biomass and Climate Data

Forest age can determine the capacity of a forest to uptake carbon from the atmosphere. However, a lack of global diagnostics that reflect the forest stage and associated disturbance regimes hampers the quantification of age-related differences in forest carbon dynamics. This study provides a new global distribution of forest age circa 2010, estimated using a machine learning approach trained with more than 40 000 plots using forest inventory, biomass and climate data. First, an evaluation against the plot-level measurements of forest age reveals that the data-driven method has a relatively good predictive capacity of classifying old-growth vs. non-old-growth (precision = 0.81 and 0.99 for old-growth and non-old-growth, respectively) forests and estimating corresponding forest age estimates (NSE = 0.6 – Nash–Sutcliffe efficiency – and RMSE = 50 years – root-mean-square error). However, there are systematic biases of overestimation in young- and underestimation in old-forest stands, respectively. Globally, we find a large variability in forest age with the old-growth forests in the tropical regions of Amazon and Congo, young forests in China, and intermediate stands in Europe. Furthermore, we find that the regions with high rates of deforestation or forest degradation (e.g. the arc of deforestation in the Amazon) are composed mainly of younger stands. Assessment of forest age in the climate space shows that the old forests are either in cold and dry regions or warm and wet regions, while young–intermediate forests span a large climatic gradient. Finally, comparing the presented forest age estimates with a series of regional products reveals differences rooted in different approaches and different in situ observations and global-scale products. Despite showing robustness in cross-validation results, additional methodological insights on further developments should as much as possible harmonize data across the different approaches. The forest age dataset presented here provides additional insights into the global distribution of forest age to better understand the global dynamics in the forest water and carbon cycles. The forest age datasets are openly available at https://doi.org/10.17871/ForestAgeBGI.2021 (Besnard et al., 2021).

Simon Besnard

Assessment of ProgPy - An Open-Source Condition Monitoring and Diagnostics Tool

Traditional maintenance programs, such as corrective and preventive strategies, may lead to high costs and operational inefficiencies. Condition Monitoring and Diagnostics (CM&D) aims to improve these maintenance strategies by enabling timely insights into equipment health and performance. However, implementation of CM&D can be challenging without a robust framework that manages data efficiently, supports interoperability and simplifies integration. To address these challenges ProgPy, an open-source Python-based prognostics tool developed by NASA Ames Research Center, offers a structured solution for broader Prognostics and Health Management (PHM) applications. Ongoing research is assessing the feasibility of implementing ProgPy as a Condition Monitoring and Diagnostics (CM&D) solution by comparing its framework to the guidelines for open CM&D systems recommended in the ISO 13374-2 standard. This evaluation aims to highlight ProgPy’s strengths and identify opportunities for improvement, thereby, contributing to its advancement as an effective tool for Prognostics and Health Management (PHM). This paper presents the results of an initial assessment of the ProgPy toolbox through a gearbox case study using open-source datasets.

Condition-Monitoring, Diagnostics, Failure, Gearbo

FLUXNET-CH4: a global, multi-ecosystem dataset and analysis of methane seasonality from freshwater wetlands

Methane (CH4) emissions from natural landscapes constitute roughly half of global CH4 contributions to the atmosphere, yet large uncertainties remain in the absolute magnitude and the seasonality of emission quantities and drivers. Eddy covariance (EC) measurements of CH4 flux are ideal for constraining ecosystem-scale CH4 emissions due to quasi-continuous and high-temporal-resolution CH4 flux measurements, coincident carbon dioxide, water, and energy flux measurements, lack of ecosystem disturbance, and increased availability of datasets over the last decade. Here, we (1) describe the newly published dataset, FLUXNET-CH4 Version 1.0, the first open-source global dataset of CH4 EC measurements (available at https://fluxnet.org/data/fluxnet-ch4-community-product/, last access: 7 April 2021). FLUXNET-CH4 includes half-hourly and daily gap-filled and non-gap-filled aggregated CH4 fluxes and meteorological data from 79 sites globally: 42 freshwater wetlands, 6 brackish and saline wetlands, 7 formerly drained ecosystems, 7 rice paddy sites, 2 lakes, and 15 uplands. Then, we (2) evaluate FLUXNET-CH4 representativeness for freshwater wetland coverage globally because the majority of sites in FLUXNET-CH4 Version 1.0 are freshwater wetlands which are a substantial source of total atmospheric CH4 emissions; and (3) we provide the first global estimates of the seasonal variability and seasonality predictors of freshwater wetland CH4 fluxes. Our representativeness analysis suggests that the freshwater wetland sites in the dataset cover global wetland bioclimatic attributes (encompassing energy, moisture, and vegetation-related parameters) in arctic, boreal, and temperate regions but only sparsely cover humid tropical regions. Seasonality metrics of wetland CH4 emissions vary considerably across latitudinal bands. In freshwater wetlands (except those between 20°S to 20°N) the spring onset of elevated CH4 emissions starts 3 d earlier, and the CH4 emission season lasts 4 d longer, for each degree Celsius increase in mean annual air temperature. On average, the spring onset of increasing CH4 emissions lags behind soil warming by 1 month, with very few sites experiencing increased CH4 emissions prior to the onset of soil warming. In contrast, roughly half of these sites experience the spring onset of rising CH4 emissions prior to the spring increase in gross primary productivity (GPP). The timing of peak summer CH4 emissions does not correlate with the timing for either peak summer temperature or peak GPP. Our results provide seasonality parameters for CH4 modeling and highlight seasonality metrics that cannot be predicted by temperature or GPP (i.e., seasonality of CH4 peak). FLUXNET-CH4 is a powerful new resource for diagnosing and understanding the role of terrestrial ecosystems and climate drivers in the global CH4 cycle, and future additions of sites in tropical ecosystems and site years of data collection will provide added value to this database. All seasonality parameters are available at https://doi.org/10.5281/zenodo.4672601 (Delwiche et al., 2021). Additionally, raw FLUXNET-CH4 data used to extract seasonality parameters can be downloaded from https://fluxnet.org/data/fluxnet-ch4-community-product/ (last access: 7 April 2021), and a complete list of the 79 individual site data DOIs is provided in Table 2 of this paper.

FLUXNET-CH4

Biological Insights at the Interface of Multiple Arabidopsis Legacy Datasets

The NASA GeneLab database includes an open-access collection of datasets yielded by space biology experiments. Six Gene Lab Data Sets (GLDS’s) performed in Arabidopsis were selected for analysis (7/17/44/121/205/213), all of which included transcriptome data from spaceflight and ground control environments. Hardware, ecotype, environmental conditions, and other experimental conditions varied, allowing the observations of overarching gene expression impacts of microgravity on plant life without focusing on effects of specific experimental conditions. Using GeneLab pre-processed datasets as the basis for the study, RNA microarray data were analyzed to identify genes that showed altered expression in microgravity when compared to control samples for each individual GLDS. All differentially expressed genes were compared to locate differentially expressed genes common between spaceflight experiments. The most noteworthy result is that not one gene shared differential expression among the six GLDS’s. However, gene expression was not influenced randomly by the microgravity environment, as there were several gene ontology terms that were significantly enriched across all experiments. These included 20 significantly enriched biological processes, and although the genes which enriched each term varied, there were many cases of specific genes common to clusters of multiple GLDS’s. Gene expression such as NAC92 and ERF011 or membrane structural element FFP6 provide insight and direction toward understanding the plant response to spaceflight. Characterizing these common processes and the shared differentially expressed genes has demonstrated potential targets for further study to understand and modulate the biological response of plants in microgravity. Life on Earth has never been subjected to the absence of gravity as a selective pressure, so observing how life forms react to a microgravity environment could provide insight to our shared fundamental biological processes. It is also feasible that the genetic modification of specific genes linked to the microgravity response could improve health and yield of space crops.

Joseph Emhof

Establishing Consensus Turbulence Statistics for Hot Subsonic Jets

Many tasks in fluids engineering require knowledge of the turbulence in jets. There is a strong, although fragmented, literature base for low order statistics, such as jet spread and other meanvelocity field characteristics. Some sources, particularly for low speed cold jets, also provide turbulence intensities that are required for validating Reynolds-averaged Navier-Stokes (RANS) Computational Fluid Dynamics (CFD) codes. There are far fewer sources for jet spectra and for space-time correlations of turbulent velocity required for aeroacoustics applications, although there have been many singular publications with various unique statistics, such as Proper Orthogonal Decomposition, designed to uncover an underlying low-order dynamical description of turbulent jet flow. As the complexity of the statistic increases, the number of flows for which the data has been categorized and assembled decreases, making it difficult to systematically validate prediction codes that require high-level statistics over a broad range of jet flow conditions. For several years, researchers at NASA have worked on developing and validating jet noise prediction codes. One such class of codes, loosely called CFD-based or statistical methods, uses RANS CFD to predict jet mean and turbulent intensities in velocity and temperature. These flow quantities serve as the input to the acoustic source models and flow-sound interaction calculations that yield predictions of far-field jet noise. To develop this capability, a catalog of turbulent jet flows has been created with statistics ranging from mean velocity to space-time correlations of Reynolds stresses. The present document aims to document this catalog and to assess the accuracies of the data, e.g. establish uncertainties for the data. This paper covers the following five tasks: Document acquisition and processing procedures used to create the particle image velocimetry (PIV) datasets. Compare PIV data with hotwire and laser Doppler velocimetry (LDV) data published in the open literature. Compare different datasets acquired at roughly the same flow conditions to establish uncertainties. Create a consensus dataset for a range of hot jet flows, including uncertainty bands. Analyze this consensus dataset for self-consistency and compare jet characteristics to those of the open literature. One final objective fulfilled by this work was the demonstration of a universal scaling for the jet flow fields, at least within the region of interest to aeroacoustics. The potential core length and the spread rate of the half-velocity radius were used to collapse of the mean and turbulent velocity fields over the first 20 jet diameters in a highly satisfying manner.

Bridges, James

A Three-Step Semi Analytical Algorithm (3SAA) for Estimating Inherent Optical Properties Over Oceanic, Coastal, and Inland Waters From Remote Sensing Reflectance

We present a three-step inverse model (3SAA) for estimating the inherent optical properties (IOPs) of surface waters from the remote sensing reflectance spectra, Rrs(). The derived IOPs include the total (a()), phytoplankton (aphy()), and colored detrital matter (acdm()), absorption coefficients, and the total (bb()) and particulate (bbp()) backscattering coefficients. The first step uses an improved neural network approach to estimate the diffuse attenuation coefficient of downwelling irradiance from Rrs. a() and bbp() are then estimated using the LS2 model (Loisel et al., 2018), which does not require spectral assumptions on IOPs and hence can assess a() and bb() at any wavelength at which Rrs() is measured. Then, an inverse optimization algorithm is combined with an optical water class (OWC) approach to assess aphy() and acdm() from anw().The proposed model is evaluated using an in situ dataset collected in open oceanic, coastal, and inland waters. Comparisons with other standard semi-analytical algorithms (QAA and GSM), as well as match-up exercises, have also been performed. The applicability of the algorithm on OLCI observations was assessed through the analysis of global IOPs spatial patterns derived from 3SAA and GSM. The good performance of 3SAA is manifested by median absolute percentage differences (MAPD) of 13%, 23%, 34% and 34% for bbp(443), anw(443), aphy(443) and acdm(443), respectively for oceanic waters. Due to the absence of spectral constraints on IOPs in the inversion of total IOPs, and the adoption of an OWC-based approach, the performance of 3SAA is only slightly degraded in bio-optical complex inland waters.

ocean color

Enabling Model Organism and Commercial Astronaut Data Access Through the NASA Open Science Data Repository

NASA’s Open Science Data Repository (OSDR) brings together omics data from NASA’s GeneLab project and non-omics data, including physiological, phenotypic, imaging, and behavioral data from NASA’s Ames Life Sciences Data Archive (ALSDA) collected from decades of space biology research, providing open and FAIR (findable, accessible, interoperable, and reusable) access of these precious data to scientists world-wide. This rich source of meticulously curated metadata and data from spaceflight and analog studies has been mined by the scientific community resulting in dozens of high impact scientific publications that reveals a complex network of molecular and physiological effects of spaceflight across living systems, from microbes to plants, to mammals. Understanding how these effects translate to the human condition is critical as we move deeper into the era of commercial space travel. However, the integration of data, specifically omics data, from astronauts is particularly challenging due to their sensitive nature. OSDR has risen to this challenge by developing a mechanism to control access to identifiable levels of omics data, such as raw sequence data, while enabling public access to processed, unidentifiable, data and associated metadata that will allow the scientific community to interrogate human astronaut data alongside data from model organisms to begin answering these critical questions. The 2021 SpaceX Inspiration4 (I4) mission collected a comprehensive atlas of biological measurements from four civilian astronauts, providing a wealth of data to characterize the effects of spaceflight on the human body. These data include both non-omics and omics assays such as direct RNA sequencing (RNA-seq), single nuclei ATAC-seq and RNA-seq, metagenomics, proteomics, and comprehensive metabolic and cytokine panels, all of which have been integrated into the OSDR system across no less than 9 studies. Each study has been carefully curated using community-backed OSDR standards for sample and assay level metadata ensuring these data are findable and accessible. In addition to hosting both raw and processed data from the principal investigator team for each assay type, the GeneLab team plans to re-process the I4 omics data using GeneLab’s standard processing pipelines. The GeneLab processed data outputs will allow for comparisons across studies on OSDR and enable visualization of these data through the OSDR data visualization platform thereby enabling data reusability and interoperability. Here we describe the robust privacy and security protocols implemented by OSDR to safeguard sensitive health data from astronauts while facilitating metadata and processed data sharing for research purposes. We further provide a road map for navigating the vast amount of data provided for each I4 study on the OSDR, including experimental design, associated experiments, payloads, and missions, data generation and analysis protocols, and associated scientific articles. Additionally, we illustrate how to interrogate the standardized metadata provided in the sample and assay tables as well as various means to download and access the data including programmatically through the GeneLab Open API (GLOpenAPI). The open access of datasets in NASA’s OSDR provides a unique opportunity for the scientific community, as well as citizen scientists and students, to continue using OSDR resources to further unlock profound insights into the consequences of space travel on the human body. Through implementation of security measures to protect sensitive human data, the OSDR seeks to strengthen the science exchange between the Biological and Physical Sciences Program and the Human Research Program, per recommendation 4-1 of the 2023-2032 Decadal Survey, and encourage further sharing and dissemination of astronaut data to provide the scientific community with the resources needed to lay the groundwork for developing targeted mitigation strategies to help withstand the rigors of long-duration spaceflight.

Amanda Marie Saravia-butler

VEDA Visualization Exploration & Data Analysis

Why? - Interdisciplinary science depends on large amount of Earth science data and computational resources - Working with these datasets is non-trivial - Big data science requires advanced distributed computing knowledge What? VEDA is an open platform that brings key Earth science datasets next to open source tools for data processing, analysis, visualization, and exploration in a managed and more accessible computing environment.

Manil Maskey

The NASA Subsonic Jet Particle Image Velocimetry (PIV) Dataset

Many tasks in fluids engineering require prediction of turbulence of jet flows. The present document documents the single-point statistics of velocity, mean and variance, of cold and hot jet flows. The jet velocities ranged from 0.5 to 1.4 times the ambient speed of sound, and temperatures ranged from unheated to static temperature ratio 2.7. Further, the report assesses the accuracies of the data, e.g., establish uncertainties for the data. This paper covers the following five tasks: (1) Document acquisition and processing procedures used to create the particle image velocimetry (PIV) datasets. (2) Compare PIV data with hotwire and laser Doppler velocimetry (LDV) data published in the open literature. (3) Compare different datasets acquired at the same flow conditions in multiple tests to establish uncertainties. (4) Create a consensus dataset for a range of hot jet flows, including uncertainty bands. (5) Analyze this consensus dataset for self-consistency and compare jet characteristics to those of the open literature. The final objective was fulfilled by using the potential core length and the spread rate of the half-velocity radius to collapse of the mean and turbulent velocity fields over the first 20 jet diameters.

Bridges, James

NASA GeneLab Multi-study Visualization Portal

NASA GeneLab has helped advance the field of Space Biology by providing a public repository where researchers can store, share, analyze and visualize the results of space flight related omics experiments. The GeneLab data visualization portal allows any user, regardless of bioinformatics knowledge or access to computational resources, to interact with the experimental data, draw their own conclusions, and gain insights about the effects of space on living systems. These tools help democratize scientific research and foster the NASA Open Science initiative. The new multi-study feature of the GeneLab visualization platform allows users to mine study metadata from RNA sequencing (RNA-seq) experiments to identify samples of interest by filtering datasets based on organism, tissue, assay technology type, and/or factor. Once samples are selected from multiple datasets, users can combine and normalize the sample data, then utilize the visualization displays, including Principal Component Analysis (PCA) plots, to assess sample distributions. Finally, users can perform differential gene expression analysis on the combined data and visualize the results through PCA plots, Volcano plots, Pair plots, Heatmap, Ideogram and Gene Set Enrichment Analysis. All user-generated results and visualizations will be available for download. Here, we present a biological study using samples from multiple GeneLab RNA-seq datasets and analyzed using the multi-study visualization platform to demonstrate inter- and intra-study variability, as well as commonly differentially expressed genes between spaceflight and ground control conditions across datasets. This new feature opens a wide range of possibilities and opportunities for further development including combining other assay technology types and integration with batch effect correction techniques and machine learning applications. Overall, this tool allows users to increase the statistical power of individual experiments, validate hypothesis, identify patterns, and opens the door to new and exciting research.

space biology

Improving Data Discovery, Analysis, and Visualizations With Cloud-Based User Services

The Global Hydrometeorology Resource Center (GHRC) Distributed Active Archive Center (DAAC) is one of 12 DAACs managed by the United States National Aeronautics and Space Administration (NASA) Earth Science Data and Information System (ESDIS) project [1]. GHRC and the other DAACs are designed to process, archive, document, and distribute NASA Earth-observing data, ranging from satellite missions to field campaigns [2]. A major goal of the DAACs is to enable science with these data. Science enabling can be difficult as datasets can be very large, use multiple formats, come from numerous platforms, and require three-dimensional visualization. GHRC is using its expertise with cloud-based technologies to develop open source and open science tools to empower users to explore, coincidentally visualize, and analyze multiple datasets. Being open source, the user community can develop visualizations for their own datasets. This presentation will expand on this objective and highlight the capabilities available to the international community now.

GHRC

A Method for Validating Causal Diagrams of Human Health Risk in Space Flight

The complexity of cause-and-effect relationships between spaceflight hazards and resulting health conditions clouds understanding of the totality of human system risk in space. In response, NASA has introduced Directed Acyclic Graphs (causal diagrams) into the human systems risk management process. These diagrams allow for a common understanding of the mechanisms that lead from unique hazards of spaceflight to the health outcomes important to agencies and astronauts. However, the paucity of available biomedical data from spaceflight creates a need for methods of validating causal models that can accommodate data from spaceflight model analogs. Here we outline one approach utilizing open-access rodent bone datasets from the Ames Life Sciences Data Archive. The properties of directed acyclic graphs themselves can provide an epistemological and statistical framework for validation of a priori causal representations of human system risk in space flight. The assumed causal connections on the graph creates sets of logical implications: variables that – if the causal diagram is correct – should be correlated, as well as sets that should be conditionally independent. By testing these implied correlations and conditional independencies both statistically and heuristically, we can provide evidence for or against specific causal pathways on the causal diagram. In addition to validation of expert-generated causal diagrams, machine learning techniques can learn the most likely structure of a causal diagram from a given dataset. Comparison with and reconciliation between machine-learned causal diagrams and expert-generated diagrams is another technique for challenging assumptions and improving our understanding of causal mechanisms. Accurately representing complex causation is essential to systemic understanding of human health risks in space travel. Having a robust system of validating causal diagrams helps us arrive at more accurate representations of causal systems. This process will be integral to developing the countermeasures necessary for extended exploration of the moon and Mars.

Robert Reynolds

Enabling Biological Discovery Through Biospecimen Sharing: The Nasa Biological Institutional Scientific Collection

Understanding biological impacts from spaceflight hazards and the subsequent development of countermeasures are a high priority to enable humanity to venture back to the Moon, and then to Mars and beyond. Experiments have been conducted with model organisms flown to space and analogous investigations terrestrially, to identify biological mechanistic impacts from spaceflight hazards and to develop mitigation countermeasures, thus contributing towards basic and applied science goals. However, sending organisms into space is a costly endeavor. To maximize scientific return, all biospecimens not required by spaceflight-relevant Principal Investigators are harvested, preserved, and archived in the NASA Biological Institutional Scientific Collection (NBISC). Biospecimens are collected and preserved according to well-established standard operating procedures to maintain scientific quality and are available on-request by the international scientific community. NBISC currently stores over 32,000 biospecimens from Shuttle, International Space Station, and ground-based space analog investigations. Tissue sharing has resulted in at least 33 publications since 2011 and 51 requests since 2016. Many requests for NBISC biospecimens come from first-time investigators who subsequently submit grants as their point-of-entry into the field of spaceflight biology and health. The NBISC biorepository is part of the NASA ‘Open Science for Life in Space’ collaborative group of projects, which includes NASA Genelab, the Space Biology Program’s Biospecimen Sharing Program, Physical Sciences Informatics, and the Ames Life Sciences Data Archive. NBISC biospecimens have been awarded to NASA Genelab, who then generated various open access science ‘omics datasets through the GeneLab Sample Processing laboratory, with resulting data widely used for biological study. Other NBISC biospecimen awards have led to studies on fecal microbiome analysis, DNA damage analysis using single-cell DNA sequencing, enzymatic-pathway identification involved in spaceflight muscle atrophy, and characterization of ocular morphological changes. Of note, NBISC is expanded to include a new Space Microbial Culture Collection (SMCC) for the collection, identification, documentation, long-term preservation, and distribution of space-related microbial isolates.

Biospecimens

Creating Benchmark Data for Artificial Intelligence and Machine Learning Space Biology Research

To identify an appropriate AI/ML approach for a specific problem, the best practice is to measure algorithm performance through the benchmarking process. A scientific benchmark consists of an AI-ready dataset and a reference implementation on a specific scientific question. The NASA Science Mission Directorate (SMD) has started the “Benchmark Initiative for AI/ML to create scientific benchmark datasets in three applications: 1) scientific benchmarking, which finds the best algorithm for a specific problem; 2) application benchmarking, which measures algorithm performance against a set of parameters; and 3) system benchmarking, which evaluates performance of hardware and software architecture. Currently, there are no standardized datasets available to benchmark AI/ML algorithms in the domain of space biology. In this work, we constructed two AI/ML-ready biological datasets from experiments in space-flown mice: cellular imaging and RNA-seq. First, radiation-exposed immune cells harbor DNA damage foci that can be fluorescently marked to visualize the amount of damage following exposure to ionizing radiation. However, such large datasets are difficult to analyze visually, due to imaging inconsistencies and human bias, and classical image processing approaches can fail on imaging artifacts. AI/ML are therefore exciting alternative, providing the speed of machines and the accuracy of humans. We have made this dataset available at https://registry.opendata.aws/bps_microscopy/. Second, high-throughput nucleic acid sequencing (DNA-seq, RNA-seq) has become widespread in biomedical research due to the growing availability and affordability of these assays. However, most sequencing datasets suffer from high dimensionality and low sample count. In this work, we used a generative adversarial network to synthesize a standardized, AI-ready, publicly available benchmark dataset for space biology RNA-seq data with sufficient space-flown and ground control mouse liver samples from NASA GeneLab. This dataset is available at https://registry.opendata.aws/bps_rnaseq/. These datasets are now fully open the Space Biology community to test their favorite AI/ML approaches.

James Casaletto

Mesoscale Phenomena and Their Contribution to the Global Response: A Focus on the Magnetotail Transition Region and Magnetosphere-Ionosphere Coupling

An important question that is being increasingly studied across subdisciplines of Heliophysics is “how do mesoscale phenomena contribute to the global response of the system?” This review paper focuses on this question within two specific but interlinked regions in Near-Earth space: the magnetotail’s transition region to the inner magnetosphere and the ionosphere. There is a concerted effort within the Geospace Environment Modeling (GEM) community to understand the degree to which mesoscale transport in the magnetotail contributes to the global dynamics of magnetic flux transport and dipolarization, particle transport and injections contributing to the storm-time ring current development, and the substorm current wedge. Because the magnetosphere-ionosphere is a tightly coupled system, it is also important to understand how mesoscale transport in the magnetotail impacts auroral precipitation and the global ionospheric system response. Groups within the Coupling, Energetics and Dynamics of Atmospheric Regions Program (CEDAR) community have also been studying how the ionosphere-thermosphere responds to these mesoscale drivers. These specific open questions are part of a larger need to better characterize and quantify mesoscale “messengers” or “conduits” of information—magnetic flux, particle flux, current, and energy—which are key to understanding the global system. After reviewing recent progress and open questions, we suggest datasets that, if developed in the future, will help answer these questions.

transition region