Search NASASearch

SEARCH · Search NASA

Results for “open datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Visual and Inertial Datasets for an eVTOL Aircraft Approach and Landing Scenario

A National Aeronautics and Space Administration (NASA) project developing computer vision algorithms for autonomous flight is producing real-world datasets with cameras mounted on aircraft. In related domains, such as autonomous driving, open datasets are key to innovation and advancement in computer vision and autonomous perception for future Advanced Air Mobility (AAM) operations. Few vision datasets, however, are publicly available in the aviation context. This paper introduces preliminary datasets containing several examples of approach and landing scenarios. The platform aircraft include a multirotor small unmanned aerial system (sUAS) and a crewed helicopter as surrogates for future electric vertical take-off and landing (eVTOL) aircraft. The dataset provides video imagery with associated inertial navigation system-global positioning system (INS-GPS) position and attitude estimates and other sensors. Surveyed locations of the visual features of the landing area are included. This dataset is the first to be released in an ongoing effort to collect and share large, diverse datasets relevant to autonomous aviation; community critique that can inform and improve future flight campaigns is welcome.

Nelson Brown

Developing Open-Source Training Materials for AI/ML and Space Biological Sciences Using NASA Cloud-Based Data

Artificial Intelligence (AI) and Machine Learning (ML) has gained significant traction in the biological and biomedical research fields, in part due to a culture of open data sharing and reuse. AI/ML methodology is well-suited to recognize and predict biological patterns from high-dimensional next-generation sequencing data (e.g. whole genome sequencing, transcriptomic sequencing), as well as from biological or medical imaging data (e.g. microscopy, computed tomography, ultrasound, magnetic resonance imaging, radiography). These methodologies hold particular promise for space biosciences research and automated space health monitoring systems. However, there are key considerations for properly training, validating, and testing a machine learning model in biological research or clinical application. Inexperienced researchers can produce models that perform poorly outside of the training dataset. Open Science principles such as data sharing and open-source code must go hand-in-hand with publicly available, high-quality training curricula in best practices, with modules centered on real-life scientific use cases and data so future AI/ML practitioners gain experience on real problems. Here we present the development of open-source training materials for AI/ML and space biosciences, as part of the NASA Transform to Open Science Training (TOPST) initiative. We develop 4 independent training programs, focused on the following topics: 1) Fundamentals of Machine Learning and Space Biosciences Domain, 2) Open Science, Artificial Intelligence, and Ethical Best Practices for Data Sharing and Analysis, 3) Using AI/ML Classification to Identify Gene Networks Affected By Space Exposure in Mouse Liver, and 4) Using Neural Networks to Find DNA Damage Patterns in Immune Cells after Radiation. All programs leverage cloud-based NASA biological datasets. The curriculum we present will enable worldwide access to training in AI/ML and scientific analysis.

James Casaletto

IMAGE Software Suite

The IMAGE Mission is generating a truely unique set of magnetospheric measurement through a first-of-its-kind complement of remote, global observations. These data are being distributed in the Universal Data Format (UDF), which consists of data, calibration, and documentation. This is an open dataset, available to all by request to the National Space Science Data Center (NSSDC) at NASA Goddard Space Flight Center. Browse data, which consists of summary observations, is also available through the NSSDC in the Common Data Format (CDF) and graphic representations of the browse data. Access to the browse data can be achieved through the NSSDC CDAWeb services or by use of NSSDC provided software tools. This presentation documents the software tools, being provided by the IMAGE team, for use in viewing and analyzing the UDF telemetry data. Like the IMAGE data, these tools are openly available. What these tools can do, how they can be obtained, and how they are expected to evolve will be discussed.

Gallagher, Dennis L.

Public-Private Partnerships to Enable Discovery, Access, and Use of NASA’s Open Earth Science Datasets

Knowledge transfer between public research institutes and private entities is an essential component of the open science movement. While both public and private institutions are making research advances in technologies, organizational boundaries can hinder knowledge transfer. Productive public-private collaboration frameworks are needed to advance research further.

Elizabeth Fancher

Algorithm Performance Dataset from NASA Open-Source Software

NASA Langley Research Center has recently developed and released the open-source software Multi Model Monte Carlo with Python (MXMCPy- LAR-19756-1) as a general capability for computing the statistics of outputs from an expensive, high-fidelity model by leveraging faster, low-fidelity models for speedup. Given a fixed computational budget and a collection of models with varying cost/accuracy, multi model Monte Carlo (MC) seeks a sample allocation strategy across the models that results in an estimator with optimal variance reduction. MXMCPy is a versatile tool that enables convenient access to many existing multi-model MC approaches (over a dozen algorithms available) within one modular and extensible package [1]. With MXMCPy, users can easily compare existing methods to determine the best choice for their particular problem,while developers have a basis for implementing and sharing new variance reduction approaches. However,there is currently very little understanding about which algorithm will perform best for a given problem (defined by the correlation between and relative cost of the available models) without a brute force search.

Geoffrey F Bomarito

Biological Research and Space Health Enabled by Machine Learning to Support Deep Space Missions

A key science goal of the NASA “Moon to Mars” campaign is to understand how biology responds to the Lunar, Martian, and deep space environments in order to advance fundamental knowledge and support human space missions. Through artificial intelligence (AI) and machine learning (ML), a paradigm shift has begun in space biosciences and engineered astronaut health systems, to enable Earth-independence and mission operations autonomy. We describe current AI/ML methods to support 1) fundamental biology, 2) in situ analytics, 3) high performance computing, 4) automated science, 5) self-driving labs, 6) remote data management, 7) integrated mission biomonitoring, and 8) a Precision Space Health system. AI/ML approaches that can be integrated to support these domains include active learning, explainable AI, adaptive learning, causal inference, knowledge graphs, federated learning, transfer learning, and large language models. Finally, we present results from several current ML projects that are underway in the space biology field to address key challenges of small sample n, high feature count, heterogeneity, and sparse data. These include 1) connecting omics to phenotypic data using an ensemble model to infer causality of rodent liver health disruption, 2) usage of explainable ML to interrogate muscular underpinnings of muscle atrophy, 3) ML models analyzing and determining directed acyclic graphs of human health risk leveraging rodent bone datasets, 4) usage of large pre-trained models connecting biomedical knowledgebases with small spaceflight datasets to understand gene-to-gene interactions, and 5) a suite of benchmarked open science datasets enabling programmers to identify best algorithms to answer space biology questions.

space biology

Making Dataset Quality Information FAIR: Supporting Open-Source Science and Enhancing (Re)Use and Trustworthiness of Scientific Data

- Quality information should be documented and readily shared within and across domains. - Sharing of dataset quality information supports open science and trustworthiness of scientific data. - Dataset quality is more than data quality. - Quality tends to be domain-specific and context-dependent. - Community guidelines provide practical steps towards FAIR dataset quality information.

Ge Peng

Biological Research and Space Health Enabled by Machine Learning to Support Deep Space Missions

A key science goal of the NASA “Moon to Mars” campaign is to understand how biology responds to the Lunar, Martian, and deep space environments in order to advance fundamental knowledge, reduce risk, and support safe, productive human space missions. Through the powerful emerging computer science approaches of artificial intelligence (AI) and machine learning (ML), a paradigm shift has begun in biomedical science and engineered astronaut health systems, to enable Earth-independence and autonomy of mission operations. We present a decadal view of AI/ML architecture to support deep space mission goals, developed in concert with leaders in the field. We describe current AI/ML methods to support 1) fundamental biology, 2) in situ analytics, 3) high performance computing hardware, 4) automated science, 5) self-driving labs, 6) remote data management, 7) integrated real-time mission biomonitoring, and 8) a Precision Space Health system. Cutting-edge AI/ML approaches that can be integrated to support these domains include active learning, explainable AI, adaptive learning, causal inference, knowledge graphs, federated learning, transfer learning, and large language models. Finally, we present results from several current ML projects that are underway in the field to address key challenges of small sample n, high feature count, heterogeneity, and sparse data. These include 1) connecting omics data to phenotypic data using an ensemble model to infer causality of spaceflight rodent liver health disruption, 2) usage of explainable ML to interrogate the muscular underpinnings of spaceflight muscle atrophy, 3) ML models analyzing and determining directed acyclic graphs of human space health risk leveraging rodent bone datasets, 4) usage of large pre-trained models connecting biomedical knowledgebases with small spaceflight datasets to understand gene-to-gene interaction networks, and 5) a suite of benchmarked open science datasets (spaceflight mouse liver; radiation DNA damage) enabling programmers to identify the best ML algorithms to answer space biological science questions.

space biology

Biological Research and Space Health Enabled by Machine Learning to Support Deep Space Missions

A key science goal of the NASA “Moon to Mars” campaign is to understand how biology responds to the Lunar, Martian, and deep space environments in order to advance fundamental knowledge, reduce risk, and support safe, productive human space missions. Through the powerful emerging computer science approaches of artificial intelligence (AI) and machine learning (ML), a paradigm shift has begun in biomedical science and engineered astronaut health systems, to enable Earth-independence and autonomy of mission operations. We present a decadal view of AI/ML architecture to support deep space mission goals, developed in concert with leaders in the field. We describe current AI/ML methods to support 1) fundamental biology, 2) in situ analytics, 3) high performance computing hardware, 4) automated science, 5) self-driving labs, 6) remote data management, 7) integrated real-time mission biomonitoring, and 8) a Precision Space Health system. Cutting-edge AI/ML approaches that can be integrated to support these domains include active learning, explainable AI, adaptive learning, causal inference, knowledge graphs, federated learning, transfer learning, and large language models. Finally, we present results from several current ML projects that are underway in the field to address key challenges of small sample n, high feature count, heterogeneity, and sparse data. These include 1) connecting omics data to phenotypic data using an ensemble model to infer causality of spaceflight rodent liver health disruption, 2) usage of explainable ML to interrogate the muscular underpinnings of spaceflight muscle atrophy, 3) ML models analyzing and determining directed acyclic graphs of human space health risk leveraging rodent bone datasets, 4) usage of large pre-trained models connecting biomedical knowledgebases with small spaceflight datasets to understand gene-to-gene interaction networks, and 5) a suite of benchmarked open science datasets (spaceflight mouse liver; radiation DNA damage) enabling programmers to identify the best ML algorithms to answer space biological science questions.

space biology

Instantaneous Photosynthetically Available Radiation (IPAR) prediction models based on Neural Network for ocean waters.

Instantaneous photosynthetically available radiation (IPAR) at the ocean surface and its vertical profile below the surface play a critical role in models to calculate net primary productivity of marine phytoplankton. In this work, we report two IPAR prediction models based on neural network (NN) approach, one for open ocean and the other for coastal waters. These models are trained, validated, and tested using a large volume of synthetic datasets for open ocean and coastal waters simulated by a radiative transfer model. Our NN models are designed to predict IPAR under a large range of atmospheric and oceanic conditions. The NN models can compute subsurface IPAR profile very accurately up to euphotic zone depth. The root mean square errors associated with the diffuse attenuation coefficient of IPAR are less than 0.011 𝑚−1 and 0.036 𝑚−1 for open ocean and coastal waters respectively. The performance of the NN models is better than presently available semi analytical models, with significant superiority in coastal waters.

PACE

Classification of Notices to Airmen using Natural Language Processing

This paper establishes the feasibility of using Natural Language Processing (NLP) to classify NOTAMs or Notices to Airmen – a pilot messaging framework to gather real-time situational awareness. Present day air mobility operations heavily rely on NOTAMs. However, pilots often have difficulty interpreting NOTAMs due to the sheer volume of inapplicable messages and unclear abbreviations. Using NLP, the presented study analyzes the accuracy of classifying NOTAMs and, thereby, the efficiency of generating actionable interpretations in real time. To this effect, efficacies of four NLP neural network architectures were analyzed, including three Recurrent Neural Networks (RNNs) with GloVe, Word2Vec, and FastText word embeddings, and one trained Bi-Directional Encoder Representations from Transformers (BERT) model. The four neural networks were trained and evaluated on three open-source datasets of varying text lengths, vocabularies, and grammars, taken from e-commerce product descriptions, social media tweets, and unstructured descriptions for data and analytics services on open data marketplaces such as NASA’s Data and Reasoning Fabric (DRF) platform. This provided cross-analysis of each neural network architecture’s performance per text type. The best performing architecture, BERT, was then fine-tuned on a collection of open-source NOTAM data. Post-training, a real-time NOTAM classification service was implemented to draw inference on new NOTAMs using the trained model, which demonstrated close to 99% accuracy in classification. This modular classification service is envisioned to be integrated with a data and analytics delivery platform, such as the DRF, thus availing real-time contextualization of NOTAMs to air mobility clients, humans, and machines for enhanced decision making.

Aiden C. Szeto

Classification of Notices to Airmen using Natural Language Processing

This paper establishes the feasibility of using Natural Language Processing (NLP) to classify NOTAMs or Notices to Airmen – a pilot messaging framework to gather real-time situational awareness. Present day air mobility operations heavily rely on NOTAMs. However, pilots often have difficulty interpreting NOTAMs due to the sheer volume of inapplicable messages and unclear abbreviations. Using NLP, the presented study analyzes the accuracy of classifying NOTAMs and, thereby, the efficiency of generating actionable interpretations in real time. To this effect, efficacies of four NLP neural network architectures were analyzed, including three Recurrent Neural Networks (RNNs) with GloVe, Word2Vec, and FastText word embeddings, and one trained Bi-Directional Encoder Representations from Transformers (BERT) model. The four neural networks were trained and evaluated on three open-source datasets of varying text lengths, vocabularies, and grammars, taken from e-commerce product descriptions, social media tweets, and unstructured descriptions for data and analytics services on open data marketplaces such as NASA’s Data and Reasoning Fabric (DRF) platform. This provided cross-analysis of each neural network architecture’s performance per text type. The best performing architecture, BERT, was then fine-tuned on a collection of open-source NOTAM data. Post-training, a real-time NOTAM classification service was implemented to draw inference on new NOTAMs using the trained model, which demonstrated close to 99% accuracy in classification. This modular classification service is envisioned to be integrated with a data and analytics delivery platform, such as the DRF, thus availing real-time contextualization of NOTAMs to air mobility clients, humans, and machines for enhanced decision making.

Aiden Szeto

New Global Characterization of Landslide Exposure

Landslides triggered by intense rainfall are hazards that impact people and infrastructure across the world, but comprehensively quantifying exposure to these hazards remains challenging. Unlike earthquakes or flooding which cover large areas, landslides occur only in highly susceptible parts of a landscape affected by intense rainfall, which may not intersect human settlement or infrastructure. Existing datasets of landslides around the world generally include only those reported to have caused impacts, leading to significant biases toward areas with higher reporting capacity, limiting how our understanding of exposure to landslides in developing countries. In this study, we use an alternative approach to estimate exposure to landslides in a homogenous fashion. We have combined a global landslide hazard proxy derived from satellite data with open-source datasets on population, roads and infrastructure to consistently estimate exposure to rapid landslide hazards around the globe. These exposure models compare favorably with existing datasets of rainfall-triggered landslide fatalities, while filling in major gaps in inventory-based estimates in parts of the world with lower reporting capacity. Our findings provide a global estimate of exposure to landslides from 2001-2019 that we suggest may be useful to disaster mitigation professionals.

earthquakes

What's New in the PSI? (Physical Sciences Informatics)

The NASA Physical Sciences Informatics (PSI) database is NASA’s archival for physical sciences research in microgravity – on ISS and also other reduced gravity platforms. The database has been making data from microgravity physical science investigations publicly available since its launch late 2014. The database was an early investment in Open Science for NASA’s Biophysics and Physical Sciences (BPS) Division of the Science Mission Directorate (SMD), who conducts fundamental and applied physical sciences yielding research data in 6 disciplines: biophysics, combustion science, complex fluids, fluid physics, fundamental physics, and materials science. PSI started initially with data from 14 microgravity investigations. Receipt of data through the years and an annual NRA to fund ground investigations extending flight datasets has increased the total to 72 investigations available with 8 new datasets in the 2020 NRA (NNH20ZDA014N). These new datasets are open for public access at the PSI website at https://www.nasa.gov/PSI.

Physical Sciences Informatics PSI database

Review of Cranked-Arrow Wing Aerodynamics Project: Its International Aeronautical Community Role

This paper provides a brief history of the F-16XL-1 aircraft, its role in the High Speed Research (HSR) program and how it was morphed into the Cranked Arrow Wing Aerodynamics Project (CAWAP). Various flight, wind-tunnel and Computational Fluid Dynamics (CFD) data sets were generated during the CAWAP. These unique and open flight datasets for surface pressures, boundary-layer profiles and skinfriction distributions, along with surface flow data, are described and sample data comparisons given. This is followed by a description of how the project became internationalized to be known as Cranked Arrow Wing Aerodynamics Project International (CAWAPI) and is concluded by an introduction to the results of a 4 year CFD predictive study of data collected at flight conditions by participating researchers.

Lamar, John E.

Overview of the Cranked-Arrow Wing Aerodynamics Project International

This paper provides a brief history of the F-16XL-1 aircraft, its role in the High Speed Research program and how it was morphed into the Cranked Arrow Wing Aerodynamics Project. Various flight, wind-tunnel and Computational Fluid Dynamics data sets were generated as part of the project. These unique and open flight datasets for surface pressures, boundary-layer profiles and skin-friction distributions, along with surface flow data, are described and sample data comparisons given. This is followed by a description of how the project became internationalized to be known as Cranked Arrow Wing Aerodynamics Project International and is concluded by an introduction to the results of a four year computational predictive study of data collected at flight conditions by participating researchers.

Obara, Clifford J.

The Cranked Arrow Wing Aerodynamics Project (CAWAP) and its Extension to the International Community as CAWAPI: Objectives and Overview

This paper provides a brief history of the F-16XL-1 aircraft, its role in the High Speed Research (HSR) program and how it was morphed into the Cranked Arrow Wing Aerodynamics Project (CAWAP). Various flight, wind-tunnel and Computational Fluid Dynamics (CFD) data sets were generated during the CAWAP. These unique and open flight datasets for surface pressures, boundary-layer profiles and skin-friction distributions, along with surface flow data, are described and sample data comparisons given. This is followed by a description of how the project became internationalized to be known as Cranked Arrow Wing Aerodynamics Project International (CAWAPI) and is concluded by an introduction to the results of a 5-year CFD predictive study of data.

Lamar, John E.