Search NASASearch

SEARCH · Search NASA

Results for “open datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Generation of random geological models using multi-randomization for machine learning

Generating high-fidelity geological models is essential for advancing machine learning (ML) methods in automated seismic interpretation. For instance, seismic images paired with corresponding fault labels are foundational for ML-based fault detection from seismic migration sections. While several open-access datasets of random geological models exist, open-source tools specifically designed to produce large volumes of such models for ML applications remain scarce. To address this gap, we present RGM (Random Geological Model), an open-source software package for efficiently generating 2D and 3D synthetic geological models tailored for ML workflows. RGM supports the creation of diverse model components, including medium property distributions (P-/S-wave velocities and density), seismic reflectivity images (i.e., synthetic migration sections), relative geological time, and discrete fault attributes such as probability, dip, strike, rake, and displacement. It also accommodates the creation of complex geological features such as salt bodies and unconformities. The model generation algorithm employs a multi-randomization strategy, yielding an effectively infinite-dimensional model space that encompasses a wide range of geological scenarios and associated seismic features. Furthermore, RGM incorporates a method to generate synthetic elastic migration images using analytical elastic reflection coefficients combined with frequency-dependent scaling. This functionality enables the creation of training datasets for ML models that leverage elastic seismic images. RGM is implemented in modern object-oriented Fortran, allowing users to flexibly control statistical parameters governing model variability. We demonstrate the capability, performance, and geological realism of the package through comprehensive 2D and 3D examples.

58 GEOSCIENCES

PPI DataHub Project Data Package: S. elongatus PCC 7942 Circadian Control Bioproduction Metabolomics (PB-DP5)

The purpose of this experiment was to evaluate how circadian clock regulation impacts carbon partitioning between storage, growth, and product synthesis in Synechococcus elongatus PCC 7942 in providing insights to strategies for enhanced bioproduction. Culture samples were collected at 0, 0.5, 1, 2, 4, 6, and 8 hours for extracellular sucrose analysis. Circadian metabolomics data was acquired using a Agilent single quadrupole gas chromatography-mass spectrometer and processed using Agilent Mass Hunter for targeted sucrose quantification. Metabolomic analysis of PCC 7942 light-dark cycle cultures transitioned to constant light revealed distinct temporal patterns in sucrose production. Processed metabolomic datasets are openly accessible from the PNNL DataHub project dataset download page and contain secondary processed GC-MS results files and supporting metadata materials linked to relevant source code information supporting data transparency and reuse.

59 BASIC BIOLOGICAL SCIENCES

PPI DataHub Project Data Package: S. elongatus PCC 7942 Circadian Control Bioproduction Transcriptomics (PB-DP3)

The purpose of this experiment was to evaluate how circadian clock regulation impacts carbon partitioning between storage, growth, and product synthesis in Synechococcus elongatus PCC 7942 in providing insights to strategies for enhanced bioproduction. Sample data was acquired using a Illumina HiSeq sequencer system and processed for RNA sequencing (RNA-Seq) expression analysis. Transcriptomic differential expression analysis revealed coordinated circadian clock-driven adjustment of the cell cycle and rewiring of energy and carbon metabolism. Processed RNA-Seq datasets are openly accessible from the PNNL DataHub project dataset download page and contain secondary processed RNA-seq results files and supporting metadata materials linked to relevant source code information supporting data transparency and reuse.

59 BASIC BIOLOGICAL SCIENCES

Algorithm Performance Dataset from NASA Open-Source Software

NASA Langley Research Center has recently developed and released the open-source software Multi Model Monte Carlo with Python (MXMCPy- LAR-19756-1) as a general capability for computing the statistics of outputs from an expensive, high-fidelity model by leveraging faster, low-fidelity models for speedup. Given a fixed computational budget and a collection of models with varying cost/accuracy, multi model Monte Carlo (MC) seeks a sample allocation strategy across the models that results in an estimator with optimal variance reduction. MXMCPy is a versatile tool that enables convenient access to many existing multi-model MC approaches (over a dozen algorithms available) within one modular and extensible package [1]. With MXMCPy, users can easily compare existing methods to determine the best choice for their particular problem,while developers have a basis for implementing and sharing new variance reduction approaches. However,there is currently very little understanding about which algorithm will perform best for a given problem (defined by the correlation between and relative cost of the available models) without a brute force search.

Geoffrey F Bomarito

A Curated Dataset of Regional Meteor Events with Simultaneous Optical and Infrasound Observations (2006–2011)

We present a curated, openly accessible dataset of 71 regional meteor events simultaneously recorded by optical and infrasound instrumentation between 2006 and 2011. These events were captured during an observational campaign using the all-sky cameras of the Southern Ontario Meteor Network and the co-located Elginfield Infrasound Array. Each entry provides optical trajectory measurements, infrasound waveforms, and atmospheric specification profiles. The integration of optical and acoustic data enables robust linkage between observed acoustic signals and specific points along meteor trajectories, offering new opportunities to examine shock wave generation, propagation, and energy deposition processes. This release fills a critical observational gap by providing the first validated, openly accessible archive of simultaneous optical–infrasound meteor observations that supports trajectory reconstruction, acoustic propagation modeling, and energy deposition analyses. By making these data openly available in a structured format, this work establishes a durable reference resource that advances reproducibility, fosters cross-disciplinary research, and underpins future developments in meteor physics, atmospheric acoustics, and planetary defense.

astrometry

Biological Research and Space Health Enabled by Machine Learning to Support Deep Space Missions

A key science goal of the NASA “Moon to Mars” campaign is to understand how biology responds to the Lunar, Martian, and deep space environments in order to advance fundamental knowledge and support human space missions. Through artificial intelligence (AI) and machine learning (ML), a paradigm shift has begun in space biosciences and engineered astronaut health systems, to enable Earth-independence and mission operations autonomy. We describe current AI/ML methods to support 1) fundamental biology, 2) in situ analytics, 3) high performance computing, 4) automated science, 5) self-driving labs, 6) remote data management, 7) integrated mission biomonitoring, and 8) a Precision Space Health system. AI/ML approaches that can be integrated to support these domains include active learning, explainable AI, adaptive learning, causal inference, knowledge graphs, federated learning, transfer learning, and large language models. Finally, we present results from several current ML projects that are underway in the space biology field to address key challenges of small sample n, high feature count, heterogeneity, and sparse data. These include 1) connecting omics to phenotypic data using an ensemble model to infer causality of rodent liver health disruption, 2) usage of explainable ML to interrogate muscular underpinnings of muscle atrophy, 3) ML models analyzing and determining directed acyclic graphs of human health risk leveraging rodent bone datasets, 4) usage of large pre-trained models connecting biomedical knowledgebases with small spaceflight datasets to understand gene-to-gene interactions, and 5) a suite of benchmarked open science datasets enabling programmers to identify best algorithms to answer space biology questions.

space biology

Making Dataset Quality Information FAIR: Supporting Open-Source Science and Enhancing (Re)Use and Trustworthiness of Scientific Data

- Quality information should be documented and readily shared within and across domains. - Sharing of dataset quality information supports open science and trustworthiness of scientific data. - Dataset quality is more than data quality. - Quality tends to be domain-specific and context-dependent. - Community guidelines provide practical steps towards FAIR dataset quality information.

Ge Peng

Biological Research and Space Health Enabled by Machine Learning to Support Deep Space Missions

A key science goal of the NASA “Moon to Mars” campaign is to understand how biology responds to the Lunar, Martian, and deep space environments in order to advance fundamental knowledge, reduce risk, and support safe, productive human space missions. Through the powerful emerging computer science approaches of artificial intelligence (AI) and machine learning (ML), a paradigm shift has begun in biomedical science and engineered astronaut health systems, to enable Earth-independence and autonomy of mission operations. We present a decadal view of AI/ML architecture to support deep space mission goals, developed in concert with leaders in the field. We describe current AI/ML methods to support 1) fundamental biology, 2) in situ analytics, 3) high performance computing hardware, 4) automated science, 5) self-driving labs, 6) remote data management, 7) integrated real-time mission biomonitoring, and 8) a Precision Space Health system. Cutting-edge AI/ML approaches that can be integrated to support these domains include active learning, explainable AI, adaptive learning, causal inference, knowledge graphs, federated learning, transfer learning, and large language models. Finally, we present results from several current ML projects that are underway in the field to address key challenges of small sample n, high feature count, heterogeneity, and sparse data. These include 1) connecting omics data to phenotypic data using an ensemble model to infer causality of spaceflight rodent liver health disruption, 2) usage of explainable ML to interrogate the muscular underpinnings of spaceflight muscle atrophy, 3) ML models analyzing and determining directed acyclic graphs of human space health risk leveraging rodent bone datasets, 4) usage of large pre-trained models connecting biomedical knowledgebases with small spaceflight datasets to understand gene-to-gene interaction networks, and 5) a suite of benchmarked open science datasets (spaceflight mouse liver; radiation DNA damage) enabling programmers to identify the best ML algorithms to answer space biological science questions.

space biology

Biological Research and Space Health Enabled by Machine Learning to Support Deep Space Missions

A key science goal of the NASA “Moon to Mars” campaign is to understand how biology responds to the Lunar, Martian, and deep space environments in order to advance fundamental knowledge, reduce risk, and support safe, productive human space missions. Through the powerful emerging computer science approaches of artificial intelligence (AI) and machine learning (ML), a paradigm shift has begun in biomedical science and engineered astronaut health systems, to enable Earth-independence and autonomy of mission operations. We present a decadal view of AI/ML architecture to support deep space mission goals, developed in concert with leaders in the field. We describe current AI/ML methods to support 1) fundamental biology, 2) in situ analytics, 3) high performance computing hardware, 4) automated science, 5) self-driving labs, 6) remote data management, 7) integrated real-time mission biomonitoring, and 8) a Precision Space Health system. Cutting-edge AI/ML approaches that can be integrated to support these domains include active learning, explainable AI, adaptive learning, causal inference, knowledge graphs, federated learning, transfer learning, and large language models. Finally, we present results from several current ML projects that are underway in the field to address key challenges of small sample n, high feature count, heterogeneity, and sparse data. These include 1) connecting omics data to phenotypic data using an ensemble model to infer causality of spaceflight rodent liver health disruption, 2) usage of explainable ML to interrogate the muscular underpinnings of spaceflight muscle atrophy, 3) ML models analyzing and determining directed acyclic graphs of human space health risk leveraging rodent bone datasets, 4) usage of large pre-trained models connecting biomedical knowledgebases with small spaceflight datasets to understand gene-to-gene interaction networks, and 5) a suite of benchmarked open science datasets (spaceflight mouse liver; radiation DNA damage) enabling programmers to identify the best ML algorithms to answer space biological science questions.

space biology

Instantaneous Photosynthetically Available Radiation (IPAR) prediction models based on Neural Network for ocean waters.

Instantaneous photosynthetically available radiation (IPAR) at the ocean surface and its vertical profile below the surface play a critical role in models to calculate net primary productivity of marine phytoplankton. In this work, we report two IPAR prediction models based on neural network (NN) approach, one for open ocean and the other for coastal waters. These models are trained, validated, and tested using a large volume of synthetic datasets for open ocean and coastal waters simulated by a radiative transfer model. Our NN models are designed to predict IPAR under a large range of atmospheric and oceanic conditions. The NN models can compute subsurface IPAR profile very accurately up to euphotic zone depth. The root mean square errors associated with the diffuse attenuation coefficient of IPAR are less than 0.011 𝑚−1 and 0.036 𝑚−1 for open ocean and coastal waters respectively. The performance of the NN models is better than presently available semi analytical models, with significant superiority in coastal waters.

PACE

Datasets of Faults in Variable Air Volume Terminal Units in a Multi-Zone Commercial Building

Faults in HVAC systems can decrease system efficiency and equipment lifespan, leading to 5%–30% of energy consumption being wasted in commercial buildings. We identified two common faults in HVAC variable air volume systems: a stuck damper fault in the variable air volume terminal unit and a discharge airflow sensor fault. We conducted three sets of damper stuck tests and two sets of airflow sensor tests, each including a fault-free scenario and scenarios with varying levels of faults, over one day. The faults were implemented in Oak Ridge National Laboratory’s two-story Flexible Research Platform building to generate a high-quality, well-controlled dataset covering fault-induced and fault-free scenarios. The test building, fault test scenarios, and data validation are described here. The open-source dataset includes 1 min intervals of weather and building data on the presence and absence of building faults. This dataset can be used to analyze the effects of HVAC system faults on system operation and indoor building conditions, and to develop or evaluate a fault detection and diagnosis algorithm.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Classification of Notices to Airmen using Natural Language Processing

This paper establishes the feasibility of using Natural Language Processing (NLP) to classify NOTAMs or Notices to Airmen – a pilot messaging framework to gather real-time situational awareness. Present day air mobility operations heavily rely on NOTAMs. However, pilots often have difficulty interpreting NOTAMs due to the sheer volume of inapplicable messages and unclear abbreviations. Using NLP, the presented study analyzes the accuracy of classifying NOTAMs and, thereby, the efficiency of generating actionable interpretations in real time. To this effect, efficacies of four NLP neural network architectures were analyzed, including three Recurrent Neural Networks (RNNs) with GloVe, Word2Vec, and FastText word embeddings, and one trained Bi-Directional Encoder Representations from Transformers (BERT) model. The four neural networks were trained and evaluated on three open-source datasets of varying text lengths, vocabularies, and grammars, taken from e-commerce product descriptions, social media tweets, and unstructured descriptions for data and analytics services on open data marketplaces such as NASA’s Data and Reasoning Fabric (DRF) platform. This provided cross-analysis of each neural network architecture’s performance per text type. The best performing architecture, BERT, was then fine-tuned on a collection of open-source NOTAM data. Post-training, a real-time NOTAM classification service was implemented to draw inference on new NOTAMs using the trained model, which demonstrated close to 99% accuracy in classification. This modular classification service is envisioned to be integrated with a data and analytics delivery platform, such as the DRF, thus availing real-time contextualization of NOTAMs to air mobility clients, humans, and machines for enhanced decision making.

Aiden C. Szeto

Classification of Notices to Airmen using Natural Language Processing

This paper establishes the feasibility of using Natural Language Processing (NLP) to classify NOTAMs or Notices to Airmen – a pilot messaging framework to gather real-time situational awareness. Present day air mobility operations heavily rely on NOTAMs. However, pilots often have difficulty interpreting NOTAMs due to the sheer volume of inapplicable messages and unclear abbreviations. Using NLP, the presented study analyzes the accuracy of classifying NOTAMs and, thereby, the efficiency of generating actionable interpretations in real time. To this effect, efficacies of four NLP neural network architectures were analyzed, including three Recurrent Neural Networks (RNNs) with GloVe, Word2Vec, and FastText word embeddings, and one trained Bi-Directional Encoder Representations from Transformers (BERT) model. The four neural networks were trained and evaluated on three open-source datasets of varying text lengths, vocabularies, and grammars, taken from e-commerce product descriptions, social media tweets, and unstructured descriptions for data and analytics services on open data marketplaces such as NASA’s Data and Reasoning Fabric (DRF) platform. This provided cross-analysis of each neural network architecture’s performance per text type. The best performing architecture, BERT, was then fine-tuned on a collection of open-source NOTAM data. Post-training, a real-time NOTAM classification service was implemented to draw inference on new NOTAMs using the trained model, which demonstrated close to 99% accuracy in classification. This modular classification service is envisioned to be integrated with a data and analytics delivery platform, such as the DRF, thus availing real-time contextualization of NOTAMs to air mobility clients, humans, and machines for enhanced decision making.

Aiden Szeto

Beyond Price-Taker: Multiscale Optimization of a Wind-Battery Integrated Energy System within the Wholesale Electricity Market

This work presents the optimization of a wind-battery IES using the multiscale optimization framework proposed in our previous work to quantify errors from the price-taker assumption. The framework, built over Prescient (an open-source package for solving production cost models), is applied to the RTS-GMLC dataset, an open-source dataset that is representative of the southwest U.S. wholesale electricity market. The framework provides detailed bidding, market clearing, and control processes of an IES, and it can quantify how the IES interacts with the market. In this work, we use the retrofit of a wind farm with a battery storage system as an example to show the difference in the market outcomes and revenues obtained from both price-taker and multiscale optimization approaches. Our work goes beyond price-taker and deep dives into quantifying IES-market interaction in optimizing IES. This framework enables users to explore how different design and operation decisions of energy systems interact with the market and provides a more accurate evaluation than the price-taker assumption.

Chen, Xinhe

New Global Characterization of Landslide Exposure

Landslides triggered by intense rainfall are hazards that impact people and infrastructure across the world, but comprehensively quantifying exposure to these hazards remains challenging. Unlike earthquakes or flooding which cover large areas, landslides occur only in highly susceptible parts of a landscape affected by intense rainfall, which may not intersect human settlement or infrastructure. Existing datasets of landslides around the world generally include only those reported to have caused impacts, leading to significant biases toward areas with higher reporting capacity, limiting how our understanding of exposure to landslides in developing countries. In this study, we use an alternative approach to estimate exposure to landslides in a homogenous fashion. We have combined a global landslide hazard proxy derived from satellite data with open-source datasets on population, roads and infrastructure to consistently estimate exposure to rapid landslide hazards around the globe. These exposure models compare favorably with existing datasets of rainfall-triggered landslide fatalities, while filling in major gaps in inventory-based estimates in parts of the world with lower reporting capacity. Our findings provide a global estimate of exposure to landslides from 2001-2019 that we suggest may be useful to disaster mitigation professionals.

earthquakes

What's New in the PSI? (Physical Sciences Informatics)

The NASA Physical Sciences Informatics (PSI) database is NASA’s archival for physical sciences research in microgravity – on ISS and also other reduced gravity platforms. The database has been making data from microgravity physical science investigations publicly available since its launch late 2014. The database was an early investment in Open Science for NASA’s Biophysics and Physical Sciences (BPS) Division of the Science Mission Directorate (SMD), who conducts fundamental and applied physical sciences yielding research data in 6 disciplines: biophysics, combustion science, complex fluids, fluid physics, fundamental physics, and materials science. PSI started initially with data from 14 microgravity investigations. Receipt of data through the years and an annual NRA to fund ground investigations extending flight datasets has increased the total to 72 investigations available with 8 new datasets in the 2020 NRA (NNH20ZDA014N). These new datasets are open for public access at the PSI website at https://www.nasa.gov/PSI.

Physical Sciences Informatics PSI database

Dataset describing two reference models for full-spectral lighting and daylight simulations together with implementations for two software systems

A dataset of two spectral lighting simulation reference models - one office and one factory hall - is presented. It aims to demonstrate and support full-spectral daylight and electric lighting simulations and facilitate evaluation of non-visual effects of light. The dataset includes Rhino CAD geometry, comprehensive spectral material and light source data and window system BSDF data. Example implementations in the two software tools, Radiance and OWL, enable reproducible workflows and support adoption in other software. The dataset is openly available on Zenodo. The office model reproduces Room 518 at the University of Innsbruck, including a west-facing façade and interior furnishings. The factory hall model follows the proposed geometry in the European standard 15193 for building energy performance. Interior reflectances in the office were measured in-situ using a handheld spectrometer. Exterior spectra and factory hall materials matching specified reflectances were obtained from an online spectral materials database. Glazing transmittance was derived from IGDB data using LBNL Optics/WINDOW. BSDFs for venetian blinds at various tilt angles, and for a diffusing pane adapted from the Complex Glazing Database, were generated in WINDOW. Luminaires in both models are specified with photometric files (Eulumdat/IES) and lamp spectra (Fluorescent 840, 4000 K LED). The provided example implementations (Radiance, OWL) include prepared input data and scripts to run first spectral simulations; example results are also included. The dataset is prepared to support reuse by researchers, designers and software developers for method validation, software engineering and comparison, and development of spectral metrics and controls.

Geisler-Moroder, David

Can protein expression be ‘solved’?

Recombinant protein expression is central to biotechnology’s application in academic exploration as well as human health, climate applications and the bioeconomy in general. However, not all proteins can be expressed in all organisms, and the field lacks a predictive model of soluble protein overexpression that could replace laborious experimental trial-and-error. Here, we discuss the state of the field and identify the lack of large, high-fidelity datasets as the primary bottleneck to progress. We review possible assays that could be used for data collection to identify a path toward an extensible experimental platform for collecting soluble recombinant protein overexpression data across organisms. We suggest that the resulting dataset should be used to train increasingly generalizable predictive models of protein expression to answer the question: “How can predictive protein expression be solved?”.

59 BASIC BIOLOGICAL SCIENCES