Search NASASearch

SEARCH · Search NASA

Results for “open datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

SpaceNet 9—Cross-Sensor Alignment of Optical and SAR Imagery

Precise registration of high-resolution synthetic aperture radar (SAR) and optical imagery is necessary for realizing the full potential and benefits of multimodal image analysis. However, two significant challenges presently exist. First, there is a lack of annotated datasets and benchmarks available for high-resolution SAR–optical image registration. Second, an assessment of efficient and reliable image registration methods that can precisely align these modalities is lacking. Here, we present a holistic description of the SpaceNet 9 Challenge and its results. We present a description of the dataset and baseline algorithm along with the results of the challenge, including a description of the winning algorithms. We release the SpaceNet 9 dataset along with open-sourcing the winning algorithms and baseline. The objective of SpaceNet 9 was to compute a dense displacement map that indicates the shift needed to align pixels in an optical image to the pixels in a SAR image. The challenge launched in April 2025 and was active for approximately two months. The top five solutions reduced image alignment error from approximately 34 m to under 13 m for public and private test data, with the best results obtaining a registration error of only 8.5 and 6.7 m on the public testing and private testing dataset, respectively. Usage of pretrained image matching models, robust outlier rejection with RANSAC, and estimating local displacement were common among the top solutions. The results of this challenge provide insight into high-resolution SAR–optical image registration and offer opportunities for future benchmarking in this domain. The baseline algorithm, winning solutions, and datasets are available at https://spacenet.ai/sn9-challenge/.

benchmark datasets

Isolating Unisolated Upsilons with Anomaly Detection in CMS Open Data

We present the first study of anti-isolated Upsilon decays to two muons (ϒ→𝜇⁺⁢𝜇⁻) in proton-proton collisions at the Large Hadron Collider. Using a machine learning (ML)-based anomaly detection strategy, we “rediscover” the ϒ in 13 TeV CMS Open Data from 2016, despite overwhelming anti-isolated backgrounds. We elevate the signal significance to 6.4⁢𝜎 using these methods, starting from 1.6⁢𝜎 using the dimuon mass spectrum alone. Moreover, we demonstrate improved sensitivity from using an ML-based estimate of the multifeature likelihood compared to traditional “cut-and-count” methods. This is the first ever detection of anti-isolated Upsilons, which can be useful in the study of heavy-flavor fragmentation in quantum chromodynamics. Our Letter demonstrates that it is possible and practical to find real signals in experimental collider data using ML-based anomaly detection, and we distill a readily accessible benchmark dataset from the CMS Open Data to facilitate future anomaly detection developments.

machine learning

PNNL INFRARED REFRACTIVE INDEX (n/k) DATASET FOR SEVEN PAH SOLIDS AT ROOM TEMPERATURE

This dataset is an open-source repository of spectral data measured at Pacific Northwest National Laboratory (PNNL). This database provides quantitative values for the complex index of refraction for seven polycyclic aromatic hydrocarbon (PAH) solids. A list of the chemicals is available in the readme file. These spectra consist of the optical constants, i.e., the real, n(ν), and imaginary, k(ν), refractive indices, over the spectral range from 7,800 to 400 cm-1 (1.28 – 25 μm). The conditions under which the individual data were acquired are described in the associated metadata files, and the user is strongly encouraged to read and understand this information to ensure the data are used appropriately for your application. Recommended Citation for Dataset Jessica M Salcido, Jeremy D. Erickson, Ashley M. Bradley, Russell G. Tonkyn, Timothy J. Johnson and Tanya L. Myers. 2026. PNNL INFRARED REFRACTIVE INDEX (n/k) DATASET FOR SEVEN PAH SOLIDS AT ROOM TEMPERATURE. [Data Set] PNNL DataHub. INSERT DOI License Information This work is marked with CC0 1.0: https://creativecommons.org/publicdomain/zero/1.0/. The authors do request that you appropriately cite the dataset when referencing or using the dataset.

Salcido, Jessica Marie Ortola

Improving Data Discovery, Analysis, and Visualizations With Cloud-Based User Services

The Global Hydrometeorology Resource Center (GHRC) Distributed Active Archive Center (DAAC) is one of 12 DAACs managed by the United States National Aeronautics and Space Administration (NASA) Earth Science Data and Information System (ESDIS) project [1]. GHRC and the other DAACs are designed to process, archive, document, and distribute NASA Earth-observing data, ranging from satellite missions to field campaigns [2]. A major goal of the DAACs is to enable science with these data. Science enabling can be difficult as datasets can be very large, use multiple formats, come from numerous platforms, and require three-dimensional visualization. GHRC is using its expertise with cloud-based technologies to develop open source and open science tools to empower users to explore, coincidentally visualize, and analyze multiple datasets. Being open source, the user community can develop visualizations for their own datasets. This presentation will expand on this objective and highlight the capabilities available to the international community now.

GHRC

Bifacial Vertical Testbed and Ground Irradiance Data in Golden, Colorado

This data was collected for Tonita et al., “Vertical bifacial photovoltaic system model validation: study with field data, various orientations, and latitudes,” for validation of optical models for vertically-oriented photovoltaics under high albedo. Ground irradiance data for vertical PV arrays modeling in agrivoltaics is also provided. The dataset is provided for further use or study as open source. For any questions on the dataset, email silvana.ovaitt@nlr.gov.

14 SOLAR ENERGY

Towards robust surrogate models: Benchmarking machine learning approaches to expediting phase field simulations of brittle fracture

Data-driven approaches have the potential to make modeling complex, nonlinear physical phenomena significantly more computationally tractable. For example, computational modeling of fracture is a core challenge where machine learning techniques have the potential to provide a much needed speedup that would enable progress in areas such as multi-scale modeling and uncertainty quantification. Currently, phase field modeling (PFM) of fracture is one such approach that offers a convenient variational formulation to model crack nucleation, branching and propagation. To date, machine learning techniques have shown promise in approximating PFM simulations. While standard fracture benchmarks represent realistic scenarios frequently observed in practice, they typically do not provide sufficiently challenging tests for data-driven methods. Here, to address this gap, we introduce a challenging dataset based on PFM simulations designed to benchmark and advance ML methods for fracture modeling. This dataset includes three energy decomposition methods, two boundary conditions, and 1000 random initial crack configurations for a total of 6000 simulations. Each sample contains 100 time steps capturing the temporal evolution of the crack field. Alongside this dataset, we also implement and evaluate Physics Informed Neural Networks (PINN), Fourier Neural Operators (FNO), and UNet models as baselines, and explore the impact of ensembling strategies on prediction accuracy. With this combination of our dataset and baseline models drawn from the literature we aim to provide a standardized and challenging benchmark for evaluating machine learning approaches to solid mechanics. Our results highlight both the promise and limitations of popular current models, and demonstrate the utility of this dataset as a testbed for advancing machine learning in fracture mechanics research.

Benchmark dataset

Drone Flight Data Logs

This dataset represents the open-air tests for the drones when testing different flight scenarios. For some flights we created and tested with a set of onboard sensors. For others we used the native logs for the drones. We recorded relevant conditions for each of the flights to examine environmental issues and weight impacts. We also looked at segmentations of flights to investigate the energy used in each type of flight.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

A Method for Validating Causal Diagrams of Human Health Risk in Space Flight

The complexity of cause-and-effect relationships between spaceflight hazards and resulting health conditions clouds understanding of the totality of human system risk in space. In response, NASA has introduced Directed Acyclic Graphs (causal diagrams) into the human systems risk management process. These diagrams allow for a common understanding of the mechanisms that lead from unique hazards of spaceflight to the health outcomes important to agencies and astronauts. However, the paucity of available biomedical data from spaceflight creates a need for methods of validating causal models that can accommodate data from spaceflight model analogs. Here we outline one approach utilizing open-access rodent bone datasets from the Ames Life Sciences Data Archive. The properties of directed acyclic graphs themselves can provide an epistemological and statistical framework for validation of a priori causal representations of human system risk in space flight. The assumed causal connections on the graph creates sets of logical implications: variables that – if the causal diagram is correct – should be correlated, as well as sets that should be conditionally independent. By testing these implied correlations and conditional independencies both statistically and heuristically, we can provide evidence for or against specific causal pathways on the causal diagram. In addition to validation of expert-generated causal diagrams, machine learning techniques can learn the most likely structure of a causal diagram from a given dataset. Comparison with and reconciliation between machine-learned causal diagrams and expert-generated diagrams is another technique for challenging assumptions and improving our understanding of causal mechanisms. Accurately representing complex causation is essential to systemic understanding of human health risks in space travel. Having a robust system of validating causal diagrams helps us arrive at more accurate representations of causal systems. This process will be integral to developing the countermeasures necessary for extended exploration of the moon and Mars.

Robert Reynolds

Enabling Biological Discovery Through Biospecimen Sharing: The Nasa Biological Institutional Scientific Collection

Understanding biological impacts from spaceflight hazards and the subsequent development of countermeasures are a high priority to enable humanity to venture back to the Moon, and then to Mars and beyond. Experiments have been conducted with model organisms flown to space and analogous investigations terrestrially, to identify biological mechanistic impacts from spaceflight hazards and to develop mitigation countermeasures, thus contributing towards basic and applied science goals. However, sending organisms into space is a costly endeavor. To maximize scientific return, all biospecimens not required by spaceflight-relevant Principal Investigators are harvested, preserved, and archived in the NASA Biological Institutional Scientific Collection (NBISC). Biospecimens are collected and preserved according to well-established standard operating procedures to maintain scientific quality and are available on-request by the international scientific community. NBISC currently stores over 32,000 biospecimens from Shuttle, International Space Station, and ground-based space analog investigations. Tissue sharing has resulted in at least 33 publications since 2011 and 51 requests since 2016. Many requests for NBISC biospecimens come from first-time investigators who subsequently submit grants as their point-of-entry into the field of spaceflight biology and health. The NBISC biorepository is part of the NASA ‘Open Science for Life in Space’ collaborative group of projects, which includes NASA Genelab, the Space Biology Program’s Biospecimen Sharing Program, Physical Sciences Informatics, and the Ames Life Sciences Data Archive. NBISC biospecimens have been awarded to NASA Genelab, who then generated various open access science ‘omics datasets through the GeneLab Sample Processing laboratory, with resulting data widely used for biological study. Other NBISC biospecimen awards have led to studies on fecal microbiome analysis, DNA damage analysis using single-cell DNA sequencing, enzymatic-pathway identification involved in spaceflight muscle atrophy, and characterization of ocular morphological changes. Of note, NBISC is expanded to include a new Space Microbial Culture Collection (SMCC) for the collection, identification, documentation, long-term preservation, and distribution of space-related microbial isolates.

Biospecimens

Creating Benchmark Data for Artificial Intelligence and Machine Learning Space Biology Research

To identify an appropriate AI/ML approach for a specific problem, the best practice is to measure algorithm performance through the benchmarking process. A scientific benchmark consists of an AI-ready dataset and a reference implementation on a specific scientific question. The NASA Science Mission Directorate (SMD) has started the “Benchmark Initiative for AI/ML to create scientific benchmark datasets in three applications: 1) scientific benchmarking, which finds the best algorithm for a specific problem; 2) application benchmarking, which measures algorithm performance against a set of parameters; and 3) system benchmarking, which evaluates performance of hardware and software architecture. Currently, there are no standardized datasets available to benchmark AI/ML algorithms in the domain of space biology. In this work, we constructed two AI/ML-ready biological datasets from experiments in space-flown mice: cellular imaging and RNA-seq. First, radiation-exposed immune cells harbor DNA damage foci that can be fluorescently marked to visualize the amount of damage following exposure to ionizing radiation. However, such large datasets are difficult to analyze visually, due to imaging inconsistencies and human bias, and classical image processing approaches can fail on imaging artifacts. AI/ML are therefore exciting alternative, providing the speed of machines and the accuracy of humans. We have made this dataset available at https://registry.opendata.aws/bps_microscopy/. Second, high-throughput nucleic acid sequencing (DNA-seq, RNA-seq) has become widespread in biomedical research due to the growing availability and affordability of these assays. However, most sequencing datasets suffer from high dimensionality and low sample count. In this work, we used a generative adversarial network to synthesize a standardized, AI-ready, publicly available benchmark dataset for space biology RNA-seq data with sufficient space-flown and ground control mouse liver samples from NASA GeneLab. This dataset is available at https://registry.opendata.aws/bps_rnaseq/. These datasets are now fully open the Space Biology community to test their favorite AI/ML approaches.

James Casaletto

Mesoscale Phenomena and Their Contribution to the Global Response: A Focus on the Magnetotail Transition Region and Magnetosphere-Ionosphere Coupling

An important question that is being increasingly studied across subdisciplines of Heliophysics is “how do mesoscale phenomena contribute to the global response of the system?” This review paper focuses on this question within two specific but interlinked regions in Near-Earth space: the magnetotail’s transition region to the inner magnetosphere and the ionosphere. There is a concerted effort within the Geospace Environment Modeling (GEM) community to understand the degree to which mesoscale transport in the magnetotail contributes to the global dynamics of magnetic flux transport and dipolarization, particle transport and injections contributing to the storm-time ring current development, and the substorm current wedge. Because the magnetosphere-ionosphere is a tightly coupled system, it is also important to understand how mesoscale transport in the magnetotail impacts auroral precipitation and the global ionospheric system response. Groups within the Coupling, Energetics and Dynamics of Atmospheric Regions Program (CEDAR) community have also been studying how the ionosphere-thermosphere responds to these mesoscale drivers. These specific open questions are part of a larger need to better characterize and quantify mesoscale “messengers” or “conduits” of information—magnetic flux, particle flux, current, and energy—which are key to understanding the global system. After reviewing recent progress and open questions, we suggest datasets that, if developed in the future, will help answer these questions.

transition region

Dataset for "A primer on forest structure measurement with lidar for ecologists"

This repository includes data and code accompanying the case study included in the manuscript "A primer on forest structure measurement with lidar for ecologists" (submitted to Ecosphere). We compiled lidar datasets from multiple platforms in a common area to: 1. Demonstrate how differences in sensor characteristics influence density and resolution of lidar data. 2. Provide open-source, co-located datasets for users to further inspect differences in lidar data. 3. Provide example code to perform basic lidar analysis. This case study is meant to allow readers to get hands-on experience with real-world data from different platforms. This case study is not meant to be a rigorous comparison of derived ecological metrics among all sensors; such comparisons can be found throughout other publications referenced throughout the main manuscript. Code includes basic functions in R commonly used to visualize and manipulate lidar data accessible with a normal laptop computer; more sophisticated algorithms for advanced users are also referenced throughout the main manuscript. Terrestrial laser scanning (TLS), mobile laser scanning (MLS), UAS laser scanning (ULS), airborne laser scanning (ALS), and spaceborne laser scanning (SLS) data were collected within the Smithsonian Environmental Research Center (SERC) forest dynamics plot in Maryland, USA. TLS, MLS, and ALS data were collected within 1 month of the 2021 growing season; ULS data were collected in November 2020 (“leaf-off” data) and July 2022 (“leaf-on” data).

54 ENVIRONMENTAL SCIENCES

Optimizing a Small RNAseq Analysis Pipeline for NASA GeneLab Using Open-Source Tools and Libraries

Small RNA sequencing (small RNAseq) is a powerful tool for studying the regulation of gene expression in various organisms. Small RNAseq has been leveraged in space biology research to study how expression of small RNAs, e.g. micro RNAs (miRNAs), small interfering RNAs (siRNAs), and piwi-interacting RNAs (piRNAs), change upon exposure to the space environment. NASA GeneLab currently hosts small RNAseq raw data derived from space-relevant experiments on the Open Science Data Repository (OSDR). To maximize the accessibility of these data to the scientific community, in addition to hosting raw data, which is only interpretable by bioinformaticians, GeneLab plans to process all small RNAseq datasets and make those processed data available to the scientific community via the OSDR. In this study, we present the development of the GeneLab standardized pipeline for processing small RNAseq datasets. Using human, plant, and synthetic small RNAseq datasets, we interrogate various open-source software and publicly available databases to evaluate their accuracy and reproducibility in each step of the pipeline. For quality control and adapter detection and trimming, we evaluated TrimGalore!, FASTX, SeqKit, and DNApi methods to optimize alignment to reference genomes. We compared BWA, Bowtie, and Bowtie2 to determine the optimal alignment tool. For each alignment tool we also assessed various reference databases, including Ensembl reference genomes and different types of small RNA reference databases, including genome, hairpin, and miRNA references from the miRbase and MirGeneDB databases. To quantify the aligned data, we compared SAMtools, HTSeq, and RSEM for counting alignment events from each alignment tool used. Finally, we evaluated various tools, including DESeq2 and EdgeR, for data normalization and subsequent differential expression analysis. We will present the results from our comparative analyses for each pipeline step and propose a consensus pipeline for processing small RNAseq data derived from various organisms exposed to the space environment.

SmallRNAseq, NASA GeneLab, quality control, adapte

Enhancing Dataset Discovery With Knowledge Graph Link Prediction Techniques

● In the evolving landscape of open science, the ability to navigate and discover pertinent datasets is increasingly significant. This primarily hinges on the presence of detailed metadata, delineating the dataset’s content, and potential spheres of application. ● The GES DISC datasets are characterized by science keywords to enable dataset discovery in web search interfaces. ● A problem may arise where a dataset lacks a science keyword that it otherwise should have. ● Machine learning techniques such as link prediction can be used to detect these missing science keywords by estimating the probability of new links forming between dataset and keyword nodes.

machine learning

Cooperative Transmission Expansion Planning Experiment Data and Results

GO WEST is an open-source power grid modeling framework for U.S. Western Interconnection, which allows users to tailor the model depending on their research study and science questions. It covers 28 balancing authorities (BA) and 12 states in U.S. Western Interconnection. GO WEST allows users to select different number of nodes and come up with a simplified network by utilizing 10,000 nodal topology of U.S. Western Interconnection created by Texas A&M University. Users can try and select different number of nodes, mathematical formulations (linear programming vs. mixed-integer linear programming), transmission line limit scaling factors, and hurdle rate scaling factors. GO WEST offers a unit commitment and economic dispatch (UC/ED) module to simulate grid operations on an hourly scale. In this sense, users can calibrate and validate their model versions by comparing model outputs to historical datasets. TEP is an open-source transmission capacity expansion model, built on GO WEST framework. It utilizes linear programming to optimize transmission capacity addition investment on existing lines within GO WEST framework. In this sense, TEP model only increases the thermal capacity of existing transmission lines and does not add new lines to the system, which leaves the topology preserved. TEP minimizes the total cost of the system which comprises the operational cost of satisfying electricity demand (i.e., generation cost), cost of loss of load (i.e., unserved energy), cost of power flow, and cost of new transmission capacity additions (i.e., investment cost). In order to use TEP model, users need to create scenarios with GO WEST framework. In this analysis, outputs from several models are used to create future inputs to GO WEST and TEP models, including GCAM-USA, TELL, CERF and reV. This dataset includes experiment inputs and outputs from three different transmission expansion scenarios (cooperative, intermediate, and individual) for 2019 and 2059. For 2019, a base scenario to illustrate the default (i.e., historical) power grid operations is also included. This study utilizes rcp45hotter_ssp3 scenario from a previous version of GCAM-USA simulations. Sources of the shapefiles in supplementary data are HIFLD Open and U.S. Energy Atlas. Please see the README file for a detailed description of the main and supplementary data.

Capacity Expansion Model

GeneLab: A Systems Biology Platform for Omics Analysis

NASA GeneLab is an open-access repository for omics datasets generated by biological experiments conducted in space or experiments relevant to spaceflight (e.g. simulated cosmic radiation, simulated microgravity, bed rest studies). The GeneLab Data Systems (GLDS) version 4.0 will be available on October 1st 2019, and will provide the latest in terms of professional state-of-the-art bioinformatics platform for the space biology and radiation community to upload their data into an omics data commons, to process their data with vetted standard workflows and to compare to existing analyses. Started in 2015 as a repository designed to archive omics data from space experiments, GeneLab has expanded its scope to all ionizing radiation omics experiments conducted on the ground and has put considerable effort in providing carefully characterized radiation metadata on all dataset. GeneLab is also providing processed data derived from the raw data covering a large spectrum of omics (genome, epigenome, transcriptome, epitranscriptome, proteome, metabolome) to help users explore important questions: 1) Which genes or proteins are expressed differently in space for various living organisms? 2) What specific DNA mutations or epigenetic changes happen in space or after exposure to ionizing radiation? and 3) How does genetics affect these responses? Processed data available on GeneLab are derived by standard data analysis workflows vetted by hundreds of scientists who volunteered to join one of the four GeneLab Analysis Working Groups (Animal AWG, Plant AWG, Microbe AWG, Multi-Omics AWG). In this presentation, we will discuss how to bridge the gap between irradiation studies performed on earth and biological experiments conducted in space since the early 1990's. We will discuss how radiation dosimetry was estimated for datasets derived from samples collected during the Space Shuttle era or on the International Space Station. Finally, we will address future strategies regarding dose monitoring in future missions into space, inter-agency efforts to unify data under one umbrella, and knowledge dissemination across the radiation research community and the space biology community.

open-science

Machine Learning Approaches to Predicting Induced Seismicity and Imaging Geothermal Reservoir Properties

This project developed machine learning (ML) methods, lab data sets, and field data to advance geothermal exploration and geothermal energy production. The work had three focus areas. One involved the development of ML methods to use microearthquakes (MEQs) for imaging geothermal reservoir properties and improving subsurface characterization – most importantly the evolution of permeability within the evolving reservoir. This part of the work included development of ML approaches for automated MEQ location, focal mechanism determination and identification of earthquake precursors. The second area focused on using MEQ signals generated by geothermal exploration and production to predict the relationship between fluid injection and seismicity. Here, we extended to reservoir scale our success in using ML to predict laboratory earthquakes and fault zone stress state. The third focus area was on lab experiments. Here, we developed new ML models for lab earthquake prediction and identification of precursors to failure to improve earthquake forecasting and early warning in geothermal settings. Major outcomes of our work include ML models that learn from MEQ signals during geothermal exploration and production to predict induced seismicity. MEQs occur naturally in connection with drilling and energy production. We developed ML methods to use the seismic waves from these events to characterize the elastic, hydraulic and poromechanical properties of reservoirs. Our work illuminated fracture geometry and the evolution of fracture permeability by incorporating seismic coda wave analysis and ML methods to relate fluid injection and seismicity. We significantly expanded laboratory earthquake prediction to include methods that use both passive measurements of microearthquakes within the lab fault zones and also active source acoustic measurements of fault zone elastic properties. These methods can now predict fault zone stress state, time to failure and the magnitude of lab earthquakes. Our work showed that repetitive stick- slip failure events during frictional sliding (the lab equivalent of earthquakes) are preceded by a cascade of micro-failure events that radiate energy in a manner that foretells unstable failure – manifest as laboratory MEQs. We documented a mapping between fracture properties and statistical attributes of elastic radiation. We extended existing works to geothermal reservoir scale and developed ML methods to determine reservoir permeability, fracture properties, and their evolution during geothermal energy production. An attractive feature of ML algorithms is their ability to handle big datasets and reveal patterns and correlations that may remain invisible to conventional analyses. Our work connected data from field, laboratory and intermediate scales to study permeability, stress, strength, fracture stiffness and geometry. At the field scale we used data from the Newberry Volcano field site, UtahFORGE, EGS Collab, and also the Bedretto underground research lab in Switzerland. These data sets are bridging the gap between the lab scale, theory, and reservoir scale. Our work produced plain language summaries to improve public understanding of DOE research. We also developed openly distributed ML and seismicity datasets for use by all researchers and we published connections between induced seismicity in geothermal areas and reservoir properties including permeability, fracture properties, and stress state. Our models are designed for the large data sets of induced seismicity typically associated with geothermal sites. We produced labeled event catalogs and used them on geothermal data to assess how ML can facilitate geothermal production and exploration. All datasets are available on the GDR Productivity: The project produced 32 publications in peer reviewed journals (two are in review). It supported the work of 6 PhD students, 40 conference presentations, 6 keynote talks at national meetings, and mentoring and professional development for 4 postdoctoral fellows.

15 GEOTHERMAL ENERGY