Search NASA⌕ Search

SEARCH · Search NASA

Results for “Dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

Quantifying radiation quality for space relevant radiation types: Fitting excess risk models to three combined HZE-irradiated mouse datasets

Radiation health risks are predominantly derived from low linear energy transfer (LET) terrestrial exposures; however, space radiation includes exposure to high-LET and high-charge, high-energy (HZE) particles. Accurately quantifying the differences in radiation quality between the space and terrestrial radiation environments is important for assessing and predicting health risks for astronauts. Weil et al. 2009 and 2014 used two different inbred mouse strains to study differences in hepatocellular carcinoma (HCC) tumorigenesis after exposures to low- and high- LET radiation[1-2]. More recently, Edmondson et al. 2020 provided valuable new tumor data in outbred mice that were exposed to low- and high-LET radiation[3]. The present study aims to rigorously investigate a relative biological effectiveness (RBE) factor by leveraging the HCC tumor data from the three datasets[1-3]. The three experiments were similarly designed, allowing the raw data to be combined into a pooled dataset to estimate excess relative risk (ERR) and excess absolute risk (EAR) models using Bayesian Poisson regression.

Lori J. Chappell↗

Development of the Ames Global Hyperspectral Synthetic Dataset

This study develops the surface BRDF (bidirectional reflectance distribution function) product of the Ames Global Hyperspectral Synthetic Dataset (AGHSD), based on the corresponding MODIS products, to support the NASA Surface Biology and Geology mission development. A main challenge in deriving a hyperspectral dataset from the multi-band satellite products is how to identify a succinct yet robust algorithm that allow us to infer BRDF at unobserved wavelengths based on the few observed bands. Using the theories of radiative transfer in vegetation canopies, we arrive at a simple equation that accurately approximates hyperspectral surface BRDF as the weighted sum of components from the soil and the vegetation. Each of the components is modeled by the product of the spectrally-dependent optical properties of a surface element (the spectra of the soil surface reflectance, the leaf single albedo, or the canopy scattering coefficient) and a spectrally-independent bidirectional scattering function. The optical properties of the soil and the vegetation can be obtained from existing spectral libraries or model simulations. The bidirectional scattering functions are represented by the Ross-Thick-Li-Sparse BRDF model, where the linear coefficients are estimated with regression analysis from the multi-band MODIS data. We validate the algorithm with simulations by Monte Carlo Ray Tracing model experiments, and the results are highly consistent with the theoretic derivation. We apply the algorithm to generate the AGHSD BRDF product at 1km and 8-day resolutions for the year of 2019. The results are biogeochemically and physically coherent and consistent, and thus serve the goal to support the science and application development of the SBG community.

Hyperspectral↗

Newly Released GPCP Version 3.2 Global Precipitation Datasets at NASA GES DISC

The Global Precipitation Climatology Project (GPCP) is the precipitation component of an internationally coordinated set of (mainly) satellite-based global products dealing with the Earth's water and energy cycles, under the auspices of the Global Water and Energy Experiment (GEWEX) Data and Assessment Panel (GDAP) of the World Climate Research Program. As the follow-on to the GPCP Version 2.X products, GPCP Version 3 (GPCP V3.2) seeks to continue the production of long, homogeneous precipitation record using modern input and calibration datasets. The GPCP V3.2 provides globally complete analyses of surface precipitation on a 0.5°x 0.5° latitude/longitude grid at both monthly and daily intervals, respectively covering 1983 to the present and June 2000 to the present. New data fields have been introduced to better characterize the precipitation, particularly including an estimate of the fraction of the precipitation that is liquid (rain) in both the Monthly and Daily, and a Quality Index for the Monthly. Compared to the operational GPCP V2.3 Monthly, the V3.2 Monthly provides a more reasonable climatology in the Southern Ocean, and increases the global average precipitation by about 4.46%, which is in line with recommendations of recent assessments. However, the two versions have comparable global and regional trends for 1983-2020. Compared to the operational One-Degree Daily (Version 1.3) product, the V3.2 Daily better represents the histogram of precipitation rates, particularly at high values. In this presentation, we will present the latest GPCP V3.2 daily and monthly datasets archived and distributed at NASA Goddard Earth Sciences (GES) Data and Information Services Center (DISC) along with examples from NASA Giovanni.

Precipitation↗

Spatially-Coordinated Airborne Data and Complementary Products for Aerosol, Gas, Cloud, and Meteorological Studies: the Nasa Activate Dataset

The NASA Aerosol Cloud meTeorology Interactions oVer the western ATlantic Experiment (ACTIVATE) produced a unique dataset for research into aerosol–cloud–meteorology interactions, with applications extending from process-based studies to multi-scale model intercomparison and improvement as well as to remote-sensing algorithm assessments and advancements. ACTIVATE used two NASA Langley Research Center aircraft, a HU-25 Falcon and King Air, to conduct systematic and spatially coordinated flights over the northwest Atlantic Ocean, resulting in 162 joint flights and 17 other single-aircraft flights between 2020 and 2022 across all seasons. Data cover 574 and 592 cumulative flights hours for the HU-25 Falcon and King Air, respectively. The HU-25 Falcon conducted profiling at different level legs below, in, and just above boundary layer clouds (< 3 km) and obtained in situ measurements of trace gases, aerosol particles, clouds, and atmospheric state parameters. Under cloud-free conditions, the HU-25 Falcon similarly conducted profiling at different level legs within and immediately above the boundary layer. The King Air (the high-flying aircraft) flew at approximately ∼ 9 km and conducted remote sensing with a lidar and polarimeter while also launching dropsondes (785 in total). Collectively, simultaneous data from both aircraft help to characterize the same vertical column of the atmosphere. In addition to individual instrument files, data from the HU-25 Falcon aircraft are combined into “merge files” on the publicly available data archive that are created at different time resolutions of interest (e.g., 1, 5, 10, 15, 30, 60 s, or matching an individual data product's start and stop times). This paper describes the ACTIVATE flight strategy, instrument and complementary dataset products, data access and usage details, and data application notes. The data are publicly accessible through https://doi.org/10.5067/SUBORBITAL/ACTIVATE/DATA001 (ACTIVATE Science Team, 2020).

Aerosol Cloud meTeorology Interactions oVer the we↗

GLORIA - A Globally Representative Hyperspectral In Situ Dataset for Optical Sensing of Water Quality

The development of algorithms for remote sensing of water quality (RSWQ) requires a large amount of in situ data to account for the bio-geo-optical diversity of inland and coastal waters. The GLObal Reflectance community dataset for Imaging and optical sensing of Aquatic environments (GLORIA) includes 7,572 curated hyperspectral remote sensing reflectance measurements at 1 nm intervals within the 350 to 900 nm wavelength range. In addition, at least one co-located water quality measurement of chlorophyll a , total suspended solids, absorption by dissolved substances, and Secchi depth, is provided. The data were contributed by researchers affiliated with 59 institutions worldwide and come from 450 different water bodies, making GLORIA the de-facto state of knowledge of in situ coastal and inland aquatic optical diversity. Each measurement is documented with comprehensive methodological details, allowing users to evaluate fitness-for-purpose, and providing a reference for practitioners planning similar measurements. We provide open and free access to this dataset with the goal of enabling scientific and technological advancement towards operational regional and global RSWQ monitoring.

remote sensing of water quality↗

A Machine Learning Ready Dataset of Acoustic Power Maps for Detection of Active Region Emergence

The development of an accurate forecast for solar eruptive activity has become increasingly important in order to prevent any potential impact on activities in space and the Earth's environment. It is therefore crucial to detect active regions before they appear on the solar surface and create early warning capabilities for upcoming Space Weather disturbances. In this work, 9TB of solar data (SDO/HMI dopplergrams, magnetograms and continuum intensity maps) involving the emergence of 61 NOAA solar active regions since 2010 were processed using the NASA HECC capabilities. An acoustic power maps time-series dataset was created (for four different frequency ranges and processed to take into account the solar sphere geometric effect ) which can be used for understanding the dynamics of the solar surface and train a variety of ML models. The calculated acoustic power maps carry precursor information associated with the decrease in continuum intensity on the solar surface, verifying older helioseismology research. Our results show that a Long Short-Term Memory (LSTMs) model, with a modest layer depth and the right hyperparameters tuned, when trained on this solar acoustic power maps dataset can predict without false negatives a drop in intensity (associated with the emergence of the active region), up to 18 hours in advance.

SMD↗

Cloud and Precipitation Analyses using Merged Datasets from Two Airborne Microwave Radiometers Covering 10–183 GHz

Microwave radiometers provide valuable insight into the structure and characteristics of clouds and precipitation. In NASA’s airborne remote-sensing arsenal, the Advanced Microwave Precipitation Radiometer (AMPR) and the Conical Scanning Millimeter-wave Imaging Radiometer (CoSMIR) have been used extensively in field campaigns throughout the world. AMPR operates with four channels between 10.7 and 85.5 GHz, while CoSMIR operates with nine channels ranging from 50.3 to 183.31 GHz. Although these datasets provide key information when used separately, the combination of these radiometers covers virtually the full range of frequencies used by the Global Precipitation Measurement (GPM) Microwave Imager (GMI), enabling suborbital observations to compare with GPM spaceborne measurements. This presentation will detail the merger of AMPR and CoSMIR data during two NASA airborne field campaigns: the Integrated Precipitation and Hydrology Experiment (IPHEx) and the Olympic Mountains Experiment / Radar Definition Experiment (OLYMPEX/RADEX). Other field campaigns, such as the Investigation of Microphysics and Precipitation for Atlantic Coast-Threatening Snowstorms (IMPACTS), may be included as well. Using the merged brightness temperature dataset, features of precipitating and non-precipitating clouds containing liquid and/or ice hydrometeors will be discussed from selected flight segments. Geophysical retrievals derived from these brightness temperatures using a one-dimensional variational (1DVAR) technique and/or multi-linear regression equations will also be employed in these analyses. Observed transitions between precipitating and non-precipitating systems will be explored in greater detail. Dropsonde data will be used to provide environmental contexts throughout each flight, and additional observations (e.g., from airborne and/or land-based radar) will be incorporated to supplement the radiometer-based results. The broader implications of these results and pathways for future work will also be discussed.

Corey G. Amiot↗

Unveiling the Transferability of PLSR Models for Leaf Trait Estimation: Lessons from a Comprehensive Analysis with a Novel Global Dataset

Leaf traits are essential for understanding many physiological and ecological processes. Partial least-squares regression (PLSR) models with leaf spectroscopy are widely applied for trait estimation, but their transferability across space, time and plant functional types (PFTs) remains unclear. We compiled a novel dataset of paired leaf traits and spectra, with 47,393 records for >700 species and eight PFTs at 101 globally-distributed locations across multiple seasons. Using this dataset, we conducted an unprecedented comprehensive analysis to assess the transferability of PLSR models in estimating leaf traits. While PLSR models demonstrate commendable performance in predicting chlorophyll content, carotenoid, leaf water and leaf mass per area prediction within their training data space, their efficacy diminishes when extrapolating to new contexts. Specifically, extrapolating to locations, seasons, and PFTs beyond the training data leads to reduced R 2 (0.12-0.49, 0.15-0.42, and 0.25-0.56) and increased NRMSE (3.58-18.24%, 6.27-11.55% and 7.0-33.12%) compared to nonspatial random cross-validation (NRCV). The results underscore the importance of incorporating greater spectral diversity in model training to boost its transferability. These findings highlight potential errors in estimating leaf traits across large spatial domains, diverse PFTs and time due to biased validation schemes and provide guidance for future field sampling strategies and remote sensing applications.

Leaf traits↗

Combining Large Datasets - Cancer Moonshot Task Group Final Summary

In February 2022, President Biden re-ignited the Cancer Moonshot with bold new goals: to reduce the cancer death rate by half within 25 years and improve the lives of people with cancer and cancer survivors. To achieve these ambitious goals, the White House convened the first-ever Cancer Cabinet, bringing together departments and agencies from across the federal government to end cancer as we know it.The Cancer Cabinet convened three task forces and supporting task groups, including the Data and Innovation Task Force, which supported the Cancer Moonshot priority to “Deliver innovation to patients and communities.” In early 2023, the “Combining Large Datasets” (CoLD) Task Group was created within the Data and Innovation Task Force. The scope of the CoLD Task Group was how federal agencies combine large datasets for broad applications across cancer prevention and control, including nutrition, epidemiology, and military/Veteran health. Within this scope, the group sought to better leverage the immense potential of data and power of data tools to increase our understanding of cancer incidence, causes, mortality, treatments, prevention, outcomes, costs, and all other aspects of the burden of cancer.

data integration↗

Nasa Genelab - Knowledge Graph Fabric Enables Deep Biomedical Analysis of Multi-Omics Datasets

The limited number of astronauts and human samples from long-duration space missions pose significant challenges for studying the health risks associated with spaceflight and developing new treatments. As a result, much of our understanding of the biological impact of space travel relies on samples from model organisms. NASA GeneLab, integrated into Open Science Data Repository (OSDR) is a centralized multi-omics resource containing almost 1000 datasets from over 500 space-related studies from human and model organism samples. Previous studies have demonstrated that human phenotypes and physiological changes caused by spaceflight can be identified by connecting gene expression data from model organisms flown in space to a biomedical knowledge graph (SPOKE). In this work, we present a data fabric connecting OSDR datasets to SPOKE that empowers biomedical analyses through the GeneLab visualization portal. This collaboration is funded by NSF’s Proto-OKN program.

data fabric↗

Advanced Interactive 3D Visualization Tool for Customizable Analyses of Tomography Datasets in Material Science

Current methods for visualizing and analyzing 3D tomography datasets in materials science often lack the interactivity and depth required for detailed structural insights. This limitation restricts a researchers' ability to accurately interpret complex data, which is critical for advancing material innovations and understanding structural properties. To address this issue, we have developed a novel, web-based interactive 3D visualization and analysis tool from the Trame framework that offers customizable features to enhance data interpretability. The tool allows users to adjust parameters such as visible range, slice planes, data rotation, and layering, providing a more detailed and dynamic view of complex structures. Its user-friendly web interface increases the accessibility and ease of use for both novice and experienced researchers, to visualize large volumetric datasets. The tool supports a diverse range of data formats, making it versatile for various research applications. Unique capabilities include real-time data manipulation, automated feature detection, context-sensitive feedback, and real-time volume calculations and distributions per sliced region or layer, alongside the ability to quickly generate high-quality screenshots and videos for presentations and reports. These advancements offer a comprehensive solution for enhanced 3D data exploration, significantly improving the analysis process and communication of results in materials science.

36 - MATERIALS SCIENCE↗

Creating a Training Dataset for Semantic Segmentation of Canal Networks for Irrigation Modernization

Canal infrastructure has provided critical irrigation water to the western United States for over a century. To continue providing vital water resources to the semi-arid West, irrigation systems must undergo maintenance and modernization. Many canal companies are resource-constrained, and because funding opportunities often require detailed knowledge of existing infrastructure, they can struggle to secure financial capital. We address this problem by creating training data for a semantic segmentation deep learning model to map canal networks throughout the western United States. To create a diverse and robust training dataset, we labelled 1-m NAIP imagery with the locations of no canals, wet canals, and dry/vegetated canals. Since creating these datasets is time consuming, we first developed a preprocessing methodology to identify canals within our four study areas. We used NAIP imagery and provided canal centerline data to buffer, standardize, and cluster the imagery, automating the labeling process as much as possible. However, this still required manual cleaning and manual classification of canal type. Challenges arose when canals were interrupted (e.g., road culverts or piped sections) or when nearby features shared similar characteristics (e.g., irrigated fields, trees, and shadows). Combining automated preprocessing with manual refinement produced four detailed canal masks to be used in the semantic segmentation model developed by Richard Tapia.

13 - HYDRO ENERGY↗

OSW Consortium 2 - Validated National Offshore Wind Resource Dataset with Uncertainty Quantification (CRADA Report)

This research has led to the development of the 2023 National Offshore Wind data set (NOW-23), which offers the latest wind resource information for offshore regions in the United States. NOW-23 supersedes, for its offshore component, the Wind Integration National Dataset (WIND) Toolkit, which was published a decade ago and is currently a primary resource for wind resource assessments and grid integration studies in the contiguous United States. By incorporating advancements in the Weather Research and Forecasting (WRF) model, NOW-23 delivers an updated and cutting-edge product to stakeholders. As part of this project, we also developed a summary of the uncertainty quantification in NOW-23, along with NOW-WAKES, a 1-year post-construction data set that quantifies expected offshore wake effects in the US Mid-Atlantic lease areas. Stakeholders can access the NOW-23 data set at https://doi.org/10.25984/1821404.

17 WIND ENERGY↗

Cambium Datasets [Slides]

This is a presentation deck used in the "Powered By Cambium" webinar, given on August 13th 2024, in which Pieter Gagnon gave an overview of NREL's Cambium datasets.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

High-Resolution Computed Tomography Dataset of Mount Simon Sandstone

The Illinois Basin is a critical structure for subsurface energy related activities and their implementation in the United States. The Mount Simon Sandstone has been identified as a storage target for permanent and transient storage of fluids in the basin. Known for its exceptional thickness, depth, porosity, and sealing properties of overlying formations, this saline reservoir is crucial for long-term subsurface energy efforts. We present an extensive Computed Tomography (CT) dataset on a high porosity and permeability zone in the lower Mount Simon Sandstone available on the Energy Data eXchange® (EDX). This publicly accessible database comprises over 500 GB of high-resolution CT scans of six core samples, with resolutions ranging from 14.8 µm to 0.7 µm per pixel. The scans include both dry sandstone samples and those saturated with multiple fluids, allowing for comparative analyses across different conditions and resolutions. Coarser scans capture the bedding structure of the sandstone, while finer resolutions reveal detailed pore infill and throat characteristics. Metadata on location, depth, and saturation state enhance usability, enabling quick identification and cross-sample comparisons. By providing a robust resource for research and collaboration, the database contributes to domestic energy advancement by supporting continued progress in the use of the subsurface for energy solutions.

characterization↗

A Curated Dataset of Regional Meteor Events with Simultaneous Optical and Infrasound Observations (2006–2011)

We present a curated, openly accessible dataset of 71 regional meteor events simultaneously recorded by optical and infrasound instrumentation between 2006 and 2011. These events were captured during an observational campaign using the all-sky cameras of the Southern Ontario Meteor Network and the co-located Elginfield Infrasound Array. Each entry provides optical trajectory measurements, infrasound waveforms, and atmospheric specification profiles. The integration of optical and acoustic data enables robust linkage between observed acoustic signals and specific points along meteor trajectories, offering new opportunities to examine shock wave generation, propagation, and energy deposition processes. This release fills a critical observational gap by providing the first validated, openly accessible archive of simultaneous optical–infrasound meteor observations that supports trajectory reconstruction, acoustic propagation modeling, and energy deposition analyses. By making these data openly available in a structured format, this work establishes a durable reference resource that advances reproducibility, fosters cross-disciplinary research, and underpins future developments in meteor physics, atmospheric acoustics, and planetary defense.

astrometry↗

Strong Regional Influence of Climatic Forcing Datasets on Global Crop Model Ensembles

We present results from the Agricultural Model Intercomparison and Improvement Project (AgMIP) Global Gridded Crop Model Intercomparison (GGCMI) Phase I, which aligned 14 global gridded crop models (GGCMs) and 11 climatic forcing datasets (CFDs) in order to understand how the selection of climate data affects simulated historical crop productivity of maize, wheat, rice and soybean. Results show that CFDs demonstrate mean biases and differences in the probability of extreme events, with larger uncertainty around extreme precipitation and in regions where observational data for climate and crop systems are scarce. Countries where simulations correlate highly with reported FAO national production anomalies tend to have high correlations across most CFDs, whose influence we isolate using multi-GGCM ensembles for each CFD. Correlations compare favorably with the climate signal detected in other studies, although production in many countries is not primarily climate-limited (particularly for rice). Bias-adjusted CFDs most often were among the highest model-observation correlations, although all CFDs produced the highest correlation in at least one top-producing country. Analysis of larger multi-CFD-multi-GGCM ensembles (up to 91 members) shows benefits over the use of smaller subset of models in some regions and farming systems, although bigger is not always better. Our analysis suggests that global assessments should prioritize ensembles based on multiple crop models over multiple CFDs as long as a top-performing CFD is utilized for the focus region.

Agricultural Model Intercomparison and Improvement↗