Search NASA⌕ Search

SEARCH · Search NASA

Results for “Dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28

Central Valley Water Resources: Improving California Groundwater Assessments using GRACE and InSAR Datasets for Water Resource Management

California’s Central Valley is one of the most productive agricultural areas in the world, producing approximately $20 billion in crops annually. The recent California droughts of 2007-2010 and 2012-2019 resulted in increased groundwater pumping in the Central Valley to adequately irrigate farmland. Overdrafting of the Central Valley aquifer results in groundwater depletion, land subsidence, and permanent loss of groundwater storage. In 2014, depletion of groundwater led the state of California to enact the Sustainable Groundwater Management Act (SGMA),requiring critically overdrafted, high, and medium priority sub-basins to reach sustainable levels of groundwater pumping and recharge by 2042. SGMA allows local Groundwater Sustainability Agencies (GSAs) the authority to create Groundwater Sustainability Plans (GSPs) at the sub-basin level. To assist California’s Department of Water Resources (DWR), this project quantified groundwater change and land subsidence in Central Valley sub-basins with sparse or unreliable well and Geographic Positioning Systems (GPS) data. This was done using NASA’s Gravity Recovery and Climate Experiment (GRACE), GRACE FollowOn (GRACE-FO), and interferograms derived from Sentinel-1 C-band Synthetic Aperture Radar (C-SAR) and Advanced Land Observing Satellite 2 (ALOS-2)Phased Array L-band Synthetic Aperture Radar 2 (PALSAR-2). Time series of the GRACE and InSAR data were compared with well and GPS data in data-dense sub-basins to determine the feasibility of these datasets for groundwater storage and subsidence monitoring. We found that GRACE and InSAR data are effective tools for determining groundwater change and land subsidence and can be used on their own to monitor sub-basins in the absence of well and GPS data

Water Resources↗

Central Valley Water Resources: Improving California Groundwater Assessments using GRACE and InSAR Datasets for Water Resource Management

California’s Central Valley is one of the most productive agricultural areas in the world, producing approximately $20 billion in crops annually. The recent California droughts of 2007-2010 and 2011-2017 resulted in increased groundwater pumping in the Central Valley to adequately irrigate farmland. Overdrafting of the Central Valley aquifer results in groundwater depletion, land subsidence, and permanent loss of groundwater storage. In 2014, depletion of groundwater led the state of California to enact the Sustainable Groundwater Management Act (SGMA), requiring critically overdrafted, high, and medium priority sub-basins to reach sustainable levels of groundwater pumping and recharge by 2042. SGMA allows local Groundwater Sustainability Agencies the authority to create Groundwater Sustainability Plans at the sub-basin level. To assist California’s Department of Water Resources, this project quantified groundwater change and land subsidence in Central Valley sub-basins with sparse or unreliable well and GPS data. This was done using NASA’s Gravity Recovery and Climate Experiment (GRACE), GRACE Follow-On (GRACE-FO), and interferograms derived from Sentinel-1 C-band Synthetic Aperture Radar (C-SAR) and Advanced Land Observing Satellite 2 (ALOS-2) Phased Array L-band Synthetic Aperture Radar 2 (PALSAR-2). Time series of the GRACE and InSAR data were compared with well and GPS data in data-dense sub-basins to determine the feasibility of these datasets for groundwater storage and subsidence monitoring. We found thatGRACE and InSAR data are effective tools for determining groundwater change and land subsidence and can be used on their own to monitor sub-basins in the absence of well and GPS data.

Water Resources↗

Creating a knowledge graph to connect scientific publications and datasets for improving discovery of GES DISC’s data and services

The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) archives and distributes to the public hundreds of Earth Science data collections. These collections are used in research, resulting in thousands of scientific papers published each year. As new users come to GES DISC for the data, it is important for them to understand how these data were used in the prior research. For this we are creating the Knowledge Graph that connects research paper citation and the data collection metadata. The relationships created in the graph have potential for the Web applications that utilize this information to directly connect the paper research to the GES DISC datasets and services. We will demonstrate these relationships using the Web application prototype.

Nathaniel Ross Crosby↗

Synthesizing Disparate LiDAR and Satellite Datasets through Deep Learning to Generate Wall-to-Wall Regional Inventories for the Complex, Mixed-Species Forests of the Eastern United States

Light detection and ranging (LiDAR) has become a commonly-used tool for generating remotely-sensed forest inventories. However, LiDAR-derived forest inventories have remained uncommon at a regional scale due to varying parameters among LiDAR data acquisitions and the availability of sufficient calibration data. Here, we present a model using a 3-D convolutional neural network (CNN), a form of deep learning capable of scanning a LiDAR point cloud, combined with coincident satellite data (spectral, phenology, and disturbance history). We compared this approach to traditional modeling used for making forest predictions from LiDAR data (height metrics and random forest) and found that the CNN had consistently lower uncertainty. We then applied the CNN to public data over six New England states in the USA, generating maps of 14 forest attributes at a 10 m resolution over 85% of the region. Aboveground biomass estimates produced a root mean square error of 36 Mg ha−1 (44%) and were within the 97.5% confidence of independent county-level estimates for 33 of 38 or 86.8% of the counties examined. CNN predictions for stem density and percentage of conifer attributes were moderately successful, while predictions for detailed species groupings were less successful. The approach shows promise for improving the prediction of forest attributes from regional LiDAR data and for combining disparate LiDAR datasets into a common framework for large-scale estimation.

Elias Ayrey↗

Quantifying radiation quality for space relevant radiation types: Fitting excess risk models to three combined HZE-irradiated mouse datasets

Radiation health risks are predominantly derived from low linear energy transfer (LET) terrestrial exposures; however, space radiation includes exposure to high-LET and high-charge, high-energy (HZE) particles. Accurately quantifying the differences in radiation quality between the space and terrestrial radiation environments is important for assessing and predicting health risks for astronauts. Weil et al. 2009 and 2014 used two different inbred mouse strains to study differences in hepatocellular carcinoma (HCC) tumorigenesis after exposures to low- and high- LET radiation. More recently, Edmundson et al. 2020 provided valuable new tumor data in outbred mice that were exposed to low- and high-LET radiation. The present study aims to rigorously investigate a relative biological effectiveness (RBE) factor by leveraging the HCC tumor data from Weil et al. 2009, Weil et al. 2014, and Edmundson et al. 2020. The three experiments were similarly designed, allowing the raw data to be combined into a pooled dataset to estimate excess relative risk (ERR) and excess absolute risk (EAR) models using Bayesian Poisson regression. These effect estimates from the pooled data provide greater power to calculate a data driven RBE. Extensive sensitivity analyses test the robustness of RBE estimates to various model assumptions. The following questions will be explored through the sensitivity analyses: • Is the shape of the dose response different for low-LET radiation and HZE radiation, indicating that RBE is a function of dose? • Does attained age modify the effect estimates differently for low-LET radiation and HZE radiation, indicating RBE is a function of attained age? • Are the effect estimates and RBE estimates different for inbred mouse strains and outbred mouse strains? • Do assumptions about differences in ERR models and EAR models change the estimated RBE? Additional studies would be needed to validate the findings from these exploratory analyses.

Lori J. Chappell↗

Learning spatial response functions from large multi-sensor AIRS and MODIS datasets

We use large datasets from the Atmospheric Infrared Sounder (AIRS) and the Moderate Resolution Imaging Spectroradiometer (MODIS) to derive AIRS spatial response functions and study their potential variations over the mission. The new reconstructed spatial response functions can be used to reduce errors in the radiances in non-uniform scenes and improve products generated using both AIRS and MODIS data. AIRS spatial response functions are distinct for each of its 2378 channels and each of its 90 scan angles. We develop the mathematical model and the optimization framework for deriving spatial response functions for two AIRS channels with low water vapor absorption and various scan angles. We quantify uncertainties in the derived reconstructions and study how they differ from pre-flight spatial response functions. We show that our approach generates reconstructions that agree with the data more accurately compared to pre-flight spatial responses. We derive spatial response functions using data collected during successive dates in order to ascertain the repeatability of the reconstructed spatial response functions. We also compare the derived spatial response functions based on data collected in the beginning, the middle, and at the current state of the mission in order to study changes in reconstructions over time.

Vese, Luminita↗

The northernmost hyperspectral FLoX sensor dataset for monitoring of high-Arctic tundra vegetation phenology and Sun-Induced Fluorescence (SIF)

A hyperspectral field sensor (FloX) was installed in Adventdalen (Svalbard, Norway) in 2019 as part of the Svalbard Integrated Arctic Earth Observing System (SIOS) for monitoring vegetation phenology and Sun-Induced Chlorophyll Fluorescence (SIF) of high-Arctic tundra. This northernmost hyperspectral sensor is located within the footprint of a tower for long-term eddy covariance flux measurements and is an integral part of an automatic environmental monitoring system on Svalbard (AsMovEn), which is also a part of SIOS. One of the measurements that this hyperspectral instrument can capture is SIF, which serves as a proxy of gross primary production (GPP) and carbon flux rates. This paper presents an overview of the data collection and processing, and the 4-year (2019–2021) datasets in processed format are available at: https://thredds.met.no/thredds/catalog/arcticdata/infranor/NINA-FLOX/raw/catalog.html associated with https://doi.org/10.21343/ZDM7-JD72 under a CC-BY-4.0 license. Results obtained from the first three years in operation showed interannual variation in SIF and other spectral vegetation indices including MERIS Terrestrial Chlorophyll Index (MTCI), EVI and NDVI. Synergistic uses of the measurements from this northernmost hyperspectral FLoX sensor, in conjunction with other monitoring systems, will advance our understanding of how tundra vegetation responds to changing climate and the resulting implications on carbon and energy balance.

hyperspectral field sensor↗

An Accelerated Life Testing Dataset for Lithium-Ion Batteries With Constant and Variable Loading Conditions

The dataset repository is organized into three main folders, each containing one group of life cycled battery packs. Within each folder individual battery packs own their dedicated csv file for continuous data logging, which are named with their respective battery pack number. The folders are named: - regular_alt_batteries: Containing one csv file for each battery pack cycled at the same load level or load range throughout lifetime - recommissioned_batteries: Containing one csv file for each battery pack cycled at different load levels at varying life stages - second_life_batteries: Containing one csv file for each second life battery pack cycled at constant current througout the second life

Li-ion Battery↗

A Global Land Cover Training Dataset From 1984 to 2020

State-of-the-art cloud computing platforms such as Google Earth Engine (GEE) enable regional-to-global land cover and land cover change mapping with machine learning algorithms. However, collection of high-quality training data, which is necessary for accurate land cover mapping, remains costly and labor-intensive. To address this need, we created a global database of nearly 2 million training units spanning the period from 1984 to 2020 for seven primary and nine secondary land cover classes. Our training data collection approach leveraged GEE and machine learning algorithms to ensure data quality and biogeographic representation. We sampled the spectral-temporal feature space from Landsat imagery to efficiently allocate training data across global ecoregions and incorporated publicly available and collaborator-provided datasets to our database. To reflect the underlying regional class distribution and post-disturbance landscapes, we strategically augmented the database. We used a machine learning-based cross-validation procedure to remove potentially mis-labeled training units. Our training database is relevant for a wide array of studies such as land cover change, agriculture, forestry, hydrology, urban development, among many others.

Radost Stanimirova↗

Neural Network Atmospheric Correction of Remote Sensing Imagery Over Water Using a Synthetic Dataset

Remote sensing atmospheric correction methods have primarily focused on imagery over land. However, accurate correction over water is important for monitoring and research of aquatic environments. More research in this area is ongoing, though one of the biggest challenges is enough quality data to develop and validate correction methods. This is especially true for neural network (NN) -based models which have shown promise in this area given enough quality data. To address this deficiency of data, we are leveraging a synthetic dataset produced by a model called SWIPE that uses radiative transfer modeling to simulate the atmospheric effects on water-leaving (WL) reflectance to estimate top-of-atmosphere (TOA) reflectance. This allows us to produce almost unlimited pairs of WL reflectance and corresponding TOA reflectance for model training across a variety of atmospheric conditions. We use two approaches for our atmospheric correction model. One uses a conditional variational autoencoder (VAE) to estimate a single WL reflectance value from a single TOA reflectance value. The second is based on a UNET architecture and estimates an array of WL reflectance values from an array of TOA reflectance values. The goal of the second method is to capture atmospheric effects that occur spatially between values within the array as compared to the first method.

deep learning↗

Homogenized Ground-Based and Profile Ozone Datasets From the TOAR-II/HEGIFTOM Project: Methods and Station Trends

Within the framework of the second phase of the Tropospheric Ozone Assessment Report (TOAR-II), it was recognized that an essential first step for deriving accurate trends from the ground-based networks that monitor ozone in the free troposphere is putting the measurements on the same basis with respect to absolute references and processing methods. The relevant procedures are referred to as “harmonization” or “homogenization”. The TOAR II working group, “HEGIFTOM” (Harmonization and Evaluation of Ground-based Instruments for Free-Tropospheric Ozone Measurements), has carried out harmonization for five types of network (Figure below) instruments (mid-1990s to 2020): ozonesondes, commercial aircraft IAGOS landing/takeoff profiles, Fourier-Transform Infrared spectrometer (FTIR), tropospheric Lidar, and Brewer/Dobson Umkehr. First, we summarize the homogenization effort for each network that provides new quality-assessed ozone profile or segment (partial column) data sets, including uncertainty estimates and quality flags. Second, with the HEGIFTOM datasets forming the basis for a global assessment of tropospheric ozone column trends, results derived with various trend detection algorithms will be presented. The ultimate goal is evaluation of the consistency of the calculated trends among different techniques at selected stations and/or regions.

ozone↗

Geophysical Retrievals and Cloud Analyses from Merged Airborne Radiometer Datasets Covering 10–684 GHz

Airborne microwave radiometers provide insight about numerous aspects of Earth’s atmosphere and yield critical validation datasets for spaceborne radiometers. Three radiometers that are important to NASA’s airborne remote-sensing arsenal include: the Advanced Microwave Precipitation Radiometer (AMPR), covering 10–85 GHz; the Conical Scanning Millimeter-wave Imaging Radiometer (CoSMIR), covering 50–183 GHz; and the Compact Scanning Submillimeter-wave Imaging Radiometer (CoSSIR), covering 170–684 GHz. The NASA field campaigns of interest to this study include: the Integrated Precipitation and Hydrology Experiment (IPHEx) in 2014, the Olympic Mountains Experiment and Radar Definition Experiment (OLYMPEX/RADEX) in 2015–2016, the Investigation of Microphysics and Precipitation for Atlantic Coast-Threatening Snowstorms (IMPACTS) in 2020–2023, and the Airborne Lightning Observatory for FEGS and TGFs (ALOFT) in 2023. To provide a more comprehensive perspective on clouds and precipitation observed during these airborne field campaigns, AMPR data were merged spatiotemporally with CoSMIR data for IPHEx, OLYMPEX/RADEX, and IMPACTS (2020 and 2022), while considering differences in instrument characteristics and operations, providing brightness temperature (Tb) values from 10–183 GHz in a common background grid throughout each flight. AMPR and CoSSIR data were similarly merged for IMPACTS (2023) and ALOFT, providing a common background grid with Tb values covering 10–684 GHz throughout each flight. These merged Tb data were employed in geophysical retrievals using the Community Radiative Transfer Model (CRTM), an Eddington radiative transfer model, and a one-dimensional variational (1DVAR) inversion method. Retrievals of cloud liquid water path were of primary interest. This presentation will include an overview of the methods for the radiometer data mergers, the radiative transfer methods, the geophysical retrievals, and detailed results from examining trends in Tb and cloud liquid water path in clouds, precipitation, and cloud-to-precipitation transition zones.

Corey G Amiot↗

Computer Vision Dataset for Aircraft Taxi Operations

The development and democratization of computer vision algorithms are contingent on the availability of high-quality datasets. In this paper, we introduce a database of forward-facing videos from taxiing aircraft as well as the time-correlated flight data at approximately 1 to10 Hz containing aircraft state data and environmental conditions. The video data is sourced from the National Aeronautics and Space Administration Airborne Science Program archive and includes over 33 hours of 4k, 1080p, and 720p video from twenty-two airports around the world. This paper describes the method of the database construction and a brief analysis of its contents.

Ryan Horn↗

The National Climate Database (NCDB): An Unbiased 100-Year Dataset for PV Modeling

In this study, we develop a statistical technique to downscale the future projection of solar irradiance for photovoltaics (PV) energy-related applications. A set of Regional Climate Model (RCM)-based projections obtained from the North American Coordinated Regional Climate Downscaling Experiment (NA-CORDEX) are used as inputs to statistical methods to generate high-resolution global horizontal irradiance (GHI) over the contiguous United States (CONUS). The main steps of the statistical downscaling method include (1) regridding RCM output (0.22 degree and daily resolutions) to handle the modeled-observed data sets on a common grid, (2) correcting bias of RCM GHI using satellite-derived observation, and (3) implementing temporal and spatial downscaling to generate GHI at 8-km and hourly resolution. Basically, complex physical processes and interactions between solar radiation and various atmospheric constituents lead solar irradiance to be highly variable and uncertain. Underrepresentation of clouds from the RCM parameterizations is the main source of error and uncertainty in modeling solar irradiance. Thus, we adapt and use the high-quality satellite-derived data from the National Solar Radiation Database (NSRDB) to analyze the bias and error of RCM GHI as well as estimate the statistical parameters for spatial and temporal downscaling. This presentation will summarize the comprehensive analysis conducted to produce and assess the results under two climate scenarios (RCP4.5 and RCP8.5). We will also present a detailed validation demonstrating the strengths of the downscaling method, a summary of the 100-year dataset from 2001-2100, and future extension of this research.

bias correction↗

VLM4Bio: A Benchmark Dataset to Evaluate Pretrained Vision-Language Models for Trait Discovery from Biological Images

Images are increasingly becoming the currency for documenting biodiversity on the planet, providing novel opportunities for accelerating scientific discoveries in the field of organismal biology, especially with the advent of large vision-language models (VLMs). We ask if pre-trained VLMs can aid scientists in answering a range of biologically relevant questions without any additional fine-tuning. In this paper, we evaluate the effectiveness of 12 state-of-the-art (SOTA) VLMs in the field of organismal biology using a novel dataset, VLM4Bio, consisting of 469K question8 answer pairs involving 30K images from three groups of organisms: fishes, birds, and butterflies, covering five biologically relevant tasks. We also explore the effects of applying prompting techniques and tests for reasoning hallucination on the performance of VLMs, shedding new light on the capabilities of current SOTA VLMs in answering biologically relevant questions using images

Maruf, M [Virginia Tech, Blacksburg]↗

High-Fidelity, Large-Scale, Realistic Dataset Development

The final report summarizes the work performed for supporting the ARPA-E Grid Optimization Competition (Challenge 2 and Challenge 3) within the stated period. Challenge 2 For the challenge period, the main responsibility of the team is to investigate, gen- erate, and deliver parts of the data sets for the competition, based on the competition model for Challenge 2, existing data sets from Challenge 1, and data source supplied by other data set teams. Challenge 3 For the challenge period, the main responsibility of the team is to propose, create, deliver, and maintain the data format during the competition period. The data format will specify how the benchmark data will be represented and communicated to competitors. It will also specify how competitors should report back the solutions. The data format will be closely aligned with the problem formulation (maintained by the formulation team) and the solution validation process (maintained by the validation team). Our team is also responsible in investigating, generating, and delivering parts of the data sets for the competition. The data sets will be created based on the competition model for Challenge 3, existing data sets from Challenge 1 and Challenge 2, and data source supplied by other data set teams.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

New Particle Formation Event Dataset at the Southern Great Plains (SGP) Observatory from 2018 to 2023

This data set contains observations of new particle formation (NPF) events collected at the U.S. Department of Energy’s Atmospheric Radiation Measurement (ARM) Southern Great Plains (SGP) observatory from 2018 to 2023. Measurements include particle number size distributions, radiance measurements, and associated meteorological variables from onsite instrumentation relevant for identifying and characterizing NPF events. Events were identified using standardized criteria and documented to support investigations of aerosol nucleation, growth dynamics, and their interactions with local atmospheric conditions. The data set provides a multi-year record that enables evaluation of seasonal and interannual variability in NPF occurrence and intensity at a mid-continental site. These data are intended to support studies of aerosol-cloud-climate interactions and model evaluation within both ARM and the broader atmospheric science community. More information can be found within the README_NPF_SGP_2018_2023.docx file.

event↗