Search NASASearch

SEARCH · Search NASA

Results for “Dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

A comparison between general circulation model simulations using two sea surface temperature datasets for January 1979

Simulations with the UCLA atmospheric general circulation model (AGCM) using two different global sea surface temperature (SST) datasets for January 1979 are compared. One of these datasets is based on Comprehensive Ocean-Atmosphere Data Set (COADS) (SSTs) at locations where there are ship reports, and climatology elsewhere; the other is derived from measurements by instruments onboard NOAA satellites. In the former dataset (COADS SST), data are concentrated along shipping routes in the Northern Hemisphere; in the latter dataset High Resolution Infrared Sounder (HIRS SST), data cover the global domain. Ensembles of five 30-day mean fields are obtained from integrations performed in the perpetual-January mode. The results are presented as anomalies, that is, departures of each ensemble mean from that produced in a control simulation with climatological SSTs. Large differences are found between the anomalies obtained using COADS and HIRS SSTs, even in the Northern Hemisphere where the datasets are most similar to each other. The internal variability of the circulation in the control simulation and the simulated atmospheric response to anomalous forcings appear to be linked in that the pattern of geopotential height anomalies obtained using COADS SSTs resembles the first empirical orthogonal function (EOF 1) in the control simulation. The corresponding pattern obtained using HIRS SSTs is substantially different and somewhat resembles EOF 2 in the sector from central North America to central Asia. To gain insight into the reasons for these results, three additional simulations are carried out with SST anomalies confined to regions where COADS SSTs are substantially warmer than HIRS SSTs. The regions correspond to warm pools in the northwest and northeast Pacific, and the northwest Atlantic. These warm pools tend to produce positive geopotential height anomalies in the northeastern part of the corresponding oceans. Both warm pools in the Pacific produce large-scale circulation anomalies with a pattern that resembles that obtained using COADS SSTs as well as EOF 1 of the control simulation; the warm pool in the Atlantic does not. These results suggest that the differences obtained with COADS SSTs and HIRS SSTs are mostly due to the differences in the datasets over the northern Pacific. There was a blocking episode near Greenland in late January 1979. Both simulations with warm SST anomalies over the northwest and northeast Pacific show a tendency toward increased incidence of North Atlantic blocking; the simulation with warm SST anomalies over the northwest Atlantic shows a tendency toward decreased incidence. These results suggest that features in both SST datasets that do not have a counterpart in the other dataset contribute signficantly to the differences between the simulated and observed fields. The results of this study imply that uncertainties in current SST distributions for the world oceans can be as important as the SST anomalies themselves in terms of their impact on the atmospheric circulation. Caution should be exercised, therefore, when linking anomalous circulation and SST patterns, especially in long-range prediction.

Ose, Tomoaki

The Transition of NASA EOS Datasets to WFO Operations: A Model for Future Technology Transfer

The collocation of a National Weather Service (NWS) Forecast Office with atmospheric scientists from NASA/Marshall Space Flight Center (MSFC) in Huntsville, Alabama has afforded a unique opportunity for science sharing and technology transfer. Specifically, the NWS office in Huntsville has interacted closely with research scientists within the SPORT (Short-term Prediction and Research and Transition) Center at MSFC. One significant technology transfer that has reaped dividends is the transition of unique NASA EOS polar orbiting datasets into NWS field operations. NWS forecasters primarily rely on the AWIPS (Advanced Weather Information and Processing System) decision support system for their day to day forecast and warning decision making. Unfortunately, the transition of data from operational polar orbiters or low inclination orbiting satellites into AWIPS has been relatively slow due to a variety of reasons. The ability to integrate these high resolution NASA datasets into operations has yielded several benefits. The MODIS (MODerate-resolution Imaging Spectrometer ) instrument flying on the Aqua and Terra satellites provides a broad spectrum of multispectral observations at resolutions as fine as 250m. Forecasters routinely utilize these datasets to locate fine lines, boundaries, smoke plumes, locations of fog or haze fields, and other mesoscale features. In addition, these important datasets have been transitioned to other WFOs for a variety of local uses. For instance, WFO Great Falls Montana utilizes the MODIS snow cover product for hydrologic planning purposes while several coastal offices utilize the output from the MODIS and AMSR-E instruments to supplement observations in the data sparse regions of the Gulf of Mexico and western Atlantic. In the short term, these datasets have benefited local WFOs in a variety of ways. In the longer term, the process by which these unique datasets were successfully transitioned to operations will benefit the planning and implementation of products and datasets derived from both NPP and NPOESS. This presentation will provide a brief overview of current WFO usage of satellite data, the transition of datasets between SPORT and the N W S , and lessons learned for future transition efforts.

Darden, C.

The SUMup Dataset: Compiled Measurements of Surface Mass Balance Components over Ice Sheets and Sea Ice with Analysis over Greenland

Increasing atmospheric temperatures over ice cover affect surface processes, including melt, snowfall, and snow density. Here, we present the Surface Mass Balance and Snow on Sea Ice Working Group (SUMup) dataset, a standardized dataset of Arctic and Antarctic observations of surface mass balance components. The July 2018 SUMup dataset consists of three subdatasets, snow/firn density (https://doi.org/10.18739/A2JH3D23R), at least near-annually resolved snow accumulation on land ice (https://doi.org/10.18739/A2DR2P790), and snow depth on sea ice (https://doi.org/10.18739/A2WS8HK6X), to monitor change and improve estimates of surface mass balance. The measurements in this dataset were compiled from field notes, papers, technical reports, and digital files. SUMup is a compiled, community-based dataset that can be and has been used to evaluate modeling efforts and remote sensing retrievals. Active submission of new or past measurements is encouraged. Analysis of the dataset shows that Greenland Ice Sheet density measurements in the top 1m do not show a strong relationship with annual temperature. At Summit Station, Greenland, accumulation and surface density measurements vary seasonally with lower values during summer months. The SUMup dataset is a dynamic, living dataset that will be updated and expanded for community use as new measurements are taken and new processes are discovered and quantified.

Montgomery, Lynn

Improving Earth Science Dataset Search with Publication

The NASA Goddard Earth Sciences Data and Information Services Center (GESDISC) archives a large number of Earth observational datasets. Thousands of the publications are created each year based on these datasets. The content of these publications can be used for discovery of the datasets based on the characteristics of applicational research. We leverage the content of these publications to retrieve the information about phenomena and domains where measurements from the datasets were utilized through linking these publications and dataset in Knowledge Graph. We retrieve phenomena and domain information using SWEET ontology and produce the set of keywords that are linked to the datasets. Further, we evaluate this link strength according to the frequency of dataset usage in the papers mentioning these keywords. We demonstrate how this linkage can improve dataset search by comparing the search results obtained from Common Metadata Repository (CMR) search and the publications based data.

Kristina Stoyanova

Calculation of top-of-atmosphere, surface and atmospheric cloud radiative kernels and feedbacks based on ISCCP-H datasets

This study aims to create observation-based cloud radiative kernel (CRK) datasets and evaluate them by direct comparison of CRK and the CRK-derived cloud feedback datasets. Based on the International Satellite Cloud Climatology Project (ISCCP) H datasets, we calculate CRKs (called FH CRKs) as 2D joint function/histogram of cloud optical depth and cloud top pressure for shortwave, longwave, and their sum, Net, at the top of atmosphere (TOA), as well as, for the first time, at the surface (SFC) and in the atmosphere (ATM). The direct comparison shows that FH agrees reasonably well with three other TOA CRK datasets. With cloud fraction change (CFC) datasets of the same histogram for doubled-CO2 simulation from 10 CFMIP1 models, we derive all the TOA, SFC and ATM cloud feedback using the FH CRKs. Our TOA cloud feedback is highly similar to the previous counterparts. Based on the comparison for the 4 CRK datasets and the 10 CFC datasets, we estimate the uncertainty budget for the CRK-derived cloud feedback and show that the CFC-associated uncertainty contributes > 98.5% of the total cloud feedback uncertainty while CRK’s is very small. Our preliminary evaluation shows that some near-zero/small cloud feedback in the TOA-alone feedback indeed results from the compensation of sizable cloud feedback of the SFC and ATM feedback, demonstrating that they help reveal some significant surface and atmospheric cloud feedback whose sum appears insignificant in TOA-alone feedback

cloud radiative kernel (CRK) datasets

Analyzing EOSDIS Dataset Research Outputs using Knowledge Graphs and Large Language Models

Datasets, unlike publications, can be updated over time, with each new version receiving a DOI but not always being linked to previous ones. This complicates tracking citations across a dataset’s lifecycle. We address this by integrating dataset versions and citations into a knowledge graph (KG), which helps trace dataset citations and analyze dataset usage in applied research. To categorize publications from various journals, we fine-tuned NASA IMPACT INDUS Large Language Model (LLM) on a labeled publication set, assigning publications to one of twenty applied research areas. By linking datasets to these research areas, we improved dataset searchability and discovery through these domains.

open-source

Optimizing tertiary storage organization and access for spatio-temporal datasets

We address in this paper data management techniques for efficiently retrieving requested subsets of large datasets stored on mass storage devices. This problem represents a major bottleneck that can negate the benefits of fast networks, because the time to access a subset from a large dataset stored on a mass storage system is much greater that the time to transmit that subset over a network. This paper focuses on very large spatial and temporal datasets generated by simulation programs in the area of climate modeling, but the techniques developed can be applied to other applications that deal with large multidimensional datasets. The main requirement we have addressed is the efficient access of subsets of information contained within much larger datasets, for the purpose of analysis and interactive visualization. We have developed data partitioning techniques that partition datasets into 'clusters' based on analysis of data access patterns and storage device characteristics. The goal is to minimize the number of clusters read from mass storage systems when subsets are requested. We emphasize in this paper proposed enhancements to current storage server protocols to permit control over physical placement of data on storage devices. We also discuss in some detail the aspects of the interface between the application programs and the mass storage system, as well as a workbench to help scientists to design the best reorganization of a dataset for anticipated access patterns.

Chen, Ling Tony

Climate Forcing Datasets for Agricultural Modeling: Merged Products for Gap-Filling and Historical Climate Series Estimation

The AgMERRA and AgCFSR climate forcing datasets provide daily, high-resolution, continuous, meteorological series over the 1980-2010 period designed for applications examining the agricultural impacts of climate variability and climate change. These datasets combine daily resolution data from retrospective analyses (the Modern-Era Retrospective Analysis for Research and Applications, MERRA, and the Climate Forecast System Reanalysis, CFSR) with in situ and remotely-sensed observational datasets for temperature, precipitation, and solar radiation, leading to substantial reductions in bias in comparison to a network of 2324 agricultural-region stations from the Hadley Integrated Surface Dataset (HadISD). Results compare favorably against the original reanalyses as well as the leading climate forcing datasets (Princeton, WFD, WFD-EI, and GRASP), and AgMERRA distinguishes itself with substantially improved representation of daily precipitation distributions and extreme events owing to its use of the MERRA-Land dataset. These datasets also peg relative humidity to the maximum temperature time of day, allowing for more accurate representation of the diurnal cycle of near-surface moisture in agricultural models. AgMERRA and AgCFSR enable a number of ongoing investigations in the Agricultural Model Intercomparison and Improvement Project (AgMIP) and related research networks, and may be used to fill gaps in historical observations as well as a basis for the generation of future climate scenarios.

Climate Forcing Data

Astronaut Photography of the Earth: A Long-Term Dataset for Earth Systems Research, Applications, and Education

The NASA Earth observations dataset obtained by humans in orbit using handheld film and digital cameras is freely accessible to the global community through the online searchable database at https://eol.jsc.nasa.gov, and offers a useful compliment to traditional ground-commanded sensor data. The dataset includes imagery from the NASA Mercury (1961) through present-day International Space Station (ISS) programs, and currently totals over 2.6 million individual frames. Geographic coverage of the dataset includes land and oceans areas between approximately 52 degrees North and South latitudes, but is spatially and temporally discontinuous. The photographic dataset includes some significant impediments for immediate research, applied, and educational use: commercial RGB films and camera systems with overlapping bandpasses; use of different focal length lenses, unconstrained look angles, and variable spacecraft altitudes; and no native geolocation information. Such factors led to this dataset being underutilized by the community but recent advances in automated and semi-automated image geolocation, image feature classification, and web-based services are adding new value to the astronaut-acquired imagery. A coupled ground software and on-orbit hardware system for the ISS is in development for planned deployment in mid-2017; this system will capture camera pose information for each astronaut photograph to allow automated, full georegistration of the data. The ground system component of the system is currently in use to fully georeference imagery collected in response to International Disaster Charter activations, and the auto-registration procedures are being applied to the extensive historical database of imagery to add value for research and educational purposes. In parallel, machine learning techniques are being applied to automate feature identification and classification throughout the dataset, in order to build descriptive metadata that will improve search capabilities. It is expected that these value additions will increase interest and use of the dataset by the global community.

Stefanov, William L.

Aircraft Engine Run-to-Failure Dataset Under Real Flight Conditions for Prognostics and Diagnostics

A key enabler of intelligent maintenance systems is the ability to predict the remaining useful lifetime (RUL) of its components, i.e., prognostics. The development of data-driven prognostics models requires datasets with run-to-failure trajectories. However, large representative run-to-failure datasets are often unavailable in real applications because failures are rare in many safety-critical systems. To foster the development of prognostics methods, we develop a new realistic dataset of run-to-failure trajectories for a fleet of aircraft engines under real flight conditions. The dataset was generated with the Commercial Modular Aero-Propulsion System Simulation (CMAPSS) model developed at NASA. The damage propagation modelling used in this dataset builds on the modelling strategy from previous work and incorporates two new levels of fidelity. First, it considers real flight conditions as recorded on board of a commercial jet. Second, it extends the degradation modelling by relating the degradation process to its operation history. This dataset also provides the health respectively fault class. Therefore, besides its applicability to prognostics problems, the dataset can be used for fault diagnostics.

CMAPPS

Visual and Inertial Datasets for an eVTOL Aircraft Approach and Landing Scenario

A National Aeronautics and Space Administration (NASA) project developing computer vision algorithms for autonomous flight is producing real-world datasets with cameras mounted on aircraft. In related domains, such as autonomous driving, open datasets are key to innovation and advancement in computer vision and autonomous perception for future Advanced Air Mobility (AAM) operations. Few vision datasets, however, are publicly available in the aviation context. This paper introduces preliminary datasets containing several examples of approach and landing scenarios. The platform aircraft include a multirotor small unmanned aerial system (sUAS) and a crewed helicopter as surrogates for future electric vertical take-off and landing (eVTOL) aircraft. The dataset provides video imagery with associated inertial navigation system-global positioning system (INS-GPS) position and attitude estimates and other sensors. Surveyed locations of the visual features of the landing area are included. This dataset is the first to be released in an ongoing effort to collect and share large, diverse datasets relevant to autonomous aviation; community critique that can inform and improve future flight campaigns is welcome.

Nelson Brown

Heuristics for Relevancy Ranking of Earth Dataset Search Results

As the Variety of Earth science datasets increases, science researchers find it more challenging to discover and select the datasets that best fit their needs. The most common way of search providers to address this problem is to rank the datasets returned for a query by their likely relevance to the user. Large web page search engines typically use text matching supplemented with reverse link counts, semantic annotations and user intent modeling. However, this produces uneven results when applied to dataset metadata records simply externalized as a web page. Fortunately, data and search provides have decades of experience in serving data user communities, allowing them to form heuristics that leverage the structure in the metadata together with knowledge about the user community. Some of these heuristics include specific ways of matching the user input to the essential measurements in the dataset and determining overlaps of time range and spatial areas. Heuristics based on the novelty of the datasets can prioritize later, better versions of data over similar predecessors. And knowledge of how different user types and communities use data can be brought to bear in cases where characteristics of the user (discipline, expertise) or their intent (applications, research) can be divined. The Earth Observing System Data and Information System has begun implementing some of these heuristics in the relevancy algorithm of its Common Metadata Repository search engine.

science data management

Relevancy Ranking of Satellite Dataset Search Results

As the Variety of Earth science datasets increases, science researchers find it more challenging to discover and select the datasets that best fit their needs. The most common way of search providers to address this problem is to rank the datasets returned for a query by their likely relevance to the user. Large web page search engines typically use text matching supplemented with reverse link counts, semantic annotations and user intent modeling. However, this produces uneven results when applied to dataset metadata records simply externalized as a web page. Fortunately, data and search provides have decades of experience in serving data user communities, allowing them to form heuristics that leverage the structure in the metadata together with knowledge about the user community. Some of these heuristics include specific ways of matching the user input to the essential measurements in the dataset and determining overlaps of time range and spatial areas. Heuristics based on the novelty of the datasets can prioritize later, better versions of data over similar predecessors. And knowledge of how different user types and communities use data can be brought to bear in cases where characteristics of the user (discipline, expertise) or their intent (applications, research) can be divined. The Earth Observing System Data and Information System has begun implementing some of these heuristics in the relevancy algorithm of its Common Metadata Repository search engine.

science data management

Development of a Knowledge Graph for Dataset Discovery and Identification at a NASA Data Center

The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) archives and distributes hundreds of Earth Science data collections to the public. These collections are used in research, resulting in the publication of thousands of scientific papers each year. As new users come to GES DISC for data, it is important for them to understand how prior research used the data. To help researchers, a knowledge graph (KG) was designed and implemented to connect publication citations with dataset metadata. The relationships created in the graph have the potential to allow the Web applications that utilize this information to directly connect the publication to the GES DISC datasets and services. These relationships are demonstrated using a web application prototype. In addition, the graph can also make connections between publications, datasets, and measurements based on the mentions of datasets and their attributes in the publications. To demonstrate this capability, a web application was created that takes the excerpt from the publication and returns a most likely dataset and measurement pairing, ranking the results based on how often these datasets and measurements were used in prior publications.

Nathaniel Crosby

Application of a Dataset-Publication Knowledge Graph for Improving Earth Science Data Search

Finding a dataset at a NASA data center that is the best fit for the researcher’s application presents a challenge, not only for a novice user but for an experienced one, due to the data complexity and a multitude of choices of the existing data. Users often search for the data based on the application they are interested in, their research domain, phenomena, research topic, etc. As existing dataset metadata may not cover these search terms, the user may not obtain the most relevant results for their purpose. This problem was addressed by leveraging the content of the titles and abstracts of the research papers that utilize NASA datasets. For this, features from the paper titles and abstracts were extracted, and then a knowledge graph (KG) was used to link these features to the datasets used in that paper. The search for the datasets was tested by querying this knowledge graph through various terms extracted from Earth Science ontologies such as Semantic Web for Earth and Environment Technology (SWEET), and it was shown that this KG search outperforms the existing search that exclusively queries the dataset metadata.

Kristina Stoyanova

Global Total Ozone Recovery Trends Attributed to Ozone-Depleting Substance (ODS) Changes Derived From Five Merged Ozone Datasets

We report on updated trends using different merged zonal mean total ozone datasets from satellite and ground-based observations for the period from 1979 to 2020. This work is an update of the trends reported in Weber et al. (2018) using the same datasets up to 2016. Merged datasets used in this study include NASA MOD v8.7 and NOAA Cohesive Data (COH) v8.6, both based on data from the series of Solar Backscatter Ultraviolet (SBUV), SBUV-2, and Ozone Mapping and Profiler Suite (OMPS) satellite instruments (1978–present), as well as the Global Ozone Monitoring Experiment (GOME)-type Total Ozone – Essential Climate Variable (GTO-ECV) and GOME-SCIAMACHY-GOME-2 (GSG) merged datasets (both 1995–present), mainly comprising satellite data from GOME, SCIAMACHY, OMI, GOME-2A, GOME-2B, and TROPOMI. The fifth dataset consists of the annual mean zonal mean data from ground-based measurements collected at the World Ozone and Ultraviolet Radiation Data Centre (WOUDC). Trends were determined by applying a multiple linear regression (MLR) to annual mean zonal mean data. The addition of 4 more years consolidated the fact that total ozone is indeed slowly recovering in both hemispheres as a result of phasing out ozone-depleting substances (ODSs) as mandated by the Montreal Protocol. The near-global (60° S–60° N) ODS-related ozone trend of the median of all datasets after 1995 was 0.4 ± 0.2 (2σ) %/decade, which is roughly a third of the decreasing rate of 1.5 ± 0.6 %/decade from 1978 until 1995. The ratio of decline and increase is nearly identical to that of the EESC (equivalent effective stratospheric chlorine or stratospheric halogen) change rates before and after 1995, confirming the success of the Montreal Protocol. The observed total ozone time series are also in very good agreement with the median of 17 chemistry climate models from CCMI-1 (Chemistry-Climate Model Initiative Phase 1) with current ODS and GHG (greenhouse gas) scenarios (REF-C2 scenario). The positive ODS-related trends in the Northern Hemisphere (NH) after 1995 are only obtained with a sufficient number of terms in the MLR accounting properly for dynamical ozone changes (Brewer–Dobson circulation, Arctic Oscillation (AO), and Antarctic Oscillation (AAO)). A standard MLR (limited to solar, Quasi-Biennial Oscillation (QBO), volcanic, and El Niño–Southern Oscillation (ENSO)) leads to zero trends, showing that the small positive ODS-related trends have been balanced by negative trend contributions from atmospheric dynamics, resulting in nearly constant total ozone levels since 2000.

Total Column Ozone Trends

Multi-Decadal Nitrogen Dioxide and Derived Products from Satellites (MINDS) Datasets Released by NASA GES DISC and Their Applications for Air Quality

Nitrogen dioxide (NO2), a pervasive air pollutant, comes from vehicles, power plants, industrial emissions, and off-road sources such as construction or lawn and gardening equipment. The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) curates many remote sensing datasets with NO2 retrievals, which have been utilized for air quality research and applications. The remotely-sensed datasets include those generated by the Ozone Monitoring Instrument (OMI) on the Aura satellite, the TROPOspheric Monitoring Instrument (TROPOMI) onboard the Copernicus Sentinel-5 Precursor (S5P), and the Ozone Mapping and Profiling Suite (OMPS) Nadir-Mapper (NM) instrument on the Suomi National Polar-orbiting Partnership (S- NPP). In collaboration with the NASA Making Earth System Data Records for Use in Research Environments (MEaSUREs) Multi-Decadal Nitrogen Dioxide and Derived Products from Satellites (MINDS) project, the GES DISC recently released MINDS datasets. The NASA MEaSUREs MINDS project aims to develop long-term NO2 global data records by adapting a consistent retrieval algorithm to multiple instrument measurements. Long-term data records will be achieved by applying consistent retrieval approaches to multiple satellite instruments, including OMI (2004 - ); the Global Ozone Monitoring Experiment (GOME, 1995-2011) onboard the second European Remote Sensing satellite (ERS-2); the Scanning Imaging Spectrometer for Atmospheric Cartography (SCIAMACHY, 2002-2012) onboard the ENVIronmental SATellite (ENVISAT); GOME-2 on the Meteorological Operational satellites (MetOp-A and MetOp-B, 2006 - ); and TROPOMI onboard the Copernicus S5P (2017 - ). The long-term record (1995 to present) of MINDS datasets makes them very useful for air quality trend studies. Some MINDS datasets with high spatial resolution of only a few kilometers can be used for air quality research and applications at regional scales. In this presentation, we will introduce all of the MINDS products and services, and demonstrate use cases of MINDS data for studying air quality. We will also present a few other NO2 datasets acquired from NASA’s Health and Air Quality Applied Sciences Team (HAQAST), to be archived and distributed by the GES DISC, and highlight some of their applications for air quality and health.

Feng Ding

Enhancing Dataset Discovery With Knowledge Graph Link Prediction Techniques

● In the evolving landscape of open science, the ability to navigate and discover pertinent datasets is increasingly significant. This primarily hinges on the presence of detailed metadata, delineating the dataset’s content, and potential spheres of application. ● The GES DISC datasets are characterized by science keywords to enable dataset discovery in web search interfaces. ● A problem may arise where a dataset lacks a science keyword that it otherwise should have. ● Machine learning techniques such as link prediction can be used to detect these missing science keywords by estimating the probability of new links forming between dataset and keyword nodes.

machine learning