Search NASA⌕ Search

Engineering topics

Gao Chen

Publications and source records attributed to Gao Chen.

54 records · Page 3

WIS and WIGOS Metadata as the Foundation for a Sustainable Framework for Global Greenhouse Gas Watch Data Exchange

Metadata (data about data) is a critical component of data discovery, description, evaluation, documentation, and preservation. Developing and propagating metadata standards has been a longstanding area of activity in WMO and beyond. The WIS2 and WIGOS metadata models are being actively developed and maintained by dedicated task teams, established under the WMO Expert Team on Metadata. The metadata representations and vocabularies are governed by well-established processes within WMO. These standards are being used in a number of metadata/data exchange activities (e.g., WMO Information System 2.0 (WIS2), WIGOS (WMDR), Climate Data Management Systems (CMDS), etc.). It should also be noted that the application of the WIS2 and WIGOS standards fully support the WMO Unified Data Policy and open data policy as well as greatly enhance the value of observations by fostering data F.A.I.R.ness. Furthermore, the WMO metadata standards can serve as the foundation for a framework that will facilitate metadata mapping between the existing schemas used in well-established data centres, e.g., WMO WDCGG (World Data Centre for Greenhouse Gases) and NOAA ObsPack (Observation Package Data Products) and to automate metadata exchange between data centres as well as with WMO. These activities will play a central role in integrating measurements sponsored by various member countries and organizations to provide a more comprehensive characterization of the temporal and spatial distribution of the greenhouse gases. At the same time, this metadata exchange can lead to member countries and partner organizations improving their current metadata collection process for data discoverability, interoperability, and (re)usability. This presentation will describe metadata activities in the context of WIS2 and WIGOS and how they apply to GGGW data integration via metadata mapping and exchange.

Gao Chen↗

The TOLNet 2.0 Website: How an API Can Promote Open Science and FAIR Principles

The Tropospheric Ozone Lidar Network (TOLNet) has generated over a decade of ozone vertical profile data products over North America. The science value of the TOLNet data has been demonstrated in numerous peer-reviewed publications on air quality and other ozone relevant research. To support the broad spectrum of data use, the TOLNet team launched a major effort to upgrade the web-based data repository aiming to enhance the data discoverability and to enable machine-to-machine data upload and download processes. Specifically, the TOLNet website included an application programming interface (API), which supports machine-to-machine data search and data download. The API also extracts selected variables from the files, which can be retrieved as JSON objects and used to create data displays without having to download or open the underlying files. The TOLNet science team members can also use the API for automated data upload, including a data file scanning feature to ensure data product integrity. To be presented will include a summary of key features of data repositories, an actual use case of machine-to-machine data access/use, as well as our journey to make TOLNet data more FAIR, i.e., more findable, accessible, interoperable, and (re)usable.

Crystal Gummo↗

Blast from the Past: ASDC Curation for NASA Suborbital Legacy Missions to Promote Data Discovery and Accessibility

NASA has an extensive history of conducting suborbital field campaigns to further advances in atmospheric sciences. Beginning with the Chemical Instrument Test and Evaluation (CITE) conducted in 1983-1984, NASA has completed many suborbital campaigns over the past three decades. Since the early 2010s, suborbital missions are typically assigned to a NASA Distributed Active Archive Center (DAAC) prior to the mission for long-term archival and distribution. Efforts are being made by NASA’s Earth Science Data and Information System (ESDIS) Project and the Airborne Data Management Group (ADMG) to assign legacy missions to DAACs for permanent archival and distribution, so that these valuable datasets remain to be available to the scientific community. NASA’s Atmospheric Science Data Center (ASDC) has been named the assigned DAAC for nearly 20 atmospheric composition legacy missions, including missions conducted as part of the Global Tropospheric Experiment (GTE) and expects to be named the assigned DAAC for more of these missions over the next few years. The primary goal of the ASDC is to provide access to the datasets as they are currently formatted to the broad user community and enhance their findability and accessibility. However, data reporting standards have evolved significantly since 1983 and the datasets span a wide variety of file formats, including text, Ames, GTE, and ICARTT (International Consortium for Atmospheric Research on Transport and Transformation), and the amount of metadata and relevant information included in the files also varies greatly and can not be readily extracted without subject matter knowledge. This has caused challenges for the ASDC’s suborbital metadata extraction pipeline in ensuring that accurate and necessary metadata is being provided for the missions by all the ASDC’s existing search mechanisms. To make the data more findable and accessible, the ASDC has begun researching ways to further enhance the datasets, including distributing value-added products (i.e. consistent file format such as ICARTT or netCDF), adding standard names from the ESDIS Standards Coordination Office (ESCO)-approved Atmospheric Composition Variable Standard Names Convention (ACVSNC), and creating outreach materials such as ArcGIS StoryMaps, User Guides, and Micro Articles, providing overviews of the missions and what type of data was collected during the missions. These efforts also help support NASA’s Open-Source Science by enhancing the FAIRness of the legacy data products. This presentation will review the ASDC’s ongoing efforts, progress made, and future plans for legacy missions.

Megan Buzanowicz↗

Extending CF Conventions to Enhance Data FAIRness for Atmospheric Composition Observations

The Hierarchical Data Format (HDF) and Network Common Data Form (NetCDF) are data file formats created to aid users in the creation or use of scientific data. These file formats are useful for handling large data volumes and hosting extensive metadata as global, group, or variable attributes and are popular with the modeling community. HDF and NetCDF files are widely used with atmospheric remote sensing data and have been used to support measurements from numerous field campaigns, from satellite to aircraft or ground and mobile based measurements. The files from airborne field studies, however, vary greatly in terms of the file structure and the amount and content of their metadata. Information relevant to the file that can be useful to the user such as the data producer, location where data was taken, variable descriptions, or information about the instrument might not be included in the file. Recently, the Measurements of Aerosols, Clouds, and their Interactions for Earth System Models (MACIE) group started a grassroots effort to develop a CF-based template for the HDF and NetCDF files for field studies, with the aim of making the data products more interoperable and usable. This template seeks to make the files more compliant to Climate and Forecast (CF) metadata conventions and to standardize the file structure and the global and variable attributes. The template would help to ensure that HDF and NetCDF files contain adequate metadata to better support their use for research, e.g., the modeling community, and to enhance the usability and interoperability of data for research communities at large. The draft template has been applied to recent field studies for various instruments and their merge files in support of the Atmosphere Observing System (AOS) project. The details of the revised template are to be presented, as well as examples of the implementation of these requirements for merge files and lidar observation data files and issues revealed during the implementation process.

Sean Leavor↗

ASDC’s Python-Based Metadata Extraction Pipeline for Suborbital Campaigns

The FAIRness of data products, especially findability and accessibility depend on rich metadata which, when extracted, can allow for proper curation. Over the past few years, the Atmospheric Science Data Center (ASDC) suborbital science support team has developed a metadata extraction pipeline to ensure the required metadata can be retrieved systematically, effectively, and efficiently to ensure the data can be used by a broad community. The development of a pipeline has presented many, but necessary, challenges to support archival and distribution of ASDC’s 30+ suborbital missions. Though sufficient metadata is provided by instrument scientists, the metadata may not be readily machine actionable due to different formats and templates. Further complicating metadata extraction, our team has found that the nature of metadata can be quite diverse given the difference in measurement types, instruments, and measurement platforms. A metadata extraction pipeline has been developed to provide an efficient, plugin-in based, method for adding new parsers, a configuration system that lets non-developers customize how files are processed, and a system for identifying and logging metadata quality issues to ensure they are readily found and addressed. The metadata extraction pipeline identifies critical pieces of metadata that are needed to promote data FAIRness, including location, file revision, measurement start/end datetime and can be easily modified to extract further information (such as variables). Given the wide-ranging datasets, the pipeline has been modified to accommodate multiple file formats, including multiple versions of ICARTT (International Consortium for Atmospheric Research on Transport and Transformation), HDF (Hierarchical Data Format), netCDF (network Common Data Form), and multiple versions of the Ames File Format. The pipeline also supports building metadata for file formats that cannot have metadata easily extracted from them, such as PDF (Portable Document Format) and GIF (Graphics Interchange Format). The pipeline has allowed our team to maintain a consistent flow of data and metadata to archival and distribution services, ensuring the ASDC meets the needs of the suborbital science community. This presentation will highlight the ASDC’s suborbital metadata extraction pipeline, its development, how it’s been modified to support data FAIRness, and plans for maintaining the pipeline and adding new features.

Abraham Porter↗

An Overview of NASA’s Airborne and Field Data Resource Center

A key recommendation from NASA’s 2022 Airborne and Field Data Workshop called for the development of a virtual Resource Center for all stakeholders across the data lifecycle of airborne and field Earth observations. The agency’s Earth Science Data and Information System (ESDIS) Project and Airborne Data Management Group (ADMG) have worked in concert to establish the newly launched NASA Airborne and Field Data Resource Center (AFDRC) to provide a single entry point for a wide assortment of information on and the effective, responsible stewardship of non-satellite observational data. The AFDRC compiles access to many existing resources, but does so in a newly organized way that integrates availability to increase efficiency and holistic understanding while simplifying users’ experience. Initially launched in fall of 2023, the AFDRC is a NASA Earthdata domain website that clarifies several previously disparate resources and provides newly updated information, including: Learning Resources: Educational resources to broaden understanding of the role airborne and field observations play in advancing understanding of our planet and NASA’s role in collecting and archiving these data. Support for data users to Find and Access Data: Advanced contextual browse/search capabilities that efficiently link researchers to data products suitable for their science objectives - this includes linking to NASA’s Catalog of Archived Suborbital Earth Science Investigations (CASEI). Working with Data: Tools specific to individual types of suborbital Earth Science data and their (inter-)disciplinary communities to provide access as well as guidance for their application. Stewardship Responsibilities: Resources for data producers with to lessen requirement burdens at the time of data transfer, and information for data stewards with consistent, authoritative guidance on best practices and agency- and/or community- specific archival procedures. This presentation will give an overview of the motivation for NASA’s AFDRC, approach for the design and content, iterative community-driven improvements, promote the use of the AFDRC, and solicit additional feedback from airborne and field data user communities.

Sara Lubkin↗

Effects of Natural Variability on the Use of Standard Deviation to Represent Measurement Uncertainties in Atmospheric Composition Studies

Measurement uncertainty is defined as a “non-negative parameter characterizing the dispersion of the quantity values being attributed to a measurand”. It is most common that the uncertainties of GAW hourly measurements, such as greenhouse gas measurements, are reported as standard deviations derived from individual sampling at a higher time resolution (e.g., 1 min). In contrast, the uncertainties of GAW measurements of reactive gases and aerosol properties are reported in percentiles covering the same probability. A quick look at hourly CO 2 data from Cape Grim, Australia yielded some interesting findings: the hourly standard deviation is, on average, more than a factor of 10 higher for the measurements under non-background conditions (over 50% observations), while the difference in average CO 2 amount fraction was less than 2 ppm. The dramatic contrast cannot be explained by the difference in measurement uncertainties, but can largely be attributed the natural variability, or episodic ambient CO 2 variation reflecting changes in meteorological conditions or emissions. These initial findings motivated a more in-depth analysis of the ground-based measurements of trace gases and aerosol properties. This analysis will be using continuous 1 min ground site observations to construct time averaged statistical indicators to evaluate whether the standard deviation is adequate to represent the dispersion, especially under marked influence by natural variability. The suitability of this representation can be determined by examining the difference between the standard deviation and percentiles encompassing the same probability. We will examine time intervals of 1 hour, 3 hours, and 24 hours, with the latter time intervals chosen to match those commonly used in model assessments. We will also investigate how natural variability can alter the probability distribution of the measurands and how adequate the quadrature propagation of uncertainties is under these conditions. The results will include several trace gases (e.g., CO 2 , CO, O 3 , and NO 2 ) with a range of measurement techniques (e.g., PANDORA, in situ), atmospheric lifetimes, and aerosol properties (e.g. scattering coefficient). The findings from this analysis should provide some useful feedback on the best practices for uncertainty reporting.

Measurement Uncertainty↗

TOLNet’s FAIR Journey: Yesterday, Today, and Tomorrow

The Tropospheric Ozone Lidar Network (TOLNet) has generated over a decade of ozone vertical profile data products over North America and contributed to several air quality focused field studies. The science value of the TOLNet data has been demonstrated in numerous peer-reviewed publications on air quality and ozone relevant research. As the broad scientific community has moved towards adopting FAIR Principles to make data more findable, accessible, interoperable, and (re)usable, the TOLNet team has been consistently making data more FAIR. This effort has many challenges, partially reflecting on the FAIR principles being domain agnostic while the implementation needs to be domain specific. The FAIR principles declare the dependence on the community standards, domain-relevant metadata, and rich metadata. This presentation uses the TOLNet data and data system as an example to explore the best practices to implement FAIR principle. Particularly, we will examine the metadata and the “richness” to support findability and usability as well as machine-to-machine actionability via API. Last year, as part of our FAIR journey, we launched the TOLNet website (https://tolnet.larc.nasa.gov/) and the API (https://tolnet.larc.nasa.gov/api/). Part of this process included extracting and cataloging metadata across the entire TOLNet mission timeframe. This enabled users to search through the mission by various metadata criteria, improving the findability and accessibility. And computers could connect directly to the TOLNet API to extract both metadata and data, providing a level of interoperability never present before for TOLNet data. On top of that, all new TOLNet data is now automatically validated using the API to ensure it complies with GEOMS standards, aiding in reusability. It takes both technology and scientists working together to make progress. The next step is to evaluate the current TOLNet offerings against NASA’s Practical Guide for Open, Free & FAIR NASA Earth Science Data Products (https://doi.org/10.5067/DOC/ESCO/ESDSWG-0002V1).

TOLNet↗