Data information system at the National Space Science Data Center
Integrated information system to support data handling activities of National Science Data Center
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Integrated information system to support data handling activities of National Science Data Center
Wildfires pose a growing concern in North America due to their harmful impacts on air quality and public health, with increased wildfire activity in recent years leading to widespread smoke plumes that can transcend borders. The exposure of New York City (NYC), the most populous city in North America, to Canadian wildfire smoke highlights the substantial implications for public health and urban environments. To better understand the impact of Canadian wildfires on air quality in NYC, satellite data from the NASA Atmospheric Science Data Center (ASDC) at Langley Research Center, along with ground-based measurements and atmospheric modeling results, are analyzed. We examine concentrations of atmospheric aerosols—particularly PM2.5 particulate matter originating from Canadian wildfires—their dispersion patterns, and the duration and intensity of smoke events impacting NYC. Data from multiple satellites, such as those from the Earth Polychromatic Imaging Camera (EPIC), are synergistically used to identify regions affected by wildfires and estimate aerosol loading. Ground-based measurements, including data from air quality monitoring stations, provide localized information for validation and calibration purposes. The findings of this study contribute to our understanding of the impact of Canadian wildfires on NYC's air quality and emphasize the importance of monitoring and prediction of transboundary smoke events using data synthesized from multiple sources, such as those provided by the ASDC. This information is crucial for policymakers, public health officials, and residents in affected areas to develop effective strategies for mitigating the health risks associated with wildfire smoke and improving air quality during wildfire seasons. The utilization of ASDC data in this research highlights the critical role of atmospheric remote sensing in addressing the challenges posed by wildfires and their consequences on regional scales.
NASA maintains an archive facility for Astronomical Science data collected from NASA's missions at the National Space Science Data Center (NSSDC) at Goddard Space Flight Center. This archive was created to insure the science data collected by NASA would be preserved and useable in the future by the science community. Through 25 years of operation there are many lessons learned, from data collection procedures, archive preservation methods, and distribution to the community. This document presents some of these more important lessons, for example: KISS (Keep It Simple, Stupid) in system development. Also addressed are some of the myths of archiving, such as 'scientists always know everything about everything', or 'it cannot possibly be that hard, after all simple data tech's do it'. There are indeed good reasons that a proper archive capability is needed by the astronomical community, the important question is how to use the existing expertise as well as the new innovative ideas to do the best job archiving this valuable science data.
The sixth annual Space and Earth Science Data Compression Workshop and the third annual Data Compression Industry Workshop were held as a single combined workshop. The workshop was held April 4, 1996 in Snowbird, Utah in conjunction with the 1996 IEEE Data Compression Conference, which was held at the same location March 31 - April 3, 1996. The Space and Earth Science Data Compression sessions seek to explore opportunities for data compression to enhance the collection, analysis, and retrieval of space and earth science data. Of particular interest is data compression research that is integrated into, or has the potential to be integrated into, a particular space or earth science data information system. Preference is given to data compression research that takes into account the scien- tist's data requirements, and the constraints imposed by the data collection, transmission, distribution and archival systems.
Citizen science (or crowdsourcing) has drawn much high-level recent and ongoing interest and support. It is poised to be applied, beyond the by-now fairly familiar use of, e.g., Twitter for natural hazards monitoring, to science research, such as augmenting the validation of NASA earth science mission data. This interest and support is seen in the 2014 National Plan for Civil Earth Observations, the 2015 White House forum on citizen science and crowdsourcing, the ongoing Senate Bill 2013 (Crowdsourcing and Citizen Science Act of 2015), the recent (August 2016) Open Geospatial Consortium (OGC) call for public participation in its newly-established Citizen Science Domain Working Group, and NASA's initiation of a new Citizen Science for Earth Systems Program (along with its first citizen science-focused solicitation for proposals). Over the past several years, we have been exploring the feasibility of extracting from the Twitter data stream useful information for application to NASA precipitation research, with both "passive" and "active" participation by the twitterers. The Twitter database, which recently passed its tenth anniversary, is potentially a rich source of real-time and historical global information for science applications. The time-varying set of "precipitation" tweets can be thought of as an organic network of rain gauges, potentially providing a widespread view of precipitation occurrence. The validation of satellite precipitation estimates is challenging, because many regions lack data or access to data, especially outside of the U.S. and in remote and developing areas. Mining the Twitter stream could augment these validation programs and, potentially, help tune existing algorithms. Our ongoing work, though exploratory, has resulted in key components for processing and managing tweets, including the capabilities to filter the Twitter stream in real time, to extract location information, to filter for exact phrases, and to plot tweet distributions. The key step is to process the "precipitation" tweets to be compatible with satellite-retrieved precipitation data. These key components for processing and managing "precipitation" tweets (and additional ones to be developed) are not limited to precipitation, nor are they limited to the Twitter social medium. Indeed, to maximize the value of our work for NASA earth science programs, these components should be generalized and be part of an overall framework for processing citizen science data for science research. In this paper, we outline such a framework.
The paper describes the European HST Science Data Archive. Particular attention is given to the flow from the HST spacecraft to the Science Data Archive at the Space Telescope European Coordinating Facility (ST-ECF); the archiving system at the ST-ECF, including the hardware and software system structure; the operations at the ST-ECF and differences with the Data Management Facility; and the current developments. A diagram of the logical structure and data flow of the system managing the European HST Science Data Archive is included.
This document provides guidelines for legal, policy, and ethical issues; standards for citizen science data collection and management; information on ensuring usability of citizen science data and communication regarding its use; and best practices for long-term archival of citizen science data.Section1 contains a detailed discussion of policy, ethical, and legal considerations influencing citizen science data collection. Section 2 considers standards for documentation, including documentation of instrumentation, procedures, and the data itself. It concludes with a discussion of how citizen science data should be attributed. Section 3 provides guidance about how to ensure citizen science data are collected and stored in a useable way. It also considers how NASA and data producers should notify the scientific community, including citizen scientists and the public, about citizen science datasets and the scientific conclusions reached using them. Finally, Section 4 provides detailed information regarding what should be archived from projects using a citizen science approach, including data and code. It provides guidance about archive location, process, and timeframe, as well as information about data access and distribution services provided by NASA that may be relevant to data producers working with citizen scientists.
The Landsat Science Data Processing System, developed by NASA for the Landsat 7 Project provides science data handling infrastructure used at the EROS Data Center Landsat 7 Data Handling Facility of the USGS Department of Interior. This paper presents an overview the designs, architectures, and details of the various systems used in the processing of the Landsat 7 Science Data.
The Telemetry and Science Data Software System (TSDSS) was designed to validate the operational health of a spacecraft, ease test verification, assist in debugging system anomalies, and provide trending data and advanced science analysis. In doing so, the system parses, processes, and organizes raw data from the Aquarius instrument both on the ground and while in space. In addition, it provides a user-friendly telemetry viewer, and an instant pushbutton test report generator. Existing ground data systems can parse and provide simple data processing, but have limitations in advanced science analysis and instant report generation. The TSDSS functions as an offline data analysis system during I&T (integration and test) and mission operations phases. After raw data are downloaded from an instrument, TSDSS ingests the data files, parses, converts telemetry to engineering units, and applies advanced algorithms to produce science level 0, 1, and 2 data products. Meanwhile, it automatically schedules upload of the raw data to a remote server and archives all intermediate and final values in a MySQL database in time order. All data saved in the system can be straightforwardly retrieved, exported, and migrated. Using TSDSS s interactive data visualization tool, a user can conveniently choose any combination and mathematical computation of interesting telemetry points from a large range of time periods (life cycle of mission ground data and mission operations testing), and display a graphical and statistical view of the data. With this graphical user interface (GUI), the data queried graphs can be exported and saved in multiple formats. This GUI is especially useful in trending data analysis, debugging anomalies, and advanced data analysis. At the request of the user, mission-specific instrument performance assessment reports can be generated with a simple click of a button on the GUI. From instrument level to observatory level, the TSDSS has been operating supporting functional and performance tests and refining system calibration algorithms and coefficients, in sync with the Aquarius/SAC-D spacecraft. At the time of this reporting, it was prepared and set up to perform anomaly investigation for mission operations preceding the Aquarius/SAC-D spacecraft launch on June 10, 2011.
An Earth Observing System global snow cover extent data products record at moderate spatial resolution (375–500 m) began in February 2000 with the Moderate-resolution Imaging Spectroradiometer (MODIS) instrument onboard the Terra satellite. The record continued with the Aqua MODIS in July 2002, the Suomi-National Polar Platform (S-NPP) Visible Infrared Imaging Radiometer Suite (VIIRS) in January 2012 and continues with the Joint Polar Satellite System-1 (JPSS-1) VIIRS, launched in November of 2017. The objective of this work is to develop a snow cover extent Earth Science Data Record (ESDR) using different satellites, sensors and algorithms. There are many issues to understand when data from different algorithms and sensors are used over a decade-scale time period to create a continuous dataset. Issues may also arise with sensor degradation and even differences in sensor band locations. In this paper we describe development of an ESDR derived from existing MODIS and VIIRS data products and demonstrate continuity among the products. The MODIS and VIIRS snow cover detection algorithms produce very similar daily snow cover maps, with 90–97% agreement in snow cover extent (SCE) in different landscapes. Differences in SCE between products ranged from 2–15% and are attributable to convolved factors of viewing geometry, pixel spread across a scan and time of observation. Compared at a common grid size of 1 km, there is a mean of 95% agreement in SCE and a difference range of 1–10% between the MODIS and VIIRS SCE maps. Mapping sensor observations to a coarser resolution grid reduces the effect of the factors convolved in the 500 m tile to tile comparisons. We conclude that the MODIS and VIIRS SCE data products are reliable constituents of a moderate-resolution ESDR.
Historically, at the end of a NASA mission, earth and space science data were stored at NASA's National Space Science Data Center (NSSDC). The original data archive consisted of both magnetic tapes and film media. As data storage technology improved, data from later missions were stored on disks and platters and higher capacity magnetic media for online accessibility. To conserve physical space at NASA archive sites and to meet disaster recovery guidelines, historical data originally stored on magnetic tapes and film were moved to the Federal Archives and Record Center (FRC) as a temporary holding area until its long-term value was determined by NASA. All records at the FRC are controlled by the NASA Records Retention Schedule (NRRS) which determines the disposal date for each record. On that date, responsible NASA parties are notified that all scheduled records should be reviewed and assessed to determine if they continue to hold significant historical, scientific or administrative value. For Earth Science data records being held at FRC, the Earth Science Data and Information System (ESDIS) Project office is the party responsible for making the value assessment that determines which records warrant preservation and which are ready for proper disposal according to NASA guidelines. Once the data's long-term value is determined, ESDIS takes definitive steps to preserve this data for future discovery and access. Deteriorating media containing historic data of value are recalled from FRC and brought back to ESDIS. Through a tedious, laborious process, digital data are recovered and restored to modern formats with improved metadata and documentation to aid discovery. The restored digital products are then incorporated into our modern online archive, and made immediately accessible to the public. In this paper, we will discuss how we identify data-at-risk, ways to minimize data loss, how we plan for recovery, how we delegate recovery activities to our archive facilities, and how we make recovered data more accessible.
A common science data processing software framework yields the benefits of reuse while remaining adaptable to address requirements that are unique to the mission.
At Goddard, engineers and scientists with a range of experience in science data systems are needed to employ new technologies and develop advances in capabilities for supporting new Earth and Space science research. Engineers with extensive experience in science data, software engineering and computer-information architectures are needed to lead and perform these activities. The increasing types and complexity of instrument data and emerging computer technologies coupled with the current shortage of computer engineers with backgrounds in science has led the need to develop a career path for science data systems engineers and architects.The current career path, in which undergraduate students studying various disciplines such as Computer Engineering or Physical Scientist, generally begins with serving on a development team in any of the disciplines where they can work in depth on existing Goddard data systems or serve with a specific NASA science team. There they begin to understand the data, infuse technologies, and begin to know the architectures of science data systems. From here the typical career involves peermentoring, on-the-job training or graduate level studies in analytics, computational science and applied science and mathematics. At the most senior level, engineers become subject matter experts and system architect experts, leading discipline-specific data centers and large software development projects. They are recognized as a subject matter expert in a science domain, they have project management expertise, lead standards efforts and lead international projects. A long career development remains necessary not only because of the breadth of knowledge required across physical sciences and engineering disciplines, but also because of the diversity of instrument data being developed today both by NASA and international partner agencies and because multidiscipline science and practitioner communities expect to have access to all types of observational data.This paper describes an approach to defining career-path guidance for college-bound high school and undergraduate engineering students, junior and senior engineers from various disciplines.
The FAIR principle (findable, accessible, interoperable, and reusable) governs the storage and sharing of NASA space biology and health data[1]. These guiding principles maximize reuse of data and the reproducibility of scientific findings. The NASA Open Science Data Repository (OSDR; an expansion of NASA GeneLab) was built on the FAIR principles and houses over 500 studies and close to 1000 datasets from decades of space life sciences experiments. OSDR embodies the FAIR principles through data governance that includes mediated, embargoed, and fully open access data. The FAIR data governance principles were recently proposed to be expanded to encompass a FAIREST framework for assessing research data repositories (FAIR + Engagement, Social connections, and Trust)[2]. FAIREST emphasizes the importance of data repositories engaging with the scientific community and gaining the trust of researchers regarding data quality. Trust also refers to the TRUST principles developed for assessment of digital repositories: Transparency, Responsibility, User Focus, Sustainability, Technology[3]. We present the “Open Science for Life in Space” Analysis Working Groups (AWGs) as evidence regarding the power of engagement, social connections, and trust which has enhanced OSDR’s capabilities and productivity. AWG members engage in two main activities. One, members provide feedback on OSDR scientific standards for data ingestion, curation, and reuse (study, subject and assay metadata; processing pipelines; dataset formats and uniformed structures for machine-readability). Two, AWG members collaborate to mine-reuse OSDR data to conduct scientific analysis. With nearly 800 active members, the AWGs have resulted in 32 publications re-using OSDR data and contributed many papers in two major special issues in Cell (2020) and Nature (2024). AWGs also serve as networking groups, facilitate social connections between researchers at all levels of experience, and also have a social online ‘Forum’ used to keep members informed on projects and opportunities. This community-centric, productive, and trustworthy data culture has resulted in a broader effect with international space agencies, academics, and the commercial space sector wanting to submit their data to OSDR. Ten studies of Inspiration 4 data were recently publicly released by OSDR, as were some JAXA human data. Coming up soon in OSDR are data submissions from the European Space Agency, Virgin Galactic PIs, and SpaceX Polaris Dawn. A major benefit of OSDR is the array of standardized and uniformly formatted data (which was developed through AWG member consensus), from which visualization tools, analysis tools, and machine learning models can be built or trained. This talk will cover the Multi-Study Visualization Tool, the Environmental Data Application, RadLab, and a UCSF-NSF funded knowledge graph biomedical health discovery tool ‘SPOKE’ currently being integrated with OSDR. OSDR also provides training programs in bioinformatics and machine learning to improve the scientific community’s awareness of data availability and to boost their ability to perform data analysis. The increasing engagement of the scientific community and the public with technologies powered by artificial intelligence (AI) heightens the need for data analysis to be transparent. The AI for Life in Space initiative leverages the data products provided in OSDR to train AI models, with an emphasis on explainable and trustworthy AI, which would not be possible without FAIR data and metadata. Overall, here we will demonstrate the importance for NASA life sciences data repositories to adhere to the FAIREST framework, by providing examples and success stories from different aspects of OSDR.
The increase in the number and volume, and sources, of globally available Earth science data measurements and datasets have afforded Earth scientists and applications researchers unprecedented opportunities to study our Earth in ever more sophisticated ways. In fact, the NASA Earth Observing System Data Information System (EOSDIS) archives have doubled from 2007 to 2014, to 9.1 PB (Ramapriyan, 2009; and https:earthdata.nasa.govaboutsystem-- performance). In addition, other US agency, international programs, field experiments, ground stations, and citizen scientists provide a plethora of additional sources for studying Earth. Co--analyzing huge amounts of heterogeneous data to glean out unobvious information is a daunting task. Earth science data analytics (ESDA) is the process of examining large amounts of data of a variety of types to uncover hidden patterns, unknown correlations and other useful information. It can include Data Preparation, Data Reduction, and Data Analysis. Through work associated with the Earth Science Information Partners (ESIP) Federation, a collection of Earth science data analytics use cases have been collected and analyzed for the purpose of extracting the types of Earth science data analytics employed, and requirements for data analytics tools and techniques yet to be implemented, based on use case needs. ESIP generated use case template, ESDA use cases, use case types, and preliminary use case analysis (this is a work in progress) will be presented.
Earth Science Data and Information Systems Overview at the Open Source Science for the Earth System Observatory Science Data Processing System Workshop.
Recognizing the significance of NASA remote sensing Earth science data in monitoring and better understanding our planet s natural environment, NASA has implemented the Decision Support Through Earth Science Research Results program (NASA ROSES solicitations). a) This successful program has yielded several monitoring, surveillance, and decision support systems through collaborations with benefiting organizations. b) The Goddard Space Flight Center (GSFC) Earth Sciences Data and Information Services Center (GES DISC) has participated in this program on two projects (one complete, one ongoing), and has had opportune ad hoc collaborations gaining much experience in the formulation, management, development, and implementation of decision support systems utilizing NASA Earth science data. c) In addition, GES DISC s understanding of Earth science missions and resulting data and information, including data structures, data usability and interpretation, data interoperability, and information management systems, enables the GES DISC to identify challenges that come with bringing science data to decision makers. d) The purpose of this presentation is to share GES DISC decision support system project experiences in regards to system sustainability, required data quality (versus timeliness), data provider understanding of how decisions are made, and the data receivers willingness to use new types of information to make decisions, as well as other topics. In addition, defining metrics that really evaluate success will be exemplified.