Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36

Quantitative Highlights of 20 years Aqua Data Archive and Data Usage

NASA’s Aqua satellite carries six Earth-observing instruments Atmospheric Infrared Sounder (AIRS), Advanced Microwave Scanning Radiometer for EOS (AMSR-E), Advanced Microwave Sounding Unit (AMSU), Clouds and the Earth’s Radiant Energy System (CERES), Humidity Sounder for Brazil (HSB) and Moderate Resolution Imaging Spectroradiometer (MODIS). Currently only four of six instruments are collecting data, two instruments that stopped transmitting data are AMSR-E that suffered a major anomaly in October 2011 and was powered off in March 2016 while as HSB failed in February 2003. NASA’s Earth Science Data and Information System (ESDIS) Project makes these data, along with derived products, available to worldwide data users. Since the launch of Aqua on May 4, 2002, more than 10,000 data products have been archived and distributed by NASA-funded Distributed Active Archive Centers (DAACs) that are part of NASA’s Earth Observing System Data and Information System (EOSDIS). At the end of the 2021 Fiscal Year with over 100,000 orbits data, about 1,000 Aqua data products constituted almost 16.5 % of the entire EOSDIS data archive volume (8.6 PB out of approximately 55.2 PB), and 7.5 PB of Aqua data were distributed to over half-a-million public users worldwide. By categorizing the Aqua data products and their distribution, we can get a quantitative assessment of Aqua data usage. NASA’s ESDIS Project has collected archive, distribution, and user information from EOSDIS data users since February 2000. These metrics are available through the ESDIS Metrics System (EMS). EMS information is stored in a relational database from which quantitative metrics of Aqua data use can be retrieved and analyzed. The purposes of this study are to: 1) perform a comprehensive investigation of the 20-year trend in the archive and distribution of Aqua data products; 2) identify and characterize data product usage over the last 20 years; and 3) identify and characterize the global user community for these data. In addition to revealing how Aqua data use has evolved over time, the results of this study provide insights on identifying the various user communities for different kinds of Earth science data products. Also, because of the enormous quantity of data handled by EOSDIS DAACs, the study provides guidance of the requirements for future data systems that will be needed to effectively and efficiently handle the ever-increasing amounts of Earth science data produced by future (and ongoing) Earth science missions.

Lalit Wanchoo↗

The Use of HDTV Format and the Electronic Theater in Presenting Earth Science

In order to maximize the public's awareness of earth science observations, earth science data must be available in multiple media formats. This talk will focus on the use High Definition TV format in presenting earth science data, The Television (HDTV) networks are mandated to completely switch over from the current TV standard (NTSC) to HDTV in the next seven years. Museums are also beginning to use HDTV format in their displays. The Visualization Analysis Laboratory at Goddard Space Flight Center has been experimenting with the use of HDTV to present earth science data. The experimental package we have developed is called the Electronic Theater (e-theater). The e-theater is a mobile presentation system used for displaying and teaching groups about earth science and the delicate interdependence between the various earth systems. The e-theater takes advantage of a double-wide screen to show the audiences high resolution data displays. The unique architecture used in this exhibit allows several data sets to be displayed at one time, demonstrating the connections between different earth systems. The data animations are manipulated in real-time during the presentation and can be paused, moved forward, backward, looped, or zoomed into, to maximize the flexibility of the presentation. Because HDTV format is used within the e-theater, the materials generated for the e-theater are made available to the news media and museums.

Summey, Barbara↗

Update on Apollo Data Restoration by the NSSDC and the PDS Lunar Data Node

The Lunar Data Node (LDN) , under the auspices of the Geosciences Node of the Planetary Data System (PDS) and the National Space Science Data Center (NSSDC), is continuing its efforts to recover and restore Apollo science data. The data being restored are in large part archived with NSSDC on older media, but unarchived data are also being recovered from other sources. They are typically on 7- or 9-track magnetic tapes, often in obsolete formats, or held on microfilm, microfiche, or paper documents. The goal of the LDN is to restore these data from their current form, which is difficult for most researchers to access, into common digital formats with all necessary supporting data (metadata) and archive the data sets with PDS. Restoration involves reading the data from the original media, deciphering the data formats to produce readable digital data and converting the data into usable tabular formats. Each set of values in the table must then be understood in terms of the quantity measured and the units used. Information on instrument properties, operational history, and calibrations is gathered and added to the data set, along with pertinent references, contacts, and other ancillary documentation. The data set then undergoes a peer review and the final validated product is archived with PDS. Although much of this effort has concentrated on data archived at NSSDC in the 1970's, we have also recovered data and information that were never sent to NSSDC. These data, retrieved from various outside sources, include raw and reduced Gamma-Ray Spectrometer data from Apollos 15 and 16, information on the Apollo 17 Lunar Ejecta And Meteorites experiment, Dust Detector data from Apollos 11, 12, 14, and I5, raw telemetry tapes from the Apollo ALSEPs, and Weekly Status Reports for all the Apollo missions. These data are currently being read or organized, and supporting data is being gathered. We are still looking for the calibrated heat flow data from Apollos 15 and 17 for the period 1975-1977, any assistance or information on these data would be welcome. NSSDC has recently been tasked to release its hard-copy archive, comprising photography, microfilm, and microfiche. The details are still being discussed, but we are concentrating on recovering the valuable lunar data from these materials while they are still readily accessible. We have identified the most critical of these data and written a LASER proposal to fund their restoration. Included in this effort are data from the Apollo 15 and 16 Mass Spectrometers and the Apollo 17 Par-UV Spectrometer and ancillary information on the Apollo 17 Surface Electrical Properties Experiment.

Williams, David R.↗

Exploring New Methods of Displaying Bit-Level Quality and Other Flags for MODIS Data

The NASA Distributed Active Archive Center (DAAC) at the National Snow and Ice Data Center (NSIDC) archives and distributes snow and sea ice products derived from the MODerate resolution Imaging Spectroradiometer (MODIS) on board NASA's Terra and Aqua satellites. All MODIS standard products are in the Earth Observing System version of the Hierarchal Data Format (HDF-EOS). The MODIS science team has packed a wealth of information into each HDF-EOS file. In addition to the science data arrays containing the geophysical product, there are often pixel-level Quality Assurance arrays which are important for understanding and interpreting the science data. Currently, researchers are limited in their ability to access and decode information stored as individual bits in many of the MODIS science products. Commercial and public domain utilities give users access, in varying degrees, to the elements inside MODIS HDF-EOS files. However, when attempting to visualize the data, users are confronted with the fact that many of the elements actually represent eight different 1-bit arrays packed into a single byte array. This project addressed the need for researchers to access bit-level information inside MODIS data files. In an previous NASA-funded project (ESDIS Prototype ID 50.0) we developed a visualization tool tailored to polar gridded HDF-EOS data set. This tool,called the Polar researchers to access, geolocate, visualize, and subset data that originate from different sources and have different spatial resolutions but which are placed on a common polar grid. The bit-level visualization function developed under this project was added to PHDIS, resulting in a versatile tool that serves a variety of needs. We call this the EOS Imaging Tool.

Khalsa, Siri Jodha Singh↗

Analysis and Review of NASA Earth Science Metadata: How Automation Plays a Role

The Analysis and Review of the Common Metadata Repository (CMR ARC) Team reviews all EOSDIS metadata. The team’s objective is to achieve consistency, correctness, and completeness for all metadata records in the CMR, as well as improve the discoverability of NASA's Earth Science data within the CMR framework. This work is currently being completed at Marshall Space Flight Center. CMR makes a single discovery point possible for NASA's Earth Science data users. The CMR team, in collaboration with three other core metadata teams, contributes to the stewardship of NASA's Earth Science data through a process of continual curation and the ongoing development of the Unified Metadata Model (UMM). A key tool now used in the curation process, referred to as the NASA CMR Dashboard, is an online curation dashboard developed in collaboration with software development company, Element 84. This tool facilitates the review of Earth Science metadata records and subsequent stakeholder collaboration on the resolution of identified issues. A key capability of the new tool is a suite of automated compliance checks written in Python 3.6 that verify the integrity of various metadata elements across multiple standards.

Staton, Patrick↗

The moderate resolution imaging spectrometer (MODIS) science and data system requirements

The Moderate Resolution Imaging Spectrometer (MODIS) has been designated as a facility instrument on the first NASA polar orbiting platform as part of the Earth Observing System (EOS) and is scheduled for launch in the late 1990s. The near-global daily coverage of MODIS, combined with its continuous operation, broad spectral coverage, and relatively high spatial resolution, makes it central to the objectives of EOS. The development, implementation, production, and validation of the core MODIS data products define a set of functional, performance, and operational requirements on the data system that operate between the sensor measurements and the data products supplied to the user community. The science requirements guiding the processing of MODIS data are reviewed, and the aspects of an operations concept for the production of data products from MODIS for use by the scientific community are discussed.

Ardanuy, Philip E.↗

Aurorasaurus Database of Real-Time, Crowd-Sourced Aurora Data for Space Weather Research

This technical report documents the details of Aurorasaurus citizen science data for the period spanning 2015 and 2016 as well as its routine data filtering protocols. Aurorasaurus citizen science data is a collection of auroral sightings submitted to the project via its website or apps and mined from social media. It is a robust data set and particularly abundant during strong geomagnetic storms when auroral precipitation models have the highest uncertainty. These data are offered to the scientific community for use through an openaccess database in its raw and scientific formats, each of which is described in detail in this technical report. Furthermore, by demonstrating its scientific utility, we aim to encourage its integration into auroral research.

Citizen science↗

Enabling Earth Science Through Cloud Computing

Cloud Computing holds tremendous potential for missions across the National Aeronautics and Space Administration. Several flight missions are already benefiting from an investment in cloud computing for mission critical pipelines and services through faster processing time, higher availability, and drastically lower costs available on cloud systems. However, these processes do not currently extend to general scientific algorithms relevant to earth science missions. The members of the Airborne Cloud Computing Environment task at the Jet Propulsion Laboratory have worked closely with the Carbon in Arctic Reservoirs Vulnerability Experiment (CARVE) mission to integrate cloud computing into their science data processing pipeline. This paper details the efforts involved in deploying a science data system for the CARVE mission, evaluating and integrating cloud computing solutions with the system and porting their science algorithms for execution in a cloud environment.

science data system↗

Radioastron flight operations

Radioastron is a space-based very-long-baseline interferometry (VLBI) mission to be operational in the mid-90's. The spacecraft and space radio telescope (SRT) will be designed, manufactured, and launched by the Russians. The United States is constructing a DSN subnet to be used in conjunction with a Russian subnet for Radioastron SRT science data acquisition, phase link, and spacecraft and science payload health monitoring. Command and control will be performed from a Russian tracking facility. In addition to the flight element, the network of ground radio telescopes which will be performing co-observations with the space telescope are essential to the mission. Observatories in 39 locations around the world are expected to participate in the mission. Some aspects of the mission that have helped shaped the flight operations concept are: separate radio channels will be provided for spacecraft operations and for phase link and science data acquisition; 80-90 percent of the spacecraft operational time will be spent in an autonomous mode; and, mission scheduling must take into account not only spacecraft and science payload constraints, but tracking station and ground observatory availability as well. This paper will describe the flight operations system design for translating the Radioastron science program into spacecraft executed events. Planning for in-orbit checkout and contingency response will also be discussed.

Altunin, V. I.↗

Data Preservation, Information Preservation, and Lifecyle of Information Management at NASA GES DISC

Data lifecycle management awareness is common today; planners are more likely to consider lifecycle issues at mission start. NASA remote sensing missions are typically subject to life cycle management plans of the Distributed Active Archive Center (DAAC), and NASA invests in these national centers for the long-term safeguarding and benefit of future generations. As stewards of older missions, it is incumbent upon us to ensure that a comprehensive enough set of information is being preserved to prevent the risk for information loss. This risk is greater when the original data experts have moved on or are no longer available. Preservation of items like documentation related to processing algorithms, pre-flight calibration data, or input-output configuration parameters used in product generation, are examples of digital artifacts that are sometimes not fully preserved. This is the grey area of information preservation; the importance of these items is not always clear and requires careful consideration. Missing important metadata about intermediate steps used to derive a product could lead to serious challenges in the reproducibility of results or conclusions. Organizations are rapidly recognizing that the focus of life-cycle preservation needs to be enlarged from the strict raw data to the more encompassing arena of information lifecycle management. By understanding what constitutes information, and the complexities involved, we are better equipped to deliver longer lasting value about the original data and derived knowledge (information) from them. The NASA Earth Science Data Preservation Content Specification is an attempt to define the content necessary for long-term preservation. It requires new lifecycle infrastructure approach along with content repositories to accommodate artifacts other than just raw data. The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) setup an open-source Preservation System capable of long-term archive of digital content to augment its raw data holding. This repository is being used for such missions as HIRDLS, UARS, TOMS, OMI, among others. We will provide a status of this implementation; report on challenges, lessons learned, and detail our plans for future evolution to include other missions and services.

data management↗

New developments in space radiation research at NASA: Annotating data using a novel radiation biology ontology

Like many interdisciplinary sciences, data producers and consumers in the field of radiation biology often use a wide variety of terminology to describe their experiments and data. Furthermore, space systems and technologies are rapidly evolving, and a shared understanding and common terminology for these is also lacking. The efficiency of research organizations can be enhanced by standardizing metadata through the use of knowledge resources like ontologies. Employing a sophisticated model such as a formal ontology to standardize metadata enables automated data acquisition processes and supports more complete, accurate meta-analysis through more efficient and complete data discovery and retrieval, particularly when using multiple data sources. Thus, we developed the Radiation Biology Ontology (RBO) in order to improved radiation biology metadata uniformity and transparency. We used open-source software (the Ontology Development Kit, Protégé and WebProtégé) and worked within the OBO Foundry framework, which includes a set of ontology development principles and practices for ontology consistency, uniformity, and accountability. The RBO has now been incorporated into two radiation research data repositories, NASA’s GeneLab omics database (https://genelab.nasa.gov), and the European Commission STORE database (https://www.storedb.org/). Continuous build integration tools allowed our international RBO collaboration to be more efficient and focus its efforts on semantic model design. Currently, the RBO contains over 300 annotated classes and individuals specific to the study of radiation on biological systems, as well as imports of many additional classes from other OBO Foundry ontologies that relate to and/or provide context for these RBO entities. We publish the RBO through the OBO Foundry, so that it is available for browsing, download, and querying through NCBI Bioportal web site and application programming interface. The NASA Ames Life Science Data Archive (ALSDA) is also in the process of adopting use of the RBO, taking NASA one step closer to a knowledge-based system for space biology data. It is our hope that the global communities of radiation research Investigators, data curators and data analysts can similarly leverage the RBO and will contribute to its further development.

radiation↗

Elevating the Quality of Space Omics Sequencing Data: Innovations and Methodologies from NASA GeneLab Sample Processing Laboratory

NASA’s GeneLab, part of the NASA Open Science Data Repository, is a space-related database that hosts a diverse range of transcriptomics, proteomics, epigenomics and genomics data. The NASA GeneLab Sample Processing Laboratory (SPL) generates omics data from biological experiments conducted aboard the International Space Station, Space Shuttle and space related ground experiments, this omics data then hosted on the GeneLab repository. Samples generated such experiments pose numerous technical challenges such as small experimental sample size, variance in dissection times, limited tissue preservation methods, prolonged storage time, and more. GeneLab SPL team had developed specialized expertise in nucleic acid extraction, library preparation and sequencing of such biological samples via extensive training and years of experience. In order to ensure data accuracy and consistency across experiments, SPL has developed standardized protocols for each species and tissue type. These protocols in conjunction with quality control metrics and data standards are crucial in generating of high-quality data. SPL protocols and standards have been developed in collaboration with the scientific community and had been made publicly available on the GeneLab portal, guaranteeing comparability of datasets across spaceflight experiments. To ensure reliability of data generation, SPL leverages cutting-edge innovations in laboratory automation for sample processing. By leveraging these state-of-the-art platforms, SPL achieves high levels of data reproducibility while significantly minimizing sources of bias and variability, especially across experiments with large numbers of samples. Over the past few years, the space biology investigator community has accessed SPL-generated data from the Open Science Data Repository for a myriad of data re-analysis and re-use studies. We observe a trend that in-house SPL-generated data consistently outperforms outsourced sequencing data in terms of technical standards, quality control metrics, timeliness of data delivery, and sequencing and reagent efficiency. Superior data generation has and will continue to enable discoveries in disease, diagnostic tools, and the biological effects of long duration spaceflight.

GeneLab↗

Elevating the Quality of Space Omics Sequencing Data: Innovations and Methodologies from NASA GeneLab Sample Processing Laboratory

NASA’s GeneLab, part of the NASA Open Science Data Repository, is a space-related database that hosts a diverse range of transcriptomics, proteomics, epigenomics and genomics data. The NASA GeneLab Sample Processing Laboratory (SPL) generates omics data from biological experiments conducted aboard the International Space Station, Space Shuttle and space related ground experiments, this omics data then hosted on the GeneLab repository. Samples generated such experiments pose numerous technical challenges such as small experimental sample size, variance in dissection times, limited tissue preservation methods, prolonged storage time, and more. GeneLab SPL team had developed specialized expertise in nucleic acid extraction, library preparation and sequencing of such biological samples via extensive training and years of experience. In order to ensure data accuracy and consistency across experiments, SPL has developed standardized protocols for each species and tissue type. These protocols in conjunction with quality control metrics and data standards are crucial in generating of high-quality data. SPL protocols and standards have been developed in collaboration with the scientific community and had been made publicly available on the GeneLab portal, guaranteeing comparability of datasets across spaceflight experiments. To ensure reliability of data generation, SPL leverages cutting-edge innovations in laboratory automation for sample processing. By leveraging these state-of-the-art platforms, SPL achieves high levels of data reproducibility while significantly minimizing sources of bias and variability, especially across experiments with large numbers of samples. Over the past few years, the space biology investigator community has accessed SPL-generated data from the Open Science Data Repository for a myriad of data re-analysis and re-use studies. We observe a trend that in-house SPL-generated data consistently outperforms outsourced sequencing data in terms of technical standards, quality control metrics, timeliness of data delivery, and sequencing and reagent efficiency. Superior data generation has and will continue to enable discoveries in disease, diagnostic tools, and the biological effects of long duration spaceflight.

GeneLab↗

Data Albums: An Event Driven Search, Aggregation and Curation Tool for Earth Science

One of the largest continuing challenges in any Earth science investigation is the discovery and access of useful science content from the increasingly large volumes of Earth science data and related information available. Approaches used in Earth science research such as case study analysis and climatology studies involve gathering discovering and gathering diverse data sets and information to support the research goals. Research based on case studies involves a detailed description of specific weather events using data from different sources, to characterize physical processes in play for a specific event. Climatology-based research tends to focus on the representativeness of a given event, by studying the characteristics and distribution of a large number of events. This allows researchers to generalize characteristics such as spatio-temporal distribution, intensity, annual cycle, duration, etc. To gather relevant data and information for case studies and climatology analysis is both tedious and time consuming. Current Earth science data systems are designed with the assumption that researchers access data primarily by instrument or geophysical parameter. Those who know exactly the datasets of interest can obtain the specific files they need using these systems. However, in cases where researchers are interested in studying a significant event, they have to manually assemble a variety of datasets relevant to it by searching the different distributed data systems. In these cases, a search process needs to be organized around the event rather than observing instruments. In addition, the existing data systems assume users have sufficient knowledge regarding the domain vocabulary to be able to effectively utilize their catalogs. These systems do not support new or interdisciplinary researchers who may be unfamiliar with the domain terminology. This paper presents a specialized search, aggregation and curation tool for Earth science to address these existing challenges. The search tool automatically creates curated "Data Albums", aggregated collections of information related to a specific science topic or event, containing links to relevant data files (granules) from different instruments; tools and services for visualization and analysis; and information about the event contained in news reports, images or videos to supplement research analysis. Curation in the tool is driven via an ontology based relevancy ranking algorithm to filter out non-relevant information and data.

Ramachandran, Rahul↗

Radio science ground data system for the Voyager-Neptune encounter, part 1

The Voyager radio science experiments at Neptune required the creation of a ground data system array that includes a Deep Space Network complex, the Parkes Radio Observatory, and the Usuda deep space tracking station. The performance requirements were based on experience with the previous Voyager encounters, as well as the scientific goals at Neptune. The requirements were stricter than those of the Uranus encounter because of the need to avoid the phase-stability problems experienced during that encounter and because the spacecraft flyby was faster and closer to the planet than previous encounters. The primary requirement on the instrument was to recover the phase and amplitude of the S- and X-band (2.3 and 8.4 GHz) signals under the dynamic conditions encountered during the occultations. The primary receiver type for the measurements was open loop with high phase-noise and frequency stability performance. The receiver filter bandwidth was predetermined based on the spacecraft's trajectory and frequency uncertainties.

Kursinski, E. R.↗

A Case Study Comparing Citizen Science Aurora Data with Global Auroral Boundaries Derived from Satellite Imagery and Empirical Models

Aurorasaurus is a citizen science project that offers a new, global data source consisting of ground-based reports of the aurora. For this case study, aurora data collected during the 17-18 March 2015 geomagnetic storm are examined to identify their conjunctions with Defense Meteorological Satellite Program (DMSP) satellite passes over the high latitude auroral regions. This unique set of aurora data can provide ground-truth validation of existing auroral precipitation models. Particularly, the solar wind driven, Oval Variation, Assessment, Tracking, Intensity, and Online Nowcasting (OVATION) Prime 2013 (OP-13) model and a Kp-dependent model of Zhang-Paxton (Z-P) are utilized for our boundary validation efforts. These two similar models are compared for the first time. Global equatorward auroral boundaries are derived from the OP 13 model and the DMSP Special Sensor Ultraviolet Spectrographic Imager (SSUSI) far ultraviolet (FUV) data using the Z-P model at a fixed flux level of 0.2 erg cm(exp -2)s(exp -1). These boundaries are then compared with citizen science reports as well as with each other. Even though there are some large differences between the global boundaries for a few cases, the average difference is about 1.5 deg in geomagnetic latitude, with OP-13 being equatorward of Z-P model. When these boundaries are compared with each other as a function of local time, no clear overall trend as a function of local time was observed. It is also found that the ground based reports are more consistent with the predictions of the OP-13 model.

Kosar, Burcu C.↗