Search NASA⌕ Search

SEARCH · Search NASA

Results for “data processing and archiving”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Fire Analysis of the Thomas Fire in California Using NASA Data in a GIS

NASA's Earth Observing System Data Information System (EOSDIS) manages Earth Observation satellites and the Distributed Active Archive Centers (DAACs), where the data is stored and processed. This poster presents the use of satellite data from NASA's inventory that have the potential for use in identification and analysis of forest fire risk and subsequent after-effects within a geographic information system (GIS). The challenge is that Earth Observation data is complicated. There is plenty of data available, however, the science teams have had a "top-down" approach: define what it is you are trying to study -select a set of satellite(s) and sensor(s), and drill down for the data.

Worldview↗

Satellite image analysis using neural networks

The tremendous backlog of unanalyzed satellite data necessitates the development of improved methods for data cataloging and analysis. Ford Aerospace has developed an image analysis system, SIANN (Satellite Image Analysis using Neural Networks) that integrates the technologies necessary to satisfy NASA's science data analysis requirements for the next generation of satellites. SIANN will enable scientists to train a neural network to recognize image data containing scenes of interest and then rapidly search data archives for all such images. The approach combines conventional image processing technology with recent advances in neural networks to provide improved classification capabilities. SIANN allows users to proceed through a four step process of image classification: filtering and enhancement, creation of neural network training data via application of feature extraction algorithms, configuring and training a neural network model, and classification of images by application of the trained neural network. A prototype experimentation testbed was completed and applied to climatological data.

Sheldon, Roger A.↗

Satellite Ocean-Color Validation Using Ships of Opportunity

The investigation s main objective is to collect from platforms of opportunity (merchant ships, research vessels) concomitant normalized water-leaving radiance and aerosol optical thickness data over the world s oceans. A global, long-term data set of these variables is needed to verify whether satellite retrievals of normalized water-leaving radiance are within acceptable error limits and, eventually, to adjust atmospheric correction schemes. To achieve this objective, volunteer officers, technicians, and scientists onboard the selected ships collect data from portable SIMBAD and Advanced SIMBAD (SIMBADA) radiometers. These instruments are specifically designed for evaluation of satellite-derived ocean color. They measure radiance in spectral bands typical of ocean-color sensors. The SIMBAD version measures in 5 spectral bands centered at 443, 490, 560, 670, and 870 nm, and the Advanced SIMBAD version in 11 spectral bands centered at 350, 380, 412, 443, 490, 510, 565, 620, 670, 750, and 870 nm. Aerosol optical thickness is obtained by viewing the sun disk like a classic sun photometer. Normalized water-leaving radiance, or marine reflectance, is obtained by viewing the ocean surface through a vertical polarizer in a specific geometry (nadir angle of 45o and relative azimuth angle of 135deg) to minimize direct sun glint and reflected sky radiation. The SIMBAD and SIMBADA data, after proper quality control and processing, are delivered to the SIMBIOS project office for inclusion in the SeaBASS archive. They complement data collected in a similar way by the Laboratoire d'Optique Atmospherique of the University of Lille, France. The SIMBAD and SIMBADA data are used to check the radiometric calibration of satellite ocean-color sensors after launch and to evaluate derived ocean-color variables (i.e., normalized water-leaving radiance, aerosol optical thickness, and aerosol type). Analysis of the SIMBAD and SIMBADA data provides information on the accuracy of satellite retrievals of normalized water-leaving radiance, an understanding of the discrepancies between satellite and in situ data, and algorithms that reduce the discrepancies, contributing to more accurate and consistent global ocean color data sets.

Frouin, Robert↗

Rapid Measurements of Aerosol Ionic Composition and 3-10 nm Particle Size Distributions On The NASA P3 To Better Quantify Processes Affecting Aerosols Advected From East Asia

The Particle Into Liquid Sample (PILS) was deployed on the NASA P3 for airborne measurements of fine particle ionic chemical composition. The data have been quality assured and reside in the NASA data archive. We have analyzed our data to characterize the sources and atmospheric processing of fine aerosol particles advected from the region during the experiments. Fine particle water-soluble potassium was found to serve as a useful aerosol tracer for biomass smoke. Ratios of PILS potassium to sulfate are used as a means of estimating the percent contribution of biomass burning to fine particle mass in mixed plumes advecting from Asia. The high correlations between K+ and NO3(sup -) and NH4(sup +)' indicated that biomass burning was a significant source of these aerosol compounds in the region. It is noteworthy that the air mass containing the highest concentrations of fine particles recorded in all of ACE-Asia and TRACE-P appeared to be advecting from the Bejing/Tientsin urban region and also had the highest K(+), NO3(sup -) and NH4(sup +) concentrations of both studies. Based on K+/SO4(sup 2-) ratio's, we estimated that the plume was composed of approx. 60% biomass burning emissions, possibly from the use of bio-fuels in the urban regions.

Weber, Rodney J.↗

KDD Services at the Goddard Earth Sciences Distributed Active Archive Center

NASA's Goddard Earth Sciences Distributed Active Archive Center (GES DAAC) processes, stores and distributes earth science data from a variety of remote sensing satellites. End users of the data range from instrument scientists to global change and climate researchers to federal agencies and foreign governments. Many of these users apply data mining techniques to large volumes of data (up to 1 TB) received from the GES DAAC. However, rapid advances in processing power are enabling increases in data processing that are outpacing tape drive performance and network capacity. As a result, the proportion of data that can be distributed to users continues to decrease. As mitigation, we are migrating more data mining and mining preparation activities into the data center in order to reduce the data volume that needs to be distributed and to offer the users a more useful and manageable product. This migration of activities faces a number of technical and human-factor challenges. As data reduction and mining algorithms are normally quite specific to the user's research needs, the user's algorithm must be integrated virtually unchanged into the archive environment. Also, the archive itself is busy with everyday data archive and distribution activities and cannot be dedicated to, or even impacted by, the mining activities. Therefore, we schedule KDD 'campaigns' (similar to reprocessing campaigns), during which we schedule a wholesale retrieval of specific data products, offering users the opportunity to extract information from the data being retrieved during the campaign.

Lynnes, Christopher↗

Analysis Ready Data in Analytics Optimized Data Stores for Analysis of Big Earth Data in the Cloud

Cloud computing offers the possibility of making the analysis of Big Data approachable for a wider community due to affordable access to computing power, an ecosystem of usable tools for parallel processing, and migration of many large datasets to archives in the cloud, allowing data-proximal computing. Generally, data analysis acceleration in the cloud comes from running multiple nodes in a split-combine-apply strategy. Data systems such as the Earth Observing System Data and Information System are in a position to "pre-split" the data by storing them in a data store that is optimized for data parallel computing, i.e., an Analytics-Optimized Data Store (AODS). A variety of approaches to AODS are possible, from highly scalable databases to scalable filesystems to data formats optimized for cloud access (e.g., zarr and cloud-optimized datasets), with the optimal choice dependent on both the types of analysis and the geospatial structure of the data. A key question is how much preprocessing of the data to do, both before splitting and as the first part of the apply step. Again, the geospatial structure of the data and the analysis type influence the decision, with the added complexity of the user type. Trans-disciplinary users who are not well-versed in the nuances of quality-filtering and georeferencing of remote sensing orbit/swath/scene data tend to ask for more highly processed data, relying on the data provider to make sensible decisions on preprocessing parameters. (This accounts for the popularity of "Level 3" gridded data, despite the lower spatial resolution it provides.) In this case, data can be preprocessed before the split, resulting in higher performance in the rest of the "apply" step, which can be transformative for use cases such as interactive data exploration at scale. Discipline researchers who are experienced with remote sensing data often prefer more flexibility in customizing the preprocessing data into Analysis Ready Data, resulting in more need for on-the-fly preprocessing.

Lynnes, Christopher↗

Life Sciences Data Archives (LSDA) in the Post-Shuttle Era

Now, more than ever before, NASA is realizing the value and importance of their intellectual assets. Principles of knowledge management-the systematic use and reuse of information, experience, and expertise to achieve a specific goal-are being applied throughout the agency. LSDA is also applying these solutions, which rely on a combination of content and collaboration technologies, to enable research teams to create, capture, share, and harness knowledge to do the things they do well, even better. In the early days of spaceflight, space life sciences data were collected and stored in numerous databases, formats, media-types and geographical locations. These data were largely unknown/unavailable to the research community. The Biomedical Informatics and Health Care Systems Branch of the Space Life Sciences Directorate at JSC and the Data Archive Project at ARC, with funding from the Human Research Program through the Exploration Medical Capability Element, are fulfilling these requirements through the systematic population of the Life Sciences Data Archive. This project constitutes a formal system for the acquisition, archival and distribution of data for HRP-related experiments and investigations. The general goal of the archive is to acquire, preserve, and distribute these data and be responsive to inquiries for the science communities. Information about experiments and data, as well as non-attributable human data and data from other species' are available on our public Web site http://lsda.jsc.nasa.gov. The Web site also includes a repository for biospecimens, and a utilization process. NASA has undertaken an initiative to develop a Shuttle Data Archive repository. The Shuttle program is nearing its end in 2010 and it is critical that the medical and research data related to the Shuttle program be captured, retained, and usable for research, lessons learned, and future mission planning. Communities of practice are groups of people who share a concern or a passion for something they do, and learn how to do it better as they interact regularly. LSDA works with the HRP community of practice to ensure that we are preserving the relevant research and data they need in the LSDA repository. An evidence-based approach to risk management is required in space life sciences. Evidence changes over time. LSDA has a pilot project with Collexis, a new type of Web-based search engine. Collexis differentiates itself from full-text search engines by making use of thesauri for information retrieval. The high-quality search is based on semantics that have been defined in a life sciences ontology. Additionally, Collexis' matching technology is unique, allowing discovery of partially matching dicuments. Users do not have to construct a complicated (Boolean) search query, but can simply enter a free text search without the risk of getting "no results". Collexis may address these issues by virtue of its retrieval and discovery capabilities across multiple repositories.

Fitts, Mary A.↗

Hyporheic-zone Processes and Stream Oxygen Dynamics: Insights from a Multiscale Reactive Transport Model: Modeling Archive

This archive contains the data and Python scripts required to reproduce the analyses and figures in the study: Gomez-Velez, J. D., Rathore, S. S., Cohen, M. J., & Painter, S. L. (2025). Hyporheic-zone Processes and Stream Oxygen Dynamics: Insights from a Multiscale Reactive Transport Model. Submitted to Water Resources Research. The analysis utilizes the subgrid model Advection Dispersion Equation with Lagrangian Subgrids (ADELS) implemented in the Advanced Terrestrial Simulator (ATS; https://amanzi.github.io/ats/stable/). In this case, the ATS and Amanzi versions are (1) ATS version 1.5.1_f5ba18f8 and (2) Amanzi version 1.6-dev_53444cca4. The repository includes a Jupyter Notebook and the necessary data (Pandas DataFrames stored as pickle files) to generate the figures for the manuscript. Additionally, it contains Python scripts to create ATS input files, run the ATS simulations, and post-process the results. Finally, it provides routines for parameter estimation using the Single-Station Metabolism (SSM) model with the Differential Evolution Adaptive Metropolis (DREAM) Markov Chain Monte Carlo (MCMC) algorithm with ZS enhancements (DREAM-ZS).

54 ENVIRONMENTAL SCIENCES↗

Globally Gridded Satellite (GridSat) Observations for Climate Studies

Geostationary satellites have provided routine, high temporal resolution Earth observations since the 1970s. Despite the long period of record, use of these data in climate studies has been limited for numerous reasons, among them: there is no central archive of geostationary data for all international satellites, full temporal and spatial resolution data are voluminous, and diverse calibration and navigation formats encumber the uniform processing needed for multi-satellite climate studies. The International Satellite Cloud Climatology Project set the stage for overcoming these issues by archiving a subset of the full resolution geostationary data at approx.10 km resolution at 3 hourly intervals since 1983. Recent efforts at NOAA s National Climatic Data Center to provide convenient access to these data include remapping the data to a standard map projection, recalibrating the data to optimize temporal homogeneity, extending the record of observations back to 1980, and reformatting the data for broad public distribution. The Gridded Satellite (GridSat) dataset includes observations from the visible, infrared window, and infrared water vapor channels. Data are stored in the netCDF format using standards that permit a wide variety of tools and libraries to quickly and easily process the data. A novel data layering approach, together with appropriate satellite and file metadata, allows users to access GridSat data at varying levels of complexity based on their needs. The result is a climate data record already in use by the meteorological community. Examples include reanalysis of tropical cyclones, studies of global precipitation, and detection and tracking of the intertropical convergence zone.

Knapp, Kenneth R.↗

ALSEP termination report

The Apollo Lunar Surface Experiments Package (ALSEP) final report was prepared when support operations were terminated September 30, 1977, and NASA discontinued the receiving and processing of scientific data transmitted from equipment deployed on the lunar surface. The ALSEP experiments (Apollo 11 to Apollo 17) are described and pertinent operational history is given for each experiment. The ALSEP data processing and distribution are described together with an extensive discussion on archiving. Engineering closeout tests and results are given, and the status and configuration of the experiments at termination are documented. Significant science findings are summarized by selected investigators. Significant operational data and recommendations are also included.

Bates, J. R.↗

Overview of the Solar-B Mission

The Solar-B mission is a collaboration between the Japan Aerospace Exploration Agency, Institute of Space and Astronautical Science, the National Aeronautics and Space Administration (NASA) and the Particle Physics and Astronomy Research Council (PPARC) of the United Kingdom and the European Space Agency. The principal scientific goals of the mission are to understand the processes of magnetic field generation, transport and ultimate dissipation of solar magnetic fields and how the release of magnetic energy is responsible for the heating and structuring of the chromosphere and corona. The scientific payload consists of three instruments: the Solar Optical Telescope that consists of the Optical Telescope Assembly and the Focal Plane Package (FPP), the X-ray Telescope and the EUV Imaging Spectrometer Each instrument is a result of the combined talents of all the members of the international team and their design and performance is described in separate papers in this session. The instruments are designed to work together as an 'observatory' simultaneously studying the target, at which the spacecraft is pointed, at different levels in the atmosphere. The spacecraft is scheduled for launch in September 2006 from the Uchinoura Space Center into a 600 km circular, sun-synchronous, polar orbit with a nominal elevation of 97.9 degrees. The orbit provides at least two morning and two evening contacts in Japan. Morning contacts are used for recovering quick look science data and the evening contacts for uploading commands. In addition ESA will provide 15 contacts per day from the Norwegian high latitude (78deg 14' N) ground station at Svalbard. The data downloads are transmitted to the ISAS Sirius database. They will be reformatted into FITS files and archived as Level 0 data on the ISAS DARTS system and made available to the scientific community. Scientific operations will be conducted from the IS AS facility located in Sagamihara, Japan. They are separated into planning, implementation and archiving. The planning process involves monthly, weekly and daily planning meetings. All scientific data will be made available after the first six month approximately one week after its collection.

Davis, John M.↗

2018 NISAR Applications Workshop: Forest and Disturbance; Workshop Report

Forest lands cover the globe and are important sources for providing ecosystem services including: carbon sequestration, biodiversity, timber, air and water quality. As such, counties around the world have dedicated programs for managing them. Accurate and timely information concerning the status of these forests (moisture, biomass, disturbance type, etc.) is essential to those Nations’ human and ecological health as well as economy. The joint NASA/US Forest Service workshop focused on arming forest land managers with observations and remote sensing information from the upcoming NASA-ISRO (Indian Space Research Organization) SAR (Synthetic Aperture Radar) (NISAR) satellite mission (expected to launch early 2022). Participants included representatives from different US Federal Agencies, private sector, and non-governmental organizations (NGO) that are key players in facilitating integration of Earth Observations (EO) into forest management and decision support workflows. They included scientists, technicians, and program managers with a responsibility for data acquisition and exploitation such as product development, delivery, and use, and capacity building. Discussions were held over two days to convey the broader forest and disturbance community information needs for various representative participants and programs and to facilitate the delivery of NISAR mission geospatial products and observational capabilities. Case studies were presented to demonstrate the current state of practice in the use of SAR remote sensing for applications of direct importance to forest and disturbance land management community. Eleven organizations presented their information requirements in response to a set of questions provided by the NASA team, then the NASA team responded by describing the degree to which NISAR could meet these requirements. Discussion ensued about needed data product specifications to increase utility (e.g., projection, latency, etc.), tools and capacity building. The general findings of this workshop were that (a) NISAR observations will be particularly useful to the global forest carbon and disturbance monitoring applications, but that certain data product design decisions (projections and radiometric and terrain corrections) need to be considered to increase utility; b) the biomass and disturbance detection algorithms meet many of the community needs, however there are other information products of value (e.g., soil moisture or disturbance classification, not just detection) and all products should be compliant with existing community standards for reporting uncertainty; c) providing SAR education to the community will be key specifically thinking about putting the information first and the SAR theory second, providing a simple guide of standard data processing steps (e.g., dB (decibel) to power conversion and speckle filtering); d) the community needs a user-friendly interface for finding free, archived data over their geographic regions of interest; e) user-friendly tools that connect to open-sources GIS (Global Information System) software (e.g., QGIS (Quantum GIS)) that include a graphical user interface (GUI) for SAR processing that enables both download and cloud processing. To integrate these findings and prepare the community before NISAR launches, it was suggested that there be a dedicated NISAR Forest and Disturbance Applications Working Group (as per the specifications in the NISAR Utilization Plan). After launch, it was decided that the community continue capacity building activities.

Stavros, Natasha↗

MODIS Snow and Ice Products from the NSIDC DAAC

The National Snow and Ice Data Center (NSIDC) Distributed Active Archive Center (DAAC) provides data and information on snow and ice processes, especially pertaining to interactions among snow, ice, atmosphere and ocean, in support of research on global change detection and model validation, and provides general data and information services to cryospheric and polar processes research community. The NSIDC DAAC is an integral part of the multi-agency-funded support for snow and ice data management services at NSIDC. The Moderate Resolution Imaging Spectroradiometer (MODIS) will be flown on the first Earth Observation System (EOS) platform (AM-1) in 1998. The MODIS Instrument Science Team is developing geophysical products from data collected by the MODIS instrument, including snow and ice products which will be archived and distributed by NSIDC DAAC. The MODIS snow and ice mapping algorithms will generate global snow, lake ice, and sea ice cover products on a daily basis. These products will augment the existing record of satellite-derived snow cover and sea ice products that began about 30 years ago. The characteristics of these products, their utility, and comparisons to other data set are discussed. Current developments and issues are summarized.

Scharfen, Greg R.↗

NASA-SETI microwave observing project: Targeted Search Element (TSE)

The Targeted Search Element (TSE) performs one of two complimentary search strategies of the NASA-SETI Microwave Observing Project (MOP): the targeted search. The principle objective of the targeted search strategy is to scan the microwave window between the frequencies of one and three gigahertz for narrowband microwave emissions eminating from the direction of 773 specifically targeted stars. The scanning process is accomplished at a minimum resolution of one or two Hertz at very high sensitivity. Detectable signals will be of a continuous wave or pulsed form and may also drift in frequency. The TSE will possess extensive radio frequency interference (RFI) mitigation and verification capability as the majority of signals detected by the TSE will be of local origin. Any signal passing through RFI classification and classifiable as an extraterrestrial intelligence (ETI) candidate will be further validated at non-MOP observatories using established protocol. The targeted search will be conducted using the capability provided by the TSE. The TSE provides six Targeted Search Systems (TSS) which independently or cooperatively perform automated collection, analysis, storage, and archive of signal data. Data is collected in 10 megahertz chunks and signal processing is performed at a rate of 160 megabits per second. Signal data is obtained utilizing the largest radio telescopes available for the Targeted Search such as those at Arecibo and Nancay or at the dedicated NASA-SETI facility. This latter facility will allow continuous collection of data. The TSE also provides for TSS utilization planning, logistics, remote operation, and for off-line data analysis and permanent archive of both the Targeted Search and Sky Survey data.

Webster, L. D.↗

High Performance Access to Archival Data Stored in HDF4 and HDF5 on Cloud Object Stores Without Reformatting the Files

Cloud computing offers numerous advantages for users of extensive Earth science data collections. These benefits encompass direct online access to data files and granules from any location, scalable access supporting parallel computing workflows, and flexible computing tools enabling innovative experimentation with processing techniques. However, older archival file formats designed for distinct computing systems hinder efficient access to decade-long time-series data when compared to data stored in modern cloud-optimized formats like Web Object Stores (WOS), exemplified by Amazon Web Services’ Simple Storage Service (S3). We describe DMR++ (Dataset Metadata Response plus plus), a technology facilitating efficient access to HDF5 (Hierarchical Data Format, version 5) and HDF4 files stored on WOS systems without requiring data reformatting. DMR++ achieves performance comparable to technologies like Zarr while preserving the original file structure, a substantial benefit considering the vast quantity of archival files held by organizations such as NASA. Moreover, DMR++ typically outperforms cloud-optimized versions of HDF5. Essentially an XML (Extensible Markup Language) document usually stored alongside the described data, DMR++ can also be generated on-the-fly but is generally created during data staging to the WOS. Archival files that use HDF4/5 often store large arrays of numerical data. The data in these files is often compressed, typically reducing their size by a factor of four or more. To achieve efficient access to portions of those arrays, they are 'chunked' into smaller sub-arrays, each individually compressed. The chunk size is a compromise, where spinning disks can efficiently access data in smaller chunks while S3 favors larger chunks. A simple optimization of aggregating smaller chunks that are stored adjacently, transferring them in a single access and then individually decompressing them will improve performance. NASA data pose an additional challenge: special Application Programmer Interface (API) libraries are often needed to compute some variables. These libraries are incompatible with WOS environments. Our solution involves storing computed values in the DMR++ document or a companion file, making them accessible like other variables and eliminating the need for specialized APIs. We outline specific optimizations for both satellite grid and swath data stored in HDF4-EOS2 (Earth Observing System).

James Gallagher↗

Recommendations for a service framework to access astronomical archives

There are a large number of astronomical archives and catalogs on-line for network access, with many different user interfaces and features. Some systems are moving towards distributed access, supplying users with client software for their home sites which connects to servers at the archive site. Many of the issues involved in defining a standard framework of services that archive/catalog suppliers can use to achieve a basic level of interoperability are described. Such a framework would simplify the development of client and server programs to access the wide variety of astronomical archive systems. The primary services that are supplied by current systems include: catalog browsing, dataset retrieval, name resolution, and data analysis. The following issues (and probably more) need to be considered in establishing a standard set of client/server interfaces and protocols: Archive Access - dataset retrieval, delivery, file formats, data browsing, analysis, etc.; Catalog Access - database management systems, query languages, data formats, synchronous/asynchronous mode of operation, etc.; Interoperability - transaction/message protocols, distributed processing mechanisms (DCE, ONC/SunRPC, etc), networking protocols, etc.; Security - user registration, authorization/authentication mechanisms, etc.; Service Directory - service registration, lookup, port/task mapping, parameters, etc.; Software - public vs proprietary, client/server software, standard interfaces to client/server functions, software distribution, operating system portability, data portability, etc. Several archive/catalog groups, notably the Astrophysics Data System (ADS), are already working in many of these areas. In the process of developing StarView, which is the user interface to the Space Telescope Data Archive and Distribution Service (ST-DADS), these issues and the work of others were analyzed. A framework of standard interfaces for accessing services on any archive system which would benefit archive user and supplier alike is proposed.

Travisano, J. J.↗

Developing a Machine-Learning-Based Processing Framework for Twitter and Other Crowdsourced Data

Crowdsourced data streams such as Twitter and other social media are important sources of real-time and historical global information for Earth science applications. At the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC), we have been exploring the Twitter data stream for its potential in augmenting the validation program of NASA's Global Precipitation Measurement (GPM) mission. To realize this potential, we need to increase the information density and enhance the quality of filtered precipitation tweets. We have implemented various components of a machine learning (ML)-based processing infrastructure for crowdsourced data that outputs, in this instance, useful and usable information derived from precipitation tweets. We have test enriched the Twitter stream with higher quality active tweets from those knowingly contributing to our effort and from existing crowdsourced programs (e.g., mPING, CoCoRaHS). We have experimented with various algorithms for processing tweets, including Naà ve Bayes, Convolutional Neural Network (CNN), Hierarchical Attention Network (HAN), and semi-supervised learning (with tri-training). Our current work focuses on (1) automated review of Earth science-related publications to determine relationships between discipline research needs and ML algorithms; (2) investigating Sequential Generative Adversarial Network (SeqGAN) for processing precipitation tweets for anomaly detection; and (3) managing crowdsourced data in a way that is compatible with existing NASA satellite data archives and using the data for ML applications. Key results include (1) network visualization of NLP-processed publications in various Earth science disciplines; (2) difference between GPM-linked, generated tweets and collected actual tweets that is small for GPM-determined light to moderate rain cases and high for GPM-determined heavy rain cases; and (3) identification of MongoDB for storing raw tweets and Zarr format for gridded tweets (compatible with GPM data). Our results have taken us a step closer to an operational ML-based tweet processing infrastructure and have already demonstrated that tweet-derived precipitation information is potentially useful for validation of Earth science satellite data.

Teng, William↗

NASA'S Earth Science Data Stewardship Activities

NASA has been collecting Earth observation data for over 50 years using instruments on board satellites, aircraft and ground-based systems. With the inception of the Earth Observing System (EOS) Program in 1990, NASA established the Earth Science Data and Information System (ESDIS) Project and initiated development of the Earth Observing System Data and Information System (EOSDIS). A set of Distributed Active Archive Centers (DAACs) was established at locations based on science discipline expertise. Today, EOSDIS consists of 12 DAACs and 12 Science Investigator-led Processing Systems (SIPS), processing data from the EOS missions, as well as the Suomi National Polar Orbiting Partnership mission, and other satellite and airborne missions. The DAACs archive and distribute the vast majority of data from NASA’s Earth science missions, with data holdings exceeding 12 petabytes The data held by EOSDIS are available to all users consistent with NASA’s free and open data policy, which has been in effect since 1990. The EOSDIS archives consist of raw instrument data counts (level 0 data), as well as higher level standard products (e.g., geophysical parameters, products mapped to standard spatio-temporal grids, results of Earth system models using multi-instrument observations, and long time series of Earth System Data Records resulting from multiple satellite observations of a given type of phenomenon). EOSDIS data stewardship responsibilities include ensuring that the data and information content are reliable, of high quality, easily accessible, and usable for as long as they are considered to be of value.

metadata↗