SEARCH · Search NASA
Results for “self-describing data”
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Determining the Completeness of the Nimbus Meteorological Data Archive
NASA launched the Nimbus series of meteorological satellites in the 1960s and 70s. These satellites carried instruments for making observations of the Earth in the visible, infrared, ultraviolet, and microwave wavelengths. The original data archive consisted of a combination of digital data written to 7-track computer tapes and on various film media. Many of these data sets are now being migrated from the old media to the GES DISC modern online archive. The process involves recovering the digital data files from tape as well as scanning images of the data from film strips. Some of the challenges of archiving the Nimbus data include the lack of any metadata from these old data sets. Metadata standards and self-describing data files did not exist at that time, and files were written on now obsolete hardware systems and outdated file formats. This requires creating metadata by reading the contents of the old data files. Some digital data files were corrupted over time, or were possibly improperly copied at the time of creation. Thus there are data gaps in the collections. The film strips were stored in boxes and are now being scanned as JPEG-2000 images. The only information describing these images is what was written on them when they were originally created, and sometimes this information is incomplete or missing. We have the ability to cross-reference the scanned images against the digital data files to determine which of these best represents the data set from the various missions, or to see how complete the data sets are. In this presentation we compared data files and scanned images from the Nimbus-2 High-Resolution Infrared Radiometer (HRIR) for September 1966 to determine whether the data and images are properly archived with correct metadata.
An interactive, discipline-independent data visualization system
The data visualization techniques used by National Space Science Data Center Graphics System are discussed. The system includes a self-describing data abstraction for the storage and manipulation of multidimensional data for discipline-independent scientific applications. The system's data format and graphics display are described. Consideration is given to techniques developed for the rapid display and manipulation of large, complex data sets and for manipulating algorithms with portable implementations which can operate on any data object or geometry. The processes of operating the system are given and examples are presented from applying the system to various types of data, including TOMS data, data from the International Satellite Cloud Climatology Project, and IRAS data.
The role of data management in discipline-independent data visualization
The common data format (CDF) is described in terms of its support applications for the database management of visualization systems. The CDF is a self-describing data abstraction technique for the storage and manipulation of multidimensional data that are based on block structures. The discipline-independent approach is designed to manage, manipulate, archive, display, and analyze data, and can be applied to heterogeneous equipment communicating different data structures over networks. An improved CDF version incorporates a hyperplane access allowing random aggregate access to subdimensional blocks within a multidimensional variable. The visualization pipeline is also discussed, which controls the flow of data and permits the visualization of different classes of data representation techniques. The system is found to accommodate a large variety of scientific data structures and large disk-based data sets.
Smartfiles: An OO approach to data file interoperability
Data files for scientific and engineering codes typically consist of a series of raw data values whose descriptions are buried in the programs that interact with these files. In this situation, making even minor changes in the file structure or sharing files between programs (interoperability) can only be done after careful examination of the data file and the I/O statement of the programs interacting with this file. In short, scientific data files lack self-description, and other self-describing data techniques are not always appropriate or useful for scientific data files. By applying an object-oriented methodology to data files, we can add the intelligence required to improve data interoperability and provide an elegant mechanism for supporting complex, evolving, or multidisciplinary applications, while still supporting legacy codes. As a result, scientists and engineers should be able to share datasets with far greater ease, simplifying multidisciplinary applications and greatly facilitating remote collaboration between scientists.
The Pleodata Language and Representation
Pleodata (Pleomorphic Data) is a markup language developed and used by the MLS group.
A Comprehensive Northern Hemisphere Particle Microphysics Data Set From the Precipitation Imaging Package
Microphysical observations of precipitating particles are critical data sources for numerical weather prediction models and remote sensing retrieval algorithms. However, obtaining coherent data sets of particle microphysics is challenging as they are often unindexed, distributed across disparate institutions, and have not undergone a uniform quality control process. This work introduces a unified, comprehensive Northern Hemisphere particle microphysical data set from the National Aeronautics and Space Administration precipitation imaging package (PIP), accessible in a standardized data format and stored in a centralized, public repository. Data is collected from 10 measurement sites spanning 34° latitude (37°N–71°N) over 10 years (2014–2023), which comprise a set of 1,070,000 precipitating minutes. The provided data set includes measurements of a suite of microphysical attributes for both rain and snow, including distributions of particle size, vertical velocity, and effective density, along with higher-order products including an approximation of volume-weighted equivalent particle densities, liquid equivalent snowfall, and rainfall rate estimates. The data underwent a rigorous standardization and quality assurance process to filter out erroneous observations to produce a self-describing, scalable, and achievable data set. Case study analyses demonstrate the capabilities of the data set in identifying physical processes like precipitation phase-changes at high temporal resolution. Bulk precipitation characteristics from a multi-site intercomparison also highlight distinct microphysical properties unique to each location. This curated PIP data set is a robust database of high-quality particle microphysical observations for constraining future precipitation retrieval algorithms, and offers new insights toward better understanding regional and seasonal differences in bulk precipitation characteristics.
AMPR CAMP2Ex Calibrated and Quality-Controlled Dataset Level 2B, Revision B
Data were acquired by the Advanced Microwave Precipitation Radiometer (AMPR) during the Cloud, Aerosol and Monsoon Processes Philippines Experiment (CAMP2Ex) field campaign in August-October of 2019. These files include the Level 2B calibrated, corrected, and geo-referenced brightness temperature for the four AMPR-observed frequencies (10, 19, 37, 85 GHz). These data are archived in a self-describing, Climate and Forecasting (CF) 1.6-compliant Version 4 Network Common Data Format (netCDF4) format.Python software has been developed for reading, plotting, and providing some additional analysis capabilities. This software is available from: https://github.com/nasa/pyampr.The AMPR instrument is explained in more detail here: https://weather.msfc.nasa.gov/ampr/.These data have been determined to be viable for publishable scientific research, and alsoshould be useful for generating quicklooks or understanding what happened during a flight. Note: AMPRis not expected to provide useful data during significant aircraft maneuvers
National Climate Assessment - Land Data Assimilation System (NCA-LDAS) Data and Services at NASA GES DISC
The National Climate Assessment-Land Data Assimilation System (NCA-LDAS) is an Integrated Terrestrial Water Analysis, and is one of NASAs contributions to the NCA of the United States. The NCA-LDAS has undergone extensive development, including multi-variate assimilation of remotely-sensed water states and anomalies as well as evaluation and verification studies, led by the Goddard Space Flight Centers Hydrological Sciences Laboratory (HSL). The resulting NCA-LDAS data have recently been released to the general public and include those from the Noah land-surface model (LSM) version 3.3 (Noah-3.3) and the Catchment LSM version Fortuna-2.5 (CLSM-F2.5). Standard LSM output variables including soil moistures temperatures, surface fluxes, snow cover depth, groundwater, and runoff are provided, as well as streamflow using a river routing system. The NCA-LDAS data are archived at and distributed by the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC). The data can be accessed via HTTP, OPeNDAP, Mirador search and download, and NASA Earth data Search. To further facilitate access and use, the NCA-LDAS data are integrated into the NASA Giovanni, for quick visualization and analysis, and into the Data Rods system, for retrieval of time series of long time periods. The temporal and spatial resolutions of the NCA-LDAS data are, respectively, daily-averages and 0.125x0.125 degree, covering North America (25N 53N; 125W 67W) and the period January 1979 to December 2015. The data files are in self-describing, machine-independent, CF-compliant netCDF-4 format.
CODARcode/MGARD
MGARD is a software providing error-controlled lossy compression and data refactoring based on multi-grid theories. It transforms floating-point scientific data into a multilevel representation, followed by quantization and lossless encoding processes, resulting in a self-describing compressed buffer. It supports diverse data topologies, error control norms, and computing architectures.
Experimenter's laboratory for visualized interactive science
The science activities of the 1990's will require the analysis of complex phenomena and large diverse sets of data. In order to meet these needs, we must take advantage of advanced user interaction techniques: modern user interface tools; visualization capabilities; affordable, high performance graphics workstations; and interoperable data standards and translator. To meet these needs, we propose to adopt and upgrade several existing tools and systems to create an experimenter's laboratory for visualized interactive science. Intuitive human-computer interaction techniques have already been developed and demonstrated at the University of Colorado. A Transportable Applications Executive (TAE+), developed at GSFC, is a powerful user interface tool for general purpose applications. A 3D visualization package developed by NCAR provides both color shaded surface displays and volumetric rendering in either index or true color. The Network Common Data Form (NetCDF) data access library developed by Unidata supports creation, access and sharing of scientific data in a form that is self-describing and network transparent. The combination and enhancement of these packages constitutes a powerful experimenter's laboratory capable of meeting key science needs of the 1990's. This proposal encompasses the work required to build and demonstrate this capability.
Experimenter's laboratory for visualized interactive science
The science activities of the 1990's will require the analysis of complex phenomena and large diverse sets of data. In order to meet these needs, we must take advantage of advanced user interaction techniques: modern user interface tools; visualization capabilities; affordable, high performance graphics workstations; and interoperatable data standards and translator. To meet these needs, we propose to adopt and upgrade several existing tools and systems to create an experimenter's laboratory for visualized interactive science. Intuitive human-computer interaction techniques have already been developed and demonstrated at the University of Colorado. A Transportable Applications Executive (TAE+), developed at GSFC, is a powerful user interface tool for general purpose applications. A 3D visualization package developed by NCAR provides both color-shaded surface displays and volumetric rendering in either index or true color. The Network Common Data Form (NetCDF) data access library developed by Unidata supports creation, access and sharing of scientific data in a form that is self-describing and network transparent. The combination and enhancement of these packages constitutes a powerful experimenter's laboratory capable of meeting key science needs of the 1990's. This proposal encompasses the work required to build and demonstrate this capability.
Microphone Phased Array NetCDF/HDF5 Archival Files: Application Program Interface Reference
An application program interface (API) has been developed for the creation and access of structured data files generated by microphone phased arrays utilized in aeroacoustics research. Two structured binary file formats are supported, namely NetCDF (Network Common Data Form) and HDF5 (Hierarchical Data Format) files. The API consists of a library of routines callable from C, Fortran or Matlab, with native versions of the API provided for each language. The libraries are divided into categories for file handling, file definition and initialization, data writing, data recovery, and error handling. The API is intended to provide a mechanism for generating self-describing binary files for long-term archiving of raw and processed data generated by phased array systems.
Compressing Aviation Data in XML Format
Design, operations and maintenance activities in aviation involve analysis of variety of aviation data. This data is typically in disparate formats making it difficult to use with different software packages. Use of a self-describing and extensible standard called XML provides a solution to this interoperability problem. XML provides a standardized language for describing the contents of an information stream, performing the same kind of definitional role for Web content as a database schema performs for relational databases. XML data can be easily customized for display using Extensible Style Sheets (XSL). While self-describing nature of XML makes it easy to reuse, it also increases the size of data significantly. Therefore, transfemng a dataset in XML form can decrease throughput and increase data transfer time significantly. It also increases storage requirements significantly. A natural solution to the problem is to compress the data using suitable algorithm and transfer it in the compressed form. We found that XML-specific compressors such as Xmill and XMLPPM generally outperform traditional compressors. However, optimal use of Xmill requires of discovery of optimal options to use while running Xmill. This, in turn, depends on the nature of data used. Manual disc0ver.y of optimal setting can require an engineer to experiment for weeks. We have devised an XML compression advisory tool that can analyze sample data files and recommend what compression tool would work the best for this data and what are the optimal settings to be used with a XML compression tool.
File Specification for M2AMIP Products: Version Number - 1.0
The Modern-Era Retrospective analysis for Research and Applications, Version 2 (MERRA-2) is an atmospheric reanalysis computed with the Goddard Earth Observing (EOS) System, Version 5.12.4 (GEOS) data assimilation system (Gelaro et al., 2017). To supplement the reanalysis, the GEOS General Circulation Model (GCM) used in MERRA-2 has been used to generate a 10-member ensemble of simulations, configured following the convention of the Atmospheric Model Intercomparison Project (AMIP; Gates et al., 1992). Each ensemble member was initialized using meteorological fields from a different date in November 1979. The AMIP simulations used the sea-surface temperature (SST) and sea-ice boundary conditions that were used in MERRA-2 (Bosilovich et al., 2016). This 10-member ensemble of AMIP simulations, denoted M2AMIP, is available for download in a group of self-describing files, which are documented in this office note. All data collections are provided on the same horizontal grid as MERRA-2. This grid has 576 points in the longitudinal direction and 361 points in the latitudinal direction, corresponding to a resolution of 0.625 degrees by 0.5 degrees. Although data collections are available at this grid, all fields are computed on a cubed-sphere grid with an approximate resolution of 50 km by 50 km and are then spatially interpolated to the latitude-longitude grid. There are no changes in the vertical grids used: variables are provided on either the native vertical grid of 72 model layers, or interpolated to 42 standard pressure levels. Unlike MERRA, no data collections are available at the vertical layer edges. More details on the grid are provided in Section 4. MERRA-2 introduced observation-based precipitation forcing for the land surface parameterization and the corresponding variable PRECTOTCORR in the MERRA-2 FLX (surface turbulent fluxes and related quantities) and LFO (land-surface forcing) collections (see Section 6; Reichle et al., 2017). While this variable is still available for M2AMIP, there was no observation-based forcing, making the value identical to the model derived precipitation, PRECTOT. Similarly, without data assimilation, the values for the analysis increments, D*DTANA, in the tendency and vertically integrated file collections are zero. The M2AMIP data are available for download online through the NASA Center for Climate Simulation (NCCS) DataPortal (https://portal.nccs.nasa.gov/datashare/gmao_m2amip/). Data are arranged in subdirectories based on ensemble member, followed by year and month. Control files that are compatible with the Grid Analysis and Display System (GrADS) are available in the ctl_daily and ctl_monthly directories for the hourly, three hourly, and monthly mean data. Control files for the monthly mean diurnal cycle can be found in the ctl_diurnal subdirectory within the directory for each individual ensemble member.