Search NASA⌕ Search

SEARCH · Search NASA

Results for “hierarchical data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Closing the Gap between FAIR Data Repositories and Hierarchical Data Formats

Many in the scientific community, particularly in publicly funded research, are pushing to adhere to more accessible data standards to maximize the findability, accessibility, interoperability, and reusability (FAIR) of scientific data, especially with the growing prevalence of machine learning augmented research. Online FAIR data repositories, such as the Open Science Framework (OSF), help facilitate the adoption of these standards by providing frameworks for storage, access, search, APIs, and other features that create organized hubs of scientific data. However, the wider acceptance of such repositories is hindered by the lack of support of hierarchical data formats, such as Technical Data Management Streaming (TDMS) and Hierarchical Data Format 5 (HDF5), that many researchers rely on to organize their datasets. Various tools and strategies should be used to allow hierarchical data formats, FAIR data repositories, and scientific organizations to work more seamlessly together. A pilot project at Los Alamos National Laboratory (LANL) addresses the disconnect between them by integrating the OSF FAIR data repository with hierarchical data renderers, extending support for additional file types in their framework. The multifaceted interactive renderer displays a tree of metadata alongside a table and plot of the data channels in the file. This allows users to quickly and efficiently load large and complex data files directly in the OSF webapp. Users who are browsing files can quickly and intuitively see the files in the way they or their colleagues structured the hierarchical form and immediately grasp their contents. This solution helps bridge the gap between hierarchical data storage techniques and FAIR data repositories, making both of them more viable options for scientific institutions like LANL which have been put off by the lack of integration between them.

97 MATHEMATICS AND COMPUTING↗

Hierarchical Data Format for Nuclear Data Sensitivities

The SCALE code system includes capabilities for sensitivity and uncertainty (S/U) analysis as part of its TSUNAMI code suite. The sensitivity of a quantity of interest (for example, an application’s $k_{eff}$) to nuclear data is stored as a profile in a text-based file, which is known as a sensitivity data file (SDF). The sensitivity profile can be used to calculate uncertainties, correlation coefficients, and similarity indices. One of the goals of the present work was to seek general performance improvements in the TSUNAMI code suite, starting with the TSUNAMI-IP code for calculating similarity indices. Through profiling, it was found that reading the text-based sensitivity files was a performance bottleneck in the TSUNAMI-IP code. In a typical TSUNAMI-IP calculation, an application might be compared to thousands of benchmarks, thus requiring the reading of thousands of SDFs. Reading of binary-based data is generally faster than reading text-based data. Hierarchical Data Format 5 (HDF5) is a binary-based format that also benefits from being portable, and it can be inspected with nonproprietary tools. This paper describes an HDF5-based file format that has been introduced for SDFs.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Hierarchical Data Format for Earth Observing System Data Product Developer's Guide

The "Hierarchical Data Format for Earth Observing System" talk will address the best practices for creating ESDIS data products. The work presented is done in support of Data Product Developers Guide Working Group with mission "to help data product developers make data usable for end users". During the presentation, we will use some examples of NASA data products and show how to modify them to make data more usable.

Data usability↗

Hierarchical Data Format (HDF) Status Update

In the "Hierarchical Data Format (HDF) Status Update" talk we will give an update on the recent and upcoming HDF5 software releases, and how to upgrade HDF5-based software to the new major releases HDF5 1.10.* to assure backward and forward compatibility for the files produces with the newest versions of the HDF5 software. We will also focus on a compression feature of HDF5 and will talk about new mechanism for storing HDF5 data in Object Store.

Virtual File Driver↗

A self-defining hierarchical data system

The Self-Defining Data System (SDS) is a system which allows the creation of self-defining hierarchical data structures in a form which allows the data to be moved between different machine architectures. Because the structures are self-defining they can be used for communication between independent modules in a distributed system. Unlike disk-based hierarchical data systems such as Starlink's HDS, SDS works entirely in memory and is very fast. Data structures are created and manipulated as internal dynamic structures in memory managed by SDS itself. A structure may then be exported into a caller supplied memory buffer in a defined external format. This structure can be written as a file or sent as a message to another machine. It remains static in structure until it is reimported into SDS. SDS is written in portable C and has been run on a number of different machine architectures. Structures are portable between machines with SDS looking after conversion of byte order, floating point format, and alignment. A Fortran callable version is also available for some machines.

Bailey, J.↗

Hierarchical Data Format for Nuclear Data Sensitivities [Slides]

An HDF5-based file format was introduced for the sensitivity data calculated by TSUNAMI. The format was defined to collect the sensitivity coefficients into hyperslabs, which optimizes file reading time and therefore improves the time-to-solution for applications. In future work, this format will be extended to store sensitivity data for depletion calculations.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Spatiotemporal Modeling of Real World Backsheets Field Survey Data: Hierarchical (Multilevel) Generalized Additive Models: Preprint

Assessing photovoltaic module backsheet durability is critical to increasing module lifetime. Lab based accelerating testing has recently failed to predict large scale failures of widely adopted polymeric materials. Field surveyed data is critical to assess the performance of component lifetime. Using a documented field survey protocol, 13 field surveys where conducted. Each measurement is encoded with it's spatial location in respect to the other modules. By combining field survey data on degradation predictors with real time satellite weather data, data-driven predictive models of backsheet degradation were trained. LOESS models were constructed to investigate the spatial dependence of measurements. It was found that micro-climatic effects like treelines, ground surface changes, and elevation changes effected the magnitude and variance of the measurements. A GAM model was created to predict the value of degradation based on measured predictors. The model includes variables on the climate of the system and the location of each measurement in the PV mounting structure. The model performed well with an adj:R2 of 0:95 for yellowness index prediction. The model was cross-validated using k-folds.

backsheet↗

Spatiotemporal Modeling of Real World Backsheets Field Survey Data: Hierarchical (Multilevel) Generalized Additive Models

Assessing photovoltaic module backsheet durability is critical to increasing module lifetime. Laboratory-based accelerating testing has recently failed to predict large scale failures of widely adopted polymeric materials. Additionally, there is a growing concern on characterizing the non-uniformity of field exposure. Therefore, data from field surveys are critical to assess the performance of component lifetimes. Using a documented field survey protocol, 19 field surveys were conducted. The focus of this survey strategy is to investigate spatial continuity in degradation modes. By combining field survey data with real-time satellite weather data, stressor / response models have been trained. Generalized additive Models (GAM) model was created to predict the value of degradation based on measured predictors. Two different GAM constructions were testing using different implementations of basis splines. The model includes variables on the environmental stressors of the system and the location of each measurement in the PV mounting structure. The incorporation of hierarchical structure into the models allowed for material specific degradation rates, while maintaining the assumption of a global trend. The model performed well with an adjusted R2 of 0.975 for yellowness index prediction.

backsheet↗

Multicast Routing of Hierarchical Data

The issue of multicast of broadband, real-time data in a heterogeneous environment, in which the data recipients differ in their reception abilities, is considered. Traditional multicast schemes, which are designed to deliver all the source data to all recipients, offer limited performance in such an environment, since they must either force the source to overcompress its signal or restrict the destination population to those who can receive the full signal. We present an approach for resolving this issue by combining hierarchical source coding techniques, which allow recipients to trade off reception bandwidth for signal quality, and sophisticated routing algorithms that deliver to each destination the maximum possible signal quality. The field of hierarchical coding is briefly surveyed and new multicast routing algorithms are presented. The algorithms are compared in terms of network utilization efficiency, lengths of paths, and the required mechanisms for forwarding packets on the resulting paths.

Shacham, Nachum↗

The Hierarchical Data Format for EOS (HDF-EOS)

HDF is a file format and a software library for data storage, management, exchange, and archiving. It is written and maintained by the National Center for Supercomputing Applications (NCSA). HDF5 has a very simple but versatile data model which is compatible with most competing formats. Through its grouping and linking mechanisms, the HDF5 data model enables complex data relationships and dependencies. HDF5 accommodates the inclusion of many common types of metadata and arbitrary types and quantities of user-defined metadata.

Ullman, Richard↗

International Satellite Cloud Climatology Project (ISCCP) Stage D1 3-Hourly Cloud Product - Revised Algorithm in Hierarchical Data Format (ISCCP_D1)

Since 1983 an international group of institutions has collected and analyzed satellite radiance measurements from up to five geostationary and two polar orbiting satellites to infer the global distribution of cloud properties and their diurnal, seasonal and interannual variations. The primary focus of the first phase of the project (1983-1995) was the elucidation of the role of clouds in the radiation budget (top of the atmosphere and surface). In the second phase of the project (1995 onwards) the analysis also concerns improving understanding of clouds in the global hydrological cycle. [Location=TROPOSPHERE] [Temporal_Coverage: Start_Date=1983-07-01; Stop_Date=] [Spatial_Coverage: Southernmost_Latitude=-90; Northernmost_Latitude=90; Westernmost_Longitude=-180; Easternmost_Longitude=180] [Data_Resolution: Latitude_Resolution=280 Km; Longitude_Resolution=280 Km; Temporal_Resolution=3 Hourly].

CLOUD LIQUID WATER PATH↗

International Satellite Cloud Climatology Project (ISCCP) Stage D2 Monthly Cloud Product - Revised Algorithm in Hierarchical Data Format (ISCCP_D2)

Since 1983 an international group of institutions has collected and analyzed satellite radiance measurements from up to five geostationary and two polar orbiting satellites to infer the global distribution of cloud properties and their diurnal, seasonal and interannual variations. The primary focus of the first phase of the project (1983-1995) was the elucidation of the role of clouds in the radiation budget (top of the atmosphere and surface). In the second phase of the project (1995 onwards) the analysis also concerns improving understanding of clouds in the global hydrological cycle. [Location=TROPOSPHERE] [Temporal_Coverage: Start_Date=1983-07-01; Stop_Date=] [Spatial_Coverage: Southernmost_Latitude=-90; Northernmost_Latitude=90; Westernmost_Longitude=-180; Easternmost_Longitude=180] [Data_Resolution: Latitude_Resolution=280 Km; Longitude_Resolution=280 Km; Temporal_Resolution=Monthly].

CLOUD TOP PRESSURE↗

Hierarchical Data Formats (HDF) Update

In this presentation, we will talk about the latest releases of HDF4 and HDF5 software and tools, new features available in HDF5, and roadmap for the HDF software. We will also solicit feedback from the users of HDF data and HDF application developers on new features and new tools. The talk will cover: Difference between 1.8 and 1.10 releases and how and when to move to the latest release Features of the recent HDF5 1.8.19, 1.10.1 and HDF 4.2.13 Overview of HDF View 3.0 and other enhancements to tools Supported compilers and systems Open discussion of new requirements and wish list of the HDF features Compression library for interoperability with h5py and Pandas and better floating-point data compression.

HDFView↗

Collaborating With Xarray to Enable Reading Hierarchical Data Files

NASA has a lot of expertise, but doesn’t need to write every single piece of code. Pangeo is a fantastic open-source community of tools for geoscience research. Xarray, a Python package, is a widely used part of this ecosystem for accessing and analyzing geoscience data.

Owen Littlejohns↗

Accelerating Multigrid-based Hierarchical Scientific Data Refactoring on GPUs

Rapid growth in scientific data and a widening gap between computational speed and I/O bandwidth make it increasingly infeasible to store and share all data produced by scientific simulations. Instead, we need methods for reducing data volumes: ideally, methods that can scale data volumes adaptively so as to enable negotiation of performance and fidelity tradeoffs in different situations. Multigrid-based hierarchical data representations hold promise as a solution to this problem, allowing for flexible conversion between different fidelities so that, for example, data can be created at high fidelity and then transferred or stored at lower fidelity via logically simple and mathematically sound operations. However, the effective use of such representations has been hindered until now by the relatively high costs of creating, accessing, reducing, and otherwise operating on such representations. We describe here highly optimized data refactoring kernels for GPU accelerators that enable efficient creation and manipulation of data in multigrid-based hierarchical forms. We demonstrate that our optimized design can achieve up to 250 TB/s aggregated data refactoring throughput—83% of theoretical peak—on 1024 nodes of the Summit supercomputer. We showcase our optimized design by applying it to a large-scale scientific visualization workflow and the MGARD lossy compression software.

Chen, Jieyang↗

Level 1 Processing of MODIS Direct Broadcast Data From Terra

In February 2000, an effort was begun to adapt the Moderate Resolution Imaging Spectroradiometer (MODIS) Level 1 production software to process direct broadcast data. Three Level 1 algorithms have been adapted and packaged for release: Level 1A converts raw (level 0) data into Hierarchical Data Format (HDF), unpacking packets into scans; Geolocation computes geographic information for the data points in the Level 1A; and the Level 1B computes geolocated, calibrated radiances from the Level 1A and Geolocation products. One useful aspect of adapting the production software is the ability to incorporate enhancements contributed by the MODIS Science Team. We have therefore tried to limit changes to the software. However, in order to process the data immediately on receipt, we have taken advantage of a branch in the geolocation software that reads orbit and altitude information from the packets themselves, rather than external ancillary files used in standard production. We have also verified that the algorithms can be run with smaller time increments (2.5 minutes) than the five-minute increments used in production. To make the code easier to build and run, we have simplified directories and build scripts. Also, dependencies on a commercial numerics library have been replaced by public domain software. A version of the adapted code has been released for Silicon Graphics machines running lrix. Perhaps owing to its origin in production, the software is rather CPU-intensive. Consequently, a port to Linux is underway, followed by a version to run on PC clusters, with an eventual goal of running in near-real-time (i.e., process a ten-minute pass in ten minutes).

Lynnes, Christopher↗

AIRS Data Subsetting Service at the Goddard Earth Sciences (GES) DISC/DAAC

The AIRS mission, as a combination of the Atmospheric Infrared Sounder (AIRS), the Advanced Microwave Sounding Unit (AMSU) and the Humidity Sounder for Brazil (HSB), brings climate research and weather prediction into 21st century. From NASA' Aqua spacecraft, the AIRS/AMSU/HSB instruments measure humidity, temperature, cloud properties and the amounts of greenhouse gases. The AIRS also reveals land and sea- surface temperatures. Measurements from these three instruments are analyzed . jointly to filter out the effects of clouds from the IR data in order to derive clear-column air-temperature profiles and surface temperatures with high vertical resolution and accuracy. Together, they constitute an advanced operational sounding data system that have contributed to improve global modeling efforts and numerical weather prediction; enhance studies of the global energy and water cycles, the effects of greenhouse gases, and atmosphere-surface interactions; and facilitate monitoring of climate variations and trends. The high data volume generated by the AIRS/AMSU/HSB instruments and the complexity of its data format (Hierarchical Data Format, HDF) are barriers to AIRS data use. Although many researchers are interested in only a fraction of the data they receive or request, they are forced to run their algorithms on a much larger data set to extract the information of interest. In order to better server its users, the GES DISC/DAAC, provider of long-term archives and distribution services as well science support for the AIRS/AMSU/HSB data products, has developed various tools for performing channels, variables, parameter, spatial and derived products subsetting, resampling and reformatting operations. This presentation mainly describes the web-enabled subsetting services currently available at the GES DISC/DAAC that provide subsetting functions for all the Level 1B and Level 2 data products from the AIRS/AMSU/HSB instruments.

Vicente, Gilberto A.↗

GMI-IPS: Processing & Visualization Software Used in ATom DC-8 Aircraft Studies

NASA's Atmospheric Tomography Mission (ATom) deployed in each of the four seasons during 2016-2018, the DC-8 aircraft in order to establish global-scale datasets intended to improve the representation of chemically reactive gases in global atmospheric chemistry models (ACMs). The Global Modeling Initiative (GMI) executed simulations for each ATom flight using the GMI Chemistry Transport Model (GMI-CTM) to provide species concentrations of chemical gases along the DC-8 flight transects. To solve the problem of translating the GMI-CTM simulation data to the unique spatial resolutions of each ATom flight, the GMI ICARTT Processing Software (GMI-IPS) was developed.The GMI-IPS is written in Python and provides data processing, flight extraction, and visualization support for aircraft research projects using ICARTT format, which is a standard format for airborne instrument data. Additionally, the GMI-IPS interpolates global gridded model data from Hierarchical Data Format (HDF) to ICARTT compatible flight transects. Software classes for instruments and collections provided by the ATom DC-8 aircraft such as MER10, MMS, etc. are derived from a common base class. Other functionality provided by the GMI-IPS are: deriving missing flight entries along a transect, reading ICARTT entries from file, and providing Python data structures for storing flight and model information, and more.The GMI-IPS is GIT source controlled, has approximately 30,000 lines of code, and supports parallelization across data collections. It delivered GMI-CTM data for more than forty distinct DC-8 aircraft flights that took place under ATom. The output ICARTT files adhere to format standard V1.1, and pass the scan utility provided by NASA LaRC Airborne Science Data for Atmospheric Composition. This presentation will include a software and methods overview, and results from ATom, including assessments using the GMI-CTM showing how well observations from ATom flight transects represent a broader region.

Damon, M. R.↗