Search NASA⌕ Search

SEARCH · Search NASA

Results for “hierarchical data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Closing the Gap between FAIR Data Repositories and Hierarchical Data Formats

Many in the scientific community, particularly in publicly funded research, are pushing to adhere to more accessible data standards to maximize the findability, accessibility, interoperability, and reusability (FAIR) of scientific data, especially with the growing prevalence of machine learning augmented research. Online FAIR data repositories, such as the Open Science Framework (OSF), help facilitate the adoption of these standards by providing frameworks for storage, access, search, APIs, and other features that create organized hubs of scientific data. However, the wider acceptance of such repositories is hindered by the lack of support of hierarchical data formats, such as Technical Data Management Streaming (TDMS) and Hierarchical Data Format 5 (HDF5), that many researchers rely on to organize their datasets. Various tools and strategies should be used to allow hierarchical data formats, FAIR data repositories, and scientific organizations to work more seamlessly together. A pilot project at Los Alamos National Laboratory (LANL) addresses the disconnect between them by integrating the OSF FAIR data repository with hierarchical data renderers, extending support for additional file types in their framework. The multifaceted interactive renderer displays a tree of metadata alongside a table and plot of the data channels in the file. This allows users to quickly and efficiently load large and complex data files directly in the OSF webapp. Users who are browsing files can quickly and intuitively see the files in the way they or their colleagues structured the hierarchical form and immediately grasp their contents. This solution helps bridge the gap between hierarchical data storage techniques and FAIR data repositories, making both of them more viable options for scientific institutions like LANL which have been put off by the lack of integration between them.

97 MATHEMATICS AND COMPUTING↗

Hierarchical Data Format for Nuclear Data Sensitivities

The SCALE code system includes capabilities for sensitivity and uncertainty (S/U) analysis as part of its TSUNAMI code suite. The sensitivity of a quantity of interest (for example, an application’s $k_{eff}$) to nuclear data is stored as a profile in a text-based file, which is known as a sensitivity data file (SDF). The sensitivity profile can be used to calculate uncertainties, correlation coefficients, and similarity indices. One of the goals of the present work was to seek general performance improvements in the TSUNAMI code suite, starting with the TSUNAMI-IP code for calculating similarity indices. Through profiling, it was found that reading the text-based sensitivity files was a performance bottleneck in the TSUNAMI-IP code. In a typical TSUNAMI-IP calculation, an application might be compared to thousands of benchmarks, thus requiring the reading of thousands of SDFs. Reading of binary-based data is generally faster than reading text-based data. Hierarchical Data Format 5 (HDF5) is a binary-based format that also benefits from being portable, and it can be inspected with nonproprietary tools. This paper describes an HDF5-based file format that has been introduced for SDFs.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Hierarchical Data Format for Nuclear Data Sensitivities [Slides]

An HDF5-based file format was introduced for the sensitivity data calculated by TSUNAMI. The format was defined to collect the sensitivity coefficients into hyperslabs, which optimizes file reading time and therefore improves the time-to-solution for applications. In future work, this format will be extended to store sensitivity data for depletion calculations.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Spatiotemporal Modeling of Real World Backsheets Field Survey Data: Hierarchical (Multilevel) Generalized Additive Models: Preprint

Assessing photovoltaic module backsheet durability is critical to increasing module lifetime. Lab based accelerating testing has recently failed to predict large scale failures of widely adopted polymeric materials. Field surveyed data is critical to assess the performance of component lifetime. Using a documented field survey protocol, 13 field surveys where conducted. Each measurement is encoded with it's spatial location in respect to the other modules. By combining field survey data on degradation predictors with real time satellite weather data, data-driven predictive models of backsheet degradation were trained. LOESS models were constructed to investigate the spatial dependence of measurements. It was found that micro-climatic effects like treelines, ground surface changes, and elevation changes effected the magnitude and variance of the measurements. A GAM model was created to predict the value of degradation based on measured predictors. The model includes variables on the climate of the system and the location of each measurement in the PV mounting structure. The model performed well with an adj:R2 of 0:95 for yellowness index prediction. The model was cross-validated using k-folds.

backsheet↗

Spatiotemporal Modeling of Real World Backsheets Field Survey Data: Hierarchical (Multilevel) Generalized Additive Models

Assessing photovoltaic module backsheet durability is critical to increasing module lifetime. Laboratory-based accelerating testing has recently failed to predict large scale failures of widely adopted polymeric materials. Additionally, there is a growing concern on characterizing the non-uniformity of field exposure. Therefore, data from field surveys are critical to assess the performance of component lifetimes. Using a documented field survey protocol, 19 field surveys were conducted. The focus of this survey strategy is to investigate spatial continuity in degradation modes. By combining field survey data with real-time satellite weather data, stressor / response models have been trained. Generalized additive Models (GAM) model was created to predict the value of degradation based on measured predictors. Two different GAM constructions were testing using different implementations of basis splines. The model includes variables on the environmental stressors of the system and the location of each measurement in the PV mounting structure. The incorporation of hierarchical structure into the models allowed for material specific degradation rates, while maintaining the assumption of a global trend. The model performed well with an adjusted R2 of 0.975 for yellowness index prediction.

backsheet↗

Accelerating Multigrid-based Hierarchical Scientific Data Refactoring on GPUs

Rapid growth in scientific data and a widening gap between computational speed and I/O bandwidth make it increasingly infeasible to store and share all data produced by scientific simulations. Instead, we need methods for reducing data volumes: ideally, methods that can scale data volumes adaptively so as to enable negotiation of performance and fidelity tradeoffs in different situations. Multigrid-based hierarchical data representations hold promise as a solution to this problem, allowing for flexible conversion between different fidelities so that, for example, data can be created at high fidelity and then transferred or stored at lower fidelity via logically simple and mathematically sound operations. However, the effective use of such representations has been hindered until now by the relatively high costs of creating, accessing, reducing, and otherwise operating on such representations. We describe here highly optimized data refactoring kernels for GPU accelerators that enable efficient creation and manipulation of data in multigrid-based hierarchical forms. We demonstrate that our optimized design can achieve up to 250 TB/s aggregated data refactoring throughput—83% of theoretical peak—on 1024 nodes of the Summit supercomputer. We showcase our optimized design by applying it to a large-scale scientific visualization workflow and the MGARD lossy compression software.

Chen, Jieyang↗

Domain-Specific Type-Safe APIs for Hierarchical Scientific Data with Modern C++

General-purpose library application programming interfaces (APIs) for self-describing hierarchical scientific data storage, such as the HDF5 and NetCDF libraries, are traditionally of runtime nature. Runtime errors for entry existence and data types are typically caught later in the development process of higher-level application-specific APIs. In this paper, we propose exploiting modern C++ metaprogramming features to add compile-time type-safety to improve the interaction with a well-defined metadata-rich scientific schema in domain-specific hierarchical datasets. We tackle two aspects of common use: (i) direct data access, (ii) flexible “in-memory” index models for efficient search and data processing. The proposed APIs use C++17’s template type auto deduction features, C++11’s enum class for type-safety and C-style preprocessor macros for generative templated code. We showcase the pros and cons of our initial work on the standard NeXus schema used for annotating and storing experimental neutron scattering data at several facilities around the world on top of HDF5. Extendable compile-time type-safe APIs are a desirable feature that could be indexed by any modern integrated development environment (IDE). Hence, such APIs can help ease the learning curve for domain scientists using a less error-prone software interaction to enhance the findability of their data without resorting to a domain-specific language (DSL).

Godoy, William↗

Hierarchical Data-Driven Protection for Microgrid with 100% Renewable Penetration: Preprint

The accurate detection and isolation of faults is critical for the reliable operation of microgrids (MGs). Traditional protection approaches are even more challenged for 100% renewable MGs because inverter-based resources (IBRs) are the only sources for fault current which are usually low and unpredictable/non-uniform. This calls for new protection scheme that can identify IBR fault responses and detect faults in MGs. Data-driven based protection can learn the pattern of IBR fault responses and make the correct decision to identify faults. Therefore, this paper presents a data-driven approach for fault localization in island MGs. The approach builds a training dataset of comprehensive fault scenarios that can be used to learn fault characteristics from processed measurements. The localization task is modeled as a binary classification problem at each relay, which simplifies the learning process. Then, a hierarchical decision mechanism is used to identify the fault location. The proposed approach is assessed using an exemplary MG with several grid-forming (GFM) and grid-following (GFL) inverters, where accurate estimation of fault location is achieved. The data-driven based protection approach developed in this paper provides a generic framework and useful guidance for power system protection engineers to achieve reliable protection for MGs with 100% renewables.

artificial intelligence↗

Solar PV, Wind Generation, and Load Forecasting Dataset for ERCOT 2018: Performance-Based Energy Resource Feedback, Optimization, and Risk Management (P.E.R.F.O.R.M.)

This report describes the Advanced Research Projects Agency-Energy Performance-Based Energy Resource Feedback, Optimization, and Risk Management (PERFORM) Electric Reliability Council of Texas (ERCOT) dataset consisting of load, solar, and wind deterministic and probabilistic forecasts at three timescales. This dataset consists of 1 year of time-coincident load, wind, and solar actuals and probabilistic forecasts for a region similar to ERCOT. All the data are stored in Hierarchical Data Format 5 (HDF5) files and have been uploaded to an Amazon Web Services repository. The ERCOT data set has 2 years (2017, 2018) of actuals and 1 year (2018) of probabilistic forecasts. These data are provided at various spatial (i.e., site-level, zone-level, and system-level) and temporal scales (i.e., day-ahead, intraday, and intra-hour). Specifically, data are provided for 125 existing wind sites, 22 existing solar sites, 139 proposed wind sites, and 204 proposed solar sites.

14 SOLAR ENERGY↗

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository

The Department of Energy's (DOE's) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and helping users access data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data. This paper provides an update on recent improvements made to the GDR's data lakes and automated data pipelines, including: (1) streamlining the data lake intake process, (2) better educating users on the process and requirements through a new data lakes page, (3) adding data lake direct access links to GDR data lake submission pages, (4) implementing a DAS data pipeline to convert DAS data uploaded in SEG-Y format to a standardized hierarchical data format v5 (HDF5), (5) extending this pipeline to encompass data in the GDR data lake, (6) adding metadata requirements for geospatial data, (7) making user interface/user experience (UX) enhancements to the data pipelines' documentation pages, and (8) improving the GDR's data standards and pipelines pages to better guide users in ensuring that their data is standardized by the GDR's automated data pipelines. 2024 Geothermal Resources Council. All rights reserved.

accessibility↗

CHESS 2025: Spectrometer orthorectified at-sensor radiance from NEON AOP imaging spectroscopy surveys

This dataset provides Level 1 (L1) orthorectified at-sensor radiance derived from measurements collected by the Imaging Spectrometer-1 (NIS-1) onboard the NEON (National Ecological Observatory Network) Airborne Observation Platform (AOP) for the 2025 Colorado Headwaters Ecological Spectroscopy Study (CHESS). NIS-1 captures light reflected from the Earth’s surface in 426 discrete wavelength bands as raw digital numbers (DNs; Level 0). These data are then calibrated to physical units (uW/cm²·sr·nm) following the processing steps described in the NEON Imaging Spectrometer Level 1B Calibrated Radiance Algorithm Theoretical Basis Document (ATBD; Gallery 2022). The data delivered here are the primary inputs for the surface reflectance product in “Custom surface reflectance, shade masks, and equivalent water thickness maps for the Colorado Headwaters Ecological Spectroscopy Study” (Carroll et al. 2026). For intertemporal comparison, the radiance data here are most directly relatable to the v2 radiance data in “NEON AOP Imaging Spectroscopy Survey of Upper East River Colorado Watersheds: Raw-Space Radiance and Observational Variable Dataset” (Goulden et al. 2018), to which the same processing methodology was applied. Together, the radiance and reflectance data enable users to exploit the unique reflection signatures of different surface objects for land cover classification, foliar trait mapping, plant vigor assessment, water content estimation, trace-element identification, and other scientific applications. The data were acquired over three study domains in the Upper Gunnison river basin: the upper East River watershed (CRBU); Almont Triangle and Taylor Canyon (ALMO); and Upper Taylor River watershed (UPTA) between 2025-06-13 and 2025-07-15. Within each domain, data are delivered by flightline as orthorectified and calibrated hyperspectral rasters in Hierarchical Data Format version 5 (HDF5) format, with radiance values provided in uW/cm²·sr·nm on a fixed, uniform Universal Transverse Mercator (UTM) grid at 1 meter spatial resolution. The radiance rasters include all 426 NIS-1 spectral bands, along with associated quality-assurance (QA) and diagnostic and ancillary layers needed for atmospheric correction workflows. Orthorectified radiance is produced from pushbroom spectrometer observations by applying NEON’s radiometric calibration (including bad pixel masking, dark subtract, dark pedestal shift correction, electronic panel ghost correction, grating ghost correction, deblur correction and flat-fielding) and spectral calibration (using spectral response function band centers and full-width at half-maximum intensity), followed by geolocation and regridding to the fixed grid. CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgement: Field and remote-sensing data acquisition was performed under a grant from the National Aeronautics and Space Administration (80NSSC24K1005). This work was also supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS Campaigns↗

Traffic safety analysis and model updating for freeways using Bayesian method

Freeway crash prediction models are the basic of traffic safety research, yet crash occurrence and the influencing factors change over time. In order to make sure the implemented safety models fit the current traffic environment, this study conducts a comparative analysis of 2017 and 2020 datasets collected from freeways in Suzhou, China. Herein, considering the spatial correlation among analysis units and the hierarchical data structure, a Bayesian conditional autoregressive negative binomial (CAR-NB) model and a Bayesian hierarchical CAR-NB (HCAR-NB) model were used to explore the safety influencing factors, and a traditional NB model was developed for further comparison. To update the HCAR-NB model from 2017 to 2020, Bayesian inference with informative priors was used to improve its goodness of fit and efficiency. Preliminary results showed that 1) the HCAR-NB model outperformed the NB model and CAR-NB model in prediction accuracy, and 2) the number of crashes was significantly correlated with average speed, speed variance, road segment length, number of lanes, and presence of ramps. The potential for safety improvement (PSI) method was applied to the modeling results to identify hotspots for the two years. The results confirmed that the hotspots spatiotemporally shifted among the freeways. The proposed crash prediction model and updating method are expected to assist implementation of informed countermeasures for freeway safety improvement.

97 MATHEMATICS AND COMPUTING↗

Counter Unmanned Aircraft System Metrics Tool

SAND2023-05387O The Counter Unmanned Aircraft System (CUAS) Metrics Tool consists of three applications: • The Data Collection Tool is designed to be used during testing and captures details that can be imported into the Figure Creator application. • The Unmanned Aircraft System (UAS) Log Converter converts different UAS logs into a standardized hierarchical data format with a timeline of geospatial data. • The Figure Creator develops relevant images for the test event. The application includes logic to ingest and correlate data from the Data Collection Tool and the UAS Log Converter to produce the chosen figures. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Mayle, Ashley↗

Model Data Archive Associated with Manuscript "Fire-altered Carbon Pools Create Disturbance Memory in Stream Dissolved Organic Carbon"

This data package supports the publication “Fire-altered Carbon Pools Create Disturbance Memory in Stream Dissolved Organic Carbon” by Li et al. (2026). The package contains processed model inputs, configuration files, restart files, simulation outputs, scripts, and visualization products used to evaluate post-fire dissolved organic carbon (DOC) dynamics in the Naches River Watershed, Washington, USA, following the 2021 Schneider Springs Fire. The modeling workflow couples ELM-BGC, the biogeochemistry-enabled Energy Exascale Earth System Model Land Model; ATS, the Advanced Terrestrial Simulator for integrated surface-subsurface hydrology; and PFLOTRAN, a reactive transport model for multicomponent aqueous geochemistry. Together, these models simulate how wildfire-induced changes in vegetation, litter, coarse woody debris, and soil organic matter influence DOC production, transport, and reaction from burned hillslopes to stream networks. The archive includes preprocessed meteorological, geospatial, hydrologic, and biogeochemical forcing data; ELM-BGC-derived DOC source terms; ATS mesh files; PFLOTRAN reactive-transport inputs; model configuration files; spin-up and transient restart files; watershed-scale diagnostic outputs; stream concentration time series; and figures or visualization files used to inspect and reproduce key results. File types include Hierarchical Data Format 5 (HDF5) files for gridded forcing and model-coupling data, model input and configuration files for ELM-BGC, ATS, and PFLOTRAN, restart and simulation-output files generated by the modeling workflow, tabular or time-series diagnostic outputs, scripts for post-processing and figure generation, and image or visualization products associated with the manuscript. Use of the package depends on the intended task. Re-running the simulations requires the relevant modeling software, including ELM-BGC, ATS, and PFLOTRAN as ATS's geochemical engine. Inspecting outputs and reproducing figures requires Python with scientific plotting libraries such as Matplotlib, and three-dimensional model outputs may be viewed with ParaView. Geographic information system files or maps may be inspected with ArcGIS Pro or comparable GIS software. The data package is intended to enable traceability, reuse, and partial reproduction of the coupled land-to-watershed hydro-biogeochemical modeling workflow used to test how wildfire disturbance affects terrestrial carbon pools and downstream DOC dynamics.

ATS↗

h5bench: A unified benchmark suite for evaluating HDF5 I/O performance on pre‐exascale platforms

Summary Parallel I/O is a critical technique for moving data between compute and storage subsystems of supercomputers. With massive amounts of data produced or consumed by compute nodes, high‐performant parallel I/O is essential. I/O benchmarks play an important role in this process; however, there is a scarcity of I/O benchmarks representative of current workloads on HPC systems. Toward creating representative I/O kernels from real‐world applications, we have created h5bench , a set of I/O kernels that exercise hierarchical data format version 5 (HDF5) I/O on parallel file systems in numerous dimensions. Our focus on HDF5 is due to the parallel I/O library's heavy usage in various scientific applications running on supercomputing systems. The various tests benchmarked in the h5bench suite include I/O operations (read and write), data locality (arrays of basic data types and arrays of structures), array dimensionality (one‐dimensional arrays, two‐dimensional meshes, three‐dimensional cubes), I/O modes (synchronous and asynchronous). In this paper, we present the observed performance of h5bench executed along several of these dimensions on existing supercomputers (Cori and Summit) and pre‐exascale platforms (Perlmutter, Theta, and Polaris). h5bench measurements can be used to identify performance bottlenecks and their root causes and evaluate I/O optimizations. As the I/O patterns of h5bench are diverse and capture the I/O behaviors of various HPC applications, this study will be helpful to the broader supercomputing and I/O community.

97 MATHEMATICS AND COMPUTING↗

Data shuffling with hierarchical tuple spaces

Methods and systems for shuffling data are described. A processor may generate pair data from source data. The processor may insert the pair data into local tuple spaces. In response to a request for a particular key, the processor may determine a presence of the requested key in a global tuple space. The processor may, in response to a presence of the requested key in the global tuple space, update the global tuple space. The update may be based on the pair data among the local tuple spaces including the existing key. The processor may, in response to an absence of the requested key in the global tuple space, insert pair data including the missing key from the local tuple spaces into the global tuple space. The processor may fetch the requested pair data, and may shuffle the fetched data to generate a dataset.

Andrade Costa, Carlos Henrique↗

Accelerating Random Forest Classification on GPU and FPGA

Random Forests (RFs) are a commonly used machine learning method for classification and regression tasks spanning a variety of application domains, including bioinformatics, business analytics, and software optimization. While prior work has focused primarily on improving performance of the training of RFs, many applications, such as malware identification, cancer prediction, and banking fraud detection, require fast RF classification. In this work, we accelerate RF classification on GPU and FPGA. In order to provide efficient support for large datasets, we propose a hierarchical memory layout suitable to the GPU/FPGA memory hierarchy. We design three RF classification code variants based on that layout, and we investigate GPU- and FPGA-specific considerations for these kernels. Our experimental evaluation, performed on an Nvidia Xp GPU and on a Xilinx Alveo U250 FPGA accelerator card using publicly available datasets on the scale of millions of samples and tens of features, covers various aspects. First, we evaluate the performance benefits of our hierarchical data structure over the standard compressed sparse row (CSR) format. Second, we compare our GPU implementation with cuML, a machine learning library targeting Nvidia GPUs. Third, we explore the performance/accuracy tradeoff resulting from the use of different tree depths in the RF. Finally, we perform a comparative performance analysis of our GPU and FPGA implementations. Our evaluation shows that for high accuracy targets, our GPU implementation yields 5-9x speedup over CSR, and up to a 2x speedup over cuML.

FPGA, Xilinx FPGA, GPU, Random Forest classificati↗