Search NASASearch

SEARCH · Search NASA

Results for “HDF5”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Forming Aggregations using Virtual Sharding: Lessons Learned from Simple Scalable Storage (S3)

Data aggregation is the ability to combine separate datasets to form a single new logical dataset provides users with a powerful abstraction. The advantage of an aggregate dataset is that the users are freed from having to understand, and incorporate into their workflow, knowledge about the (ad hoc) organization of the constituent datasets. However, aggregating large numbers of files can be computationally complex with data server systems performing many repetitive operations. As part of the authors work on subsetting data stored on Amazon Web Service (AWS) Simple Storage Service (S3), we developed technology to read portions of otherwise monolithic data files. This enables the formation of virtual shards for user in subsetting data stored in HDF5 (hierarchical data format, version 5) files. This same tool can be used to form aggregations that combine data stored in many HDF5 files when those files are stored on S3. The nature of the virtual sharding and the algorithm that exploits it for subsetting is such that it can also be used for aggregation with the need for many of the repetitive operations required by the per file aggregation techniques. We will present timing information that demonstrates the flexibility of this approach. However, the lessons learned is that while this is a useful result in and of itself, these very same techniques can be applied in other contexts where data are stored in services and on media other than S3. For example, this same technique can be applied to data stored on spinning disk. Pushing the envelope for S3 forced a reexamination of our data access techniques which lead to unexpected positive benefits.

Gallagher, James

New Features of the NEQAIR Radiation Code

The longest-lived code for predicting shock layer radiation, NEQAIR, is now in its 5th decade of service. Substantial changes to the code have been made over the previous decade, the most recent report of which was at the 5th Workshop on Radiation in High Temperature Gases in 2014, for the version referred to as NEQAIR14. This paper will review some of the improvements made to the NEQAIR code since then, which is now at v15.2. Some of these features are discussed briefly below. NEQAIR15 and subsequent versions have enabled parallel evaluation of multiple lines of sight. This is accomplished by utilizing the HDF5 file format and placing multiple lines into a single file, LOS.h5, which is used for both input and output. This approach enables straightforward parallel execution both over the number of lines of sight and the number of points per line. For large problems, runtime reduces linearly with the number of nodes deployed since each line is processed independently by a subset of MPI ranks. Three applications of the multi-line solver are discussed. The first has to do with performing loosely coupled radiation-flowfield solutions. In this case the computed absorption and emission coefficients are used to evaluate the total energy absorbed or emitted at each point, allowing evaluation of the volumetric source term in the flowfield. The second computation is for obtaining heat flux from nonuniform flows, which require integration over spherical co-ordinates. These are of particular interest for evaluating radiation on the vehicle backshell. This 3D option improves the angular integration scheme and allows adaptive line selection that together reduce the number of lines required by about an order of magnitude. The final application is for remote observation, which is essentially the 3D integration problem over a small solid angle. For all three of these computations, data can be stored in the HDF5 file which allows a NEQAIR run to be restarted when it times out, or to add atmospheric absorption or instrument scan functions. An additional level of parallelism is enabled in NEQAIR15.2 using GPU routines. The GPU parallelism has realized up to 8x speed-up when running on a single core but diminishes as CPU parallelism is increased. For running multi-line simulations, it may be easier to reserve a large number of CPU nodes than to obtain the number of GPU nodes required for similar performance. A GUI, known as NEQTPY, allows for reading and creating input files, running NEQAIR, and displaying results. A significant feature of NEQTPY is the ability to perform spectral fits to data. The fits can operate on a single line spectrum (radiance vs. wavelength) or a 3D input file with multiple columns of data. Other new features include improved constants, additional species, more detailed non-Boltzmann modelling, advanced user controls, the ability to read and calculate spectra from HITRAN datafiles, photodissociation and photoionization cross-sections. A “fast” automatic grid option may reduce the size and time of spectral calculations while still maintaining good accuracy for total heat flux.

Brett A Cruden

Utah FORGE: Well 16B(78)-32 Distributed Temperature Sensing Data from April and May 2024

This dataset includes Neubrex Energy Services fiber optic distributed temperature sensing (DTS) data from well 16B(78)-32 during stimulation and circulation, including interaction with well 16A(78)-32, during April and May 2024. The DTS data are stored in HDF5 file format and are accompanied by a PowerPoint report on the study. All times in this dataset are in UTC. Depths are in MD relative to Kelly Bushing Height, and temperatures are in degrees Fahrenheit. All DTS measurements were made using a Yokogawa 3000DTSX Distributed Temperature Sensing Interrogator Unit, with a spatial sampling interval of 3.28 feet and a temporal sampling rate of 129 seconds. The third-party Pressure-Temperature Gauge data should be used with caution after April 20, 2024, as its performance is not considered reliable beyond this date.

15 GEOTHERMAL ENERGY

Utah FORGE: Fiber Optic Cumulative Strain Change and Strain Change Rate Data From Well 16A Stimulation at Well 16B

This dataset includes Rayleigh Frequency Shift (RFS) Distributed Strain Sensing (DSS) cumulative strain change and change rate data. The data was acquired during the stimulation of Utah FORGE Well 16A(78)-32 in April 2024 via fiber installed in Well 16B(78)-32. The fiber optic data was acquired using Neubrex SR7000 RFS DSS Distributed Strain sensing instruments and is saved here in the format of HDF5 files (.h5 extension). The spatial sampling on the full wellbore profiles is 0.20 centimeters. The data is the far field strain change response from a baseline profile made down the 16B well on April 3, 2024, so each strain value represents the strain change or strain change rate at each depth relative to the baseline reference profile. The data arrays for each type share the same dimensions (number of channels and time stamps).

15 GEOTHERMAL ENERGY

Utah FORGE: Well 16B(78)-32 DTS, RFS DSS Strain, and Absolute Strain Circulation Test Fiber Optic Data - August 2024

This dataset contains processed fiber optic measurements collected during the extended cross-well circulation test at the Utah FORGE site in August 2024. The data was acquired from the 16B(78)-32 well using distributed fiber optic sensing (DFOS) technology, including distributed temperature sensing (DTS) on multimode fiber and distributed strain sensing (DSS) measurements on single-mode fiber. The dataset includes Brillouin absolute strain, Rayleigh frequency shift (RFS) DSS strain change, RFS DSS strain change rate, and temperature (DTS) data, all stored in HDF5 format. Time coordinates are provided in UTC, and depth measurements are given in measured depth relative to the rotary kelly bushing (MD RKB) in feet. The spatial sampling on the RFS DSS strain data and the Brillouin Absolute Total Strain data is 20 cm and the spatial sampling on the DTS data is 1m. The dataset is accompanied by a report from Neubrex, which provides further documentation.

15 GEOTHERMAL ENERGY

Incorporating ISO Metadata Using HDF Product Designer

The need to store in HDF5 files increasing amounts of metadata of various complexity is greatly overcoming the capabilities of the Earth science metadata conventions currently in use. Data producers until now did not have much choice but to come up with ad hoc solutions to this challenge. Such solutions, in turn, pose a wide range of issues for data managers, distributors, and, ultimately, data users. The HDF Group is experimenting on a novel approach of using ISO 19115 metadata objects as a catch-all container for all the metadata that cannot be fitted into the current Earth science data conventions. This presentation will showcase how the HDF Product Designer software can be utilized to help data producers include various ISO metadata objects in their products.

metadata

Leveraging the Cloud for HDF1 Software Testing

In this talk we will discuss how we leverage the Cloud for HDF software daily regression testing including testing of the HDF5 parallel library on the Cloud cluster using Orange FS.

CI testing

Collision Tracking in OpenMC: Methods and Applications in Neutron Noise, Neutron Imaging, Time-of-Flight, and Multiplicity Counting

We present the development and application of a collision tracking feature within the OpenMC Monte Carlo particle transport code, designed for diverse applications such as neutron spectroscopy, scatter camera system, neutron noise, and multiplicity counting simulations. This feature enables the tracking of individual particle collisions, with potential applications in nuclear nonproliferation, reactor physics, and nuclear security. Additionally, the feature holds potential for the calibration of neutron detectors, specifically in converting light output into energy deposited within the detectors. The implementation consists of a set of filters—such as reaction type, energy, cell, and material—that constrain the set of collisions that are tracked, extensions to the Python API to enable simple input specification, and support for writing either OpenMC’s native HDF5-based format or the Monte Carlo particle list format. This feature was added to the official OpenMC release in version 0.15.3. In this work, the feature will be applied to showcase scenarios such as time-of-flight simulations, scatter-camera imaging for neutron source localization, neutron-noise analysis to extract integral kinetic parameters such as the prompt decay constant α, and multiplicity counting to estimate the mass of special nuclear materials. Ultimately, this feature aims to expand the application scope of open-source Monte Carlo particle transport codes such as OpenMC.

Monte Carlo code

DaYu: Optimizing Distributed Scientific Workflows by Decoding Dataflow Semantics and Dynamics

The combination of ever-growing scientific datasets and distributed workflow complexity creates I/O performance bottlenecks due to data volume, velocity, and variety. Although the increasing use of descriptive data formats (e.g., HDF5, netCDF) helps organize these datasets, it also creates obscure bottlenecks due to the need to translate high level operations into file addresses and then into low-level I/O operations. To address this challenge, we introduce DaYu, a method and toolset for analyzing (a) semantic relationships between logical datasets and file addresses, (b) how dataset operations translate into I/O, and (c) the combination across entire workflows. DaYu's analysis and visualization enables identification of critical bottlenecks and reasoning about remediation. We describe our methodology and propose optimization guidelines. Evaluation on scientific workflows demonstrates up to 3.7x performance improvements in I/O time for obscure bottlenecks. The time and storage overhead for DaYu's time-ordered data is typically under 0.2% of runtime and 0.25% of data volume, respectively.

Tang, Meng

Data-Driven Protection Software to classify fault locations by protective zone in distribution systems with high PV penetration

The software contains (a) the source codes to generate Point-on-Wave (PoW) transient data for any feeder model in Alternative Transient Program (ATP) format. Codes provide options to change different steady state settings, including the loading condition and PV capacity and transient state setting like faults type, location and initiation time (b) data post-processing source code to converted data from native format to COMTRADE, csv, HDF5 (c) Docker container to train CNN to classify fault locations by protective zone. The container takes dataset and other training parameters (sampling rate, training epochs, batch size etc) as input to train CNN. The container writes back the trained CNN model, training and testing metrics and plots to the local workstation

Ramesh, Meghana

Dirty Word Scanner

SAND2025-09142O Dirty Word Scanner helps prevent the accidental inclusion of sensitive terms by scaning files in repositories to catch "dirty words" before they are committed. While there are existing solutions focused on passwords and API keys, this tool offers additional features tailored to specific security needs. It will function as a standalone tool, incorporating advanced capabilities from similar tools to provide a comprehensive solution. This tool can unpack HDF5 files and examine their contents. It can display image, audio, and visual files to the user and request a manual determination of whether they are safe. It can also detect arbitrary binary files and ask the user to verify that they're safe. The tool enables sophisticated whitelisting of strings and regular expressions for cases where a term is sensitive in certain contexts but not in others. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Gates, Jason [Sandia National Lab. (SNL-CA), Liver

Knowledge Oriented Graph Unified Transformer (KOGUT) v0.1

KOGUT — Knowledge Oriented Graph Unified Transformer KOGUT implements the Relational Graph Transformer (RelGT) architecture for knowledge graph link prediction in biological domains, with a primary focus on microbial growth media prediction. While the original RelGT (arXiv:2505.10960) targets relational tables, time series, and multi-table databases, KOGUT adapts this architecture for heterogeneous biological knowledge graphs, providing first-in-class AI predictive models for microbial cultivation. Key Adaptations Beyond Original RelGT: - Knowledge Graph Focus: Applied to biological KGs with semantic node types (taxa, chemicals, media, phenotypes, environments) versus generic relational database tables, trained on the KG-Microbe knowledge graph (1.3M entities, 2.9M edges, 24 relation types). - Multimodal Node Encoding: Integrates node labels, categories, descriptions, and synonyms from KG metadata through learned embedding layers—adapting relational column features to graph node attributes with textual semantics. - Extended K-Hop Subgraph Strategy: Optimized neighborhood sampling (3-hop default, configurable up to 200 nodes) tuned for sparse biological networks, building on the original local-global attention framework with biological relation preservation. - Biolink Predicate Preservation: Type-specific transformations for 24 biological edge semantics (occurs_in, consumes, produces, has_phenotype, subclass_of) beyond standard relational foreign keys, enabling multi-relation link prediction. - Inductive Learning Support: Enables zero-shot predictions for novel taxa through feature-based embeddings (temperature, oxygen requirements, gram stain, cell shape), extending the original transductive relational benchmark scope to uncultured microorganisms. CheapSOTA Performance Optimizations (This Distribution): - VQ-EMA Centroid Attention: Vector quantization with exponential moving average for improved global context modeling (+5-10% MRR improvement). - HDF5 Precomputed Data Loading: One-time preprocessing of k-hop subgraphs to eliminate redundant graph traversals (2-5× training speedup). - Distributed Data Parallel Training: Multi-GPU support for scaling to larger knowledge graphs (tested on 4× NVIDIA A100 GPUs at NERSC Perlmutter). - Mixed Precision Training: Automatic mixed precision (AMP) for memory efficiency and faster training. Advantages Over Standard Knowledge Graph Embedding Models: Combines RelGT's proven multi-element tokenization (features, type, hop, structure) with graph-native biological representations, enabling interpretable link prediction across heterogeneous entities that standard embedding models (TransE, RotatE, ComplEx) and table-based transformers cannot directly model. Achieves near-perfect performance on microbial growth media prediction (MRR: 0.9966, Precision@1: 0.9932, Hit@10: 1.0000) while maintaining explainability through attention-based reasoning over biological pathways. Training Data: - KG-Microbe merged knowledge graph: 1,379,337 nodes, 2,960,472 edges - 24 biological relation types including taxonomic hierarchies, metabolic interactions, phenotype associations, and environmental relationships - Primary prediction task: Growth media suitability for microbial taxa (biolink:occurs_in, 50K edges) - Multi-relation capability: Predicts links for any of the 24 relation types, including chemical consumption/production, phenotype associations, and taxonomic classification Citation: Original RelGT Architecture: Dwivedi et al., "Relational Graph Transformer", arXiv:2505.10960, 2025 KOGUT Implementation: Knowledge Oriented Graph Unified Transformer for Microbial Growth Media Prediction Developed at Lawrence Berkeley National Laboratory (LBNL) Trained on NERSC Perlmutter supercomputer

Joachimiak, Marcin [Lawrence Berkeley National Lab

Scientific Core Library Stack (SCLS) v2026

SCLS (Scientific Core Library Stack) is an opinionated build and packaging system for scientific computing libraries developed at Lawrence Berkeley National Laboratory. It produces a coherent, reproducible stack of numerical libraries — including BLAS/LAPACK, MPI, sparse direct and iterative solvers, graph partitioners, and parallel I/O libraries (e.g., PETSc, SLEPc, HDF5, NetCDF, MUMPS, OpenBLAS) — that work together without manual repair by downstream scientific software. From a single recipe-and-flavor model, SCLS produces native RPM packages for RHEL-family Linux, DEB packages for Debian/Ubuntu, direct Unix-style prefix installs for HPC and locked-down environments, and native macOS builds. Multiple build "flavors" (e.g., GCC+OpenBLAS, GCC+MKL, Intel+MKL, debug) coexist in distinct prefixes on the same host. Compared to general-purpose meta-build frameworks, SCLS is deliberately curated rather than infinitely configurable. It enforces deterministic, audit-friendly behavior: explicit build dependencies, no silent feature autodetection, a clear open-source license policy, and rpath-based runtime linkage so installs integrate cleanly with standard package-manager workflows.

Messe, Christian [Lawrence Berkeley National Labor

One Million Open-source Cislunar Orbits.

The dataset contains one million integrated cislunar orbit trajectories, for a time span of up to six years. The data was generated on LLNL HPC systems and is saved in the form of both HDF5 and CSVs.

Yeager, T

Dust Survival in Galactic Winds

This repository contains three-dimensional volumetric data from an Eulerian hydrodynamical simulation (conducted on a uniform Cartesian grid) generated by the Cholla hydrodynamics code. The datasets contain snapshots (full-grid, projections, and slices) in the HDF5 format of a multi-phase medium in which a hot, diffuse, dust-free background wind accelerates a cool, dense cloud of gas and dust. This scenario is intended to represent a supernova-driven galactic outflow, in which hot supernova winds are thought to accelerate cool interstellar medium material out of the galactic disk into the surrounding circumgalactic medium. There are three separate datasets for simulations corresponding to three cloud evolutionary scenarios: long-term cloud survival (surv), marginal cloud survival (disr), and cloud destruction (dest). Projection and slice images of the simulations are also included in this repository.

79 ASTRONOMY AND ASTROPHYSICS

Cholla Galactic OutfLow Simulations (CGOLS)

These datasets contain full hydro-field snapshots from the galactic outflow simulations in the CGOLS suite, models I-V. The datasets were generated using the Cholla hydrodynamics code (https://github.com/cholla-hydro/cholla); descriptions of the models are in the associated publications (Schneider & Robertson 2018, ApJ; Schneider et al. 2018, ApJ; Schneider et al. 2020, ApJ; and Schneider & Mao, 2024, ApJ). Each hdf5 dataset is numbered according to the simulation time of the snapshot, in Myr. Fields include density, x momentum, y momentum, z momentum, total energy, and thermal energy (for models I - III), as well as a passive scalar field (models IV and V). 2 dimensional density and temperature projections, as well as slices along each midplane are also included if they exist.

79 ASTRONOMY AND ASTROPHYSICS

Simulated Microstructures for Laser Powder Bed Fusion Additive Manufacturing Using Myna, AdditiveFOAM, and ExaCA

This dataset provides sample datasets containing voxelized, three-dimensional representations of simulated grain structures and crystallographic orientations that can result from laser powder bed fusion additive manufacturing. The six microstructure files each contain approximately 1 cubic millimeter of material (1 mm x 1 mm cross-section over 26 simulated layers of deposition). Some of the microstructures have columnar grains that extend across nearly the entire simulation domain, while others have more equiaxed or truncated columnar grains. The process conditions to generate these microstructures were from the Peregrine v2023-10 dataset (10.13139/ORNLNCCS/2008021). The codes used are publicly available and released under open-source licenses. Myna (https://github.com/ORNL-MDF/Myna) was used for configuration of the cases from the Peregrine v2023-10 HDF5 dataset and to run the simulation workflow. AdditiveFOAM (https://github.com/ORNL/AdditiveFOAM) was used to simulate the melt pool and generate solidification conditions. And ExaCA (https://github.com/LLNL/ExaCA ) was used to simulate the three-dimensional microstructures.

36 MATERIALS SCIENCE

Spin-phonon coupling in AFM transition-metal mono-oxide

Time-of-flight INS measurements were performed on single crystal NiO with the Wide Angular Range Chopper Spectrometer (ARCS) at the Spallation Neutron Source. Experiments were performed on NiO single crystal mounted in an aluminum can and cooled using a closed-cycle helium refrigerator. Measurements were conducted at T = 100 K and 650 K, with the [HHL] scattering plane aligned horizontally. A Fermi chopper with slit spacing of 1.52mm, spinning at 300 Hz, was used to select an incident neutron energy of 100 meV. All datasets were normalized to a vanadium standard to correct for detector efficiency and solid angle coverage. The data sets include the .nxs files, the generated .hdf5 files (for use with Phonon Explorer), and Python scripts used to create them.

36 MATERIALS SCIENCE