Search NASA⌕ Search

SEARCH · Search NASA

Results for “Scientific software”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

R&D to Ensure a Scientific Basis for Qualification Tests and Standards (Final Report)

Project return on investment in a photovoltaic (PV) system depends increasingly on maintaining high energy yields, and the system lifetime is a major factor in levelized cost of electricity (LCOE). Thus, the rate of PV deployment and the success of these assets depends upon reliable long-term power generation. The overarching objective of this program is to improve photovoltaic (PV) module reliability via development of tests and standards. Where reliability problems or risk are discovered, we can design tests to ensure that these liabilities don't affect future generations of products. Customers can use these tests to understand which products are susceptible to certain degradation mechanisms, and manufacturers can use the tests to design unwanted characteristics out of their products. The work under this program identifies PV reliability needs, performs characterization that provides scientific understanding of targeted degradation mechanisms, and translates those data into practical and predictive test protocols and standards. Major accomplishments include: A model for polarization-type potential induced degradation (PID-p) was developed and validated against experimental data. NREL is currently leading a new edition of IEC 62804-1 for PID detection. PID-p can cause large losses in current and voltage for some module designs on cloudy days. Finite element modeling (FEM) and experiment was used to determine when cells crack in a module. It was shown that cells in landscape orientation are much more likely to crack than those on portrait orientation. Shortly thereafter, the first products with portrait-oriented cells were introduced. Studies of how to test for light and elevated temperature degradation (LeTID) culminated with the publication of IEC TS 63342. Software to predict the progression of LeTID was developed, validated, and made publicly available. Field validated tests and international standards for durability of PV module coatings abrasion, backsheets, and encapsulants were developed. Examples are IEC 62788-1-1, IEC 62788-2 ED2, IEC TS 62788-7-2, IEC 62788-7-3 ED1, IEC 63209-2. NREL led the development a high-temperature testing technical specification, and published guidelines that enable installers to determine whether higher-temperature testing is needed, simply based on location and mounting configuration. In a number of our case studies, variations in the bills of materials or workmanship have been associated with variations in reliability. These observations emphasize the importance of quality assurance to reliability. A framework for criticality (i.e. Pareto) analysis was developed and published. The framework helps us and other researchers determine what problems should be addressed for reliability research to have the biggest industry impact. NREL continues to participate actively in international standards development and stakeholder engagement activities, including organizing an annual PV Reliability Workshop. These activities are important for ensuring we address issues that are relevant and timely, and that we convey our results to those who may benefit.

14 SOLAR ENERGY↗

Transmission electron microscopy with in-situ ion irradiation: Facilities and community

Whilst there is a clear scientific and technological need for the technical capabilities of transmission electron microscopes with in-situ ion irradiation, it also requires a collaborative community of international researchers to support such facilities in successfully meeting this demand. Instruments of this type serve to provide fundamental understanding of the mechanisms which drive changes in materials important to nuclear fission and fusion energy, the semiconductor industry, quantum information systems, space travel, astronomy, geology and many more applications. As these areas continue to evolve and the instrumentation possibilities expand, the capacity of in-situ ion irradiation facilities must also develop hand-in-hand with the user community to deliver an ever-greater diversity of high-fidelity extreme-environment experimentation. Future directions for the field, such as miniaturization from MEMS/microfluidic devices and advanced controls with ML-based analysis, continuously emerge to advance both the hardware and software which support the coupling of TEMs with ion beams. This review sets out to provide up-to-date insights into the community and advancement of current, and development of future, facilities which have the potential to further unlock access to the nanoscale exploration of coupled extreme environments crucial to many of the important science and engineering challenges we face today.

In-situ irradiation↗

JARVIS-Leaderboard: a large scale benchmark of materials design methods

Abstract Lack of rigorous reproducibility and validation are significant hurdles for scientific development across many fields. Materials science, in particular, encompasses a variety of experimental and theoretical approaches that require careful benchmarking. Leaderboard efforts have been developed previously to mitigate these issues. However, a comprehensive comparison and benchmarking on an integrated platform with multiple data modalities with perfect and defect materials data is still lacking. This work introduces JARVIS-Leaderboard, an open-source and community-driven platform that facilitates benchmarking and enhances reproducibility. The platform allows users to set up benchmarks with custom tasks and enables contributions in the form of dataset, code, and meta-data submissions. We cover the following materials design categories: Artificial Intelligence (AI), Electronic Structure (ES), Force-fields (FF), Quantum Computation (QC), and Experiments (EXP). For AI, we cover several types of input data, including atomic structures, atomistic images, spectra, and text. For ES, we consider multiple ES approaches, software packages, pseudopotentials, materials, and properties, comparing results to experiment. For FF, we compare multiple approaches for material property predictions. For QC, we benchmark Hamiltonian simulations using various quantum algorithms and circuits. Finally, for experiments, we use the inter-laboratory approach to establish benchmarks. There are 1281 contributions to 274 benchmarks using 152 methods with more than 8 million data points, and the leaderboard is continuously expanding. The JARVIS-Leaderboard is available at the website: https://pages.nist.gov/jarvis_leaderboard/

36 MATERIALS SCIENCE↗

From PINNs to PIKANs: recent advances in physics-informed machine learning

Physics-Informed Neural Networks (PINNs) have emerged as a key tool in Scientific Machine Learning since their introduction in 2017, enabling the efficient solution of ordinary and partial differential equations using sparse measurements. Over the past few years, significant advancements have been made in the training and optimization of PINNs, covering aspects such as network architectures, adaptive refinement, domain decomposition, and the use of adaptive weights and activation functions. A notable recent development is the Physics-Informed Kolmogorov-Arnold Networks (PIKANS), which leverage a representation model originally proposed by Kolmogorov in 1957, offering a promising alternative to traditional PINNs. In this review, we provide a comprehensive overview of the latest advancements in PINNs, focusing on improvements in network design, feature expansion, optimization techniques, uncertainty quantification, and theoretical insights. We also survey key applications across a range of fields, including biomedicine, fluid and solid mechanics, geophysics, dynamical systems, heat transfer, chemical engineering, and beyond. Lastly, we review computational frameworks and software tools developed by both academia and industry to support PINN research and applications.

Kolmogorov-Arnold networks↗

PFLOTRAN modeling data and scripts associated with “Refining the Hydrogeologic Framework of a Large River Corridor Model Using Waterborne Transient Electromagnetics”

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the publication “Refining the Hydrogeologic Framework of a Large River Corridor Model Using Waterborne Transient Electromagnetics” submitted to Water Resources Research (Terry et al. 2025). The data package contains the groundwater modeling dataset from PFLOTRAN software. It includes the python script for mesh generation, boundary condition setting, PFLOTRAN input deck formation and postprocessing. It couples groundwater flow and species transport for Hanford Reach river corridor and pipelines the model generation and processing. This model can be used to easily generate the model and analysis for Hanford site. It can also be adjusted to other hydrologic area with ease. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. The data package consists of 6 folders: (1) “data” contains all necessary data as input and intermediate data for processing; (2) “mesh” contains all mesh related files to generate mesh in Hanford Reach river corridor; (3) “model_run” contains the generated script for PFLOTRAN modeling; (4) “notebooks” contains all the Python script to generate the model; (5) “output” contains all the output from the computation; (6) “postprocessing” contains the Python script to generate scientific figure for manuscript. All files are .csv (comma-separated values), .h5 (HDF5 format), .in (input files), .ipynb (Jupyter notebooks), .p (Python pickle), .png (images), .PNG (images), .py (Python scripts), .pyc (Python bytecode), .r (R scripts), .sh (shell scripts), .txt (text files), .vtu (3D mesh/visualization format), .xz (compressed archive), or .zip (compressed archive).

54 ENVIRONMENTAL SCIENCES↗

Recent Developments in DFTB+, a Software Package for Efficient Atomistic Quantum Mechanical Simulations

DFTB+ is a flexible, open-source software package developed by its community, designed for fast and efficient atomistic quantum mechanical simulations. It employs various methods that approximate density functional theory (DFT), such as density functional-based tight binding (DFTB) and the extended tight binding (xTB) approach allowing simulations of large systems over extended time scales with reasonable accuracy, while being significantly faster than traditional ab initio methods. In recent years, several new extensions of the DFTB method have been developed and implemented in the DFTB+ program package in order to improve the accuracy and generality of the available simulation results. In this paper, we review those enhancements, show several use case examples and discuss the strengths and limitations of its features.

36 MATERIALS SCIENCE↗

Science Uses Deployment Operations-Advanced Wireless: Exploring Open Radio Access Network Technologies for Energy Science

Open Radio Access Network is emerging as a solution to the increasing demand for more flexible, cost-effective, and advanced mobile network infrastructures. This evolution is driven by advancements in wireless technologies and the growing complexity of deploying and managing these networks. O-RAN represents a significant shift in wireless technology, building upon the 3rd Generation Partnership Project framework to foster openness, flexibility, and interoperability. By decoupling hardware and software components, Open Radio Access Network enables a multi-vendor ecosystem that encourages innovation and diverse solutions. Open Radio Access Network's potential extends beyond traditional wireless applications, with growing interest in its role in advancing energy systems, particularly in the context of smart grids, microgrids, and the integration of renewable energy sources. While the role of open-wireless technologies in driving energy transformation is increasingly recognized, further exploration is needed. Vendors and utilities are investigating how Open Radio Access Network technologies can optimize energy use cases and improve the performance of 5G and beyond applications. This report outlines efforts under the Science Uses Deployment Operations Advance Wireless project, a collaboration between the National Laboratory of the Rockies' Cybersecurity Research Center, Argonne National Laboratory, Lawrence Berkeley National Laboratory, and the Department of Energy's Energy Science Network research and operations staff. The focus of this project is on due diligence, through testing and evaluation, preparing for the deployment of advanced wireless infrastructure for scientific use cases, with an emphasis on Open Radio Access Network technology, its components, integrations, and its ability to support vertical stack application across the energy sector. Additionally, the report highlights the value cases for utilities, underscoring how adopting open wireless standards can accelerate the evolution of energy systems, foster innovation, and improve the integration of critical energy technologies.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Modeling performance of data collection systems for high-energy physics

Exponential increases in scientific experimental data are outpacing silicon technology progress, necessitating heterogeneous computing systems—particularly those utilizing machine learning (ML)—to meet future scientific computing demands. The growing importance and complexity of heterogeneous computing systems require systematic modeling to understand and predict the effective roles for ML. We present a model that addresses this need by framing the key aspects of data collection pipelines and constraints and combining them with the important vectors of technology that shape alternatives, computing metrics that allow complex alternatives to be compared. For instance, a data collection pipeline may be characterized by parameters such as sensor sampling rates and the overall relevancy of retrieved samples. Alternatives to this pipeline are enabled by development vectors including ML, parallelization, advancing CMOS, and neuromorphic computing. By calculating metrics for each alternative such as overall F1 score, power, hardware cost, and energy expended per relevant sample, our model allows alternative data collection systems to be rigorously compared. We apply this model to the Compact Muon Solenoid experiment and its planned high luminosity-large hadron collider upgrade, evaluating novel technologies for the data acquisition system (DAQ), including ML-based filtering and parallelized software. The results demonstrate that improvements to early DAQ stages significantly reduce resources required later, with a power reduction of 60% and increased relevant data retrieval per unit power (from 0.065 to 0.31 samples/kJ). However, we predict that further advances will be required in order to meet overall power and cost constraints for the DAQ.

Olin-Ammentorp, Wilkie (ORCID:0000000224729862)↗

Nanotomography for Quantitative 3D Particle Reconstruction

Particulates are ubiquitous across fuel cycle operations and carry critical information about particle formation, processing, and potential proliferation-related activities. Traditional analytical techniques, including micro-Raman spectroscopy and standard electron microscopy, are often limited in spatial resolution or dimensionality, particularly when used to examine metallic or submicron-scale features. Understanding particle morphology, phase distribution, and internal porosity is essential for constraining formation conditions, thermodynamic environments, and material transport behavior. In this report, we demonstrate the application of plasma focused ion beam nanotomography to reconstruct micron-scale particulates at nanoscale resolution. Using high-resolution backscattered electron imaging and Avizo software, we obtained 3D reconstructions that enabled quantitative analysis of particle morphology, phase composition, and internal voids. Representative examples include a Ta particle with a large central void and a composite particle with embedded tetrahedral crystalline structures. These reconstructions reveal structural and compositional details that are inaccessible through conventional 2D imaging. The results demonstrate that nanotomography provides both qualitative and quantitative insights into particle formation and behavior. Using nanotomography, porosity and phase distributions can be quantified to inform models of particle density, transport, and solidification conditions. Beyond technical insights, the workflow developed here establishes a transferable capability for analyzing heterogeneous particles and has potential applications in bulk materials studies via x-ray computed tomography or other volumetric imaging modalities. Ongoing efforts are focused on optimizing the workflow to process multiple particles simultaneously, increasing throughput and statistical robustness. Overall, this work illustrates the power of nanotomography as a tool for connecting particulate morphology to formation mechanisms, composition, and transport, thereby strengthening analytical capabilities for nuclear forensics, fuel cycle analysis, and related scientific investigations.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Benchmark for two-dimensional large scale coherent structures in partially magnetized E × B plasmas—community collaboration & lessons learned

Low-temperature plasmas (LTPs) are essential to both fundamental scientific research and critical industrial applications. As in many areas of science, numerical simulations have become a vital tool for uncovering new physical phenomena and guiding technological development. Code benchmarking remains crucial for verifying implementations and evaluating performance. This work continues the Landmark benchmark initiative, a series specifically designed to support the verification of LTP codes. In this study, seventeen simulation codes from a collaborative community of nineteen international institutions modeled a partially magnetized E × B Penning discharge. The emergence of large scale coherent structures, or rotating plasma spokes, endows this configuration with an enormous range of time scales, making it particularly challenging to simulate. The codes showed excellent agreement on the rotation frequency of the spoke as well as key plasma properties, including time-averaged ion density, plasma potential, and electron temperature profiles. Achieving this level of agreement came with challenges, and we share lessons learned on how to conduct future benchmarking campaigns. Comparing code implementations, computational hardware, and simulation runtimes also revealed interesting trends, which are summarized with the aim of guiding future plasma simulation software development.

benchmarking↗

Integrated fluorescence light microscopy-guided cryo-focused ion beam-milling for in situ montage cryo-ET

Cryogenic-electron tomography (cryo-ET) permits the in situ visualization of biological macromolecules at the molecular level. Owing to the variable thickness of cells, tissues and organisms, frozen specimens may need to be thinned by cryo-focused ion beam (FIB) milling to produce thin (<500 nm) cryo-lamellae suitable for cryo-ET. Locating regions of interest remains a challenge because untargeted milling can lead to inadvertent ablation and removal of regions of interest. Correlative light and electron microscopy, combined with cryo-FIB milling, can guide the identification of labeled targets in the cellular milieu. Multiple transfers between cryo-imaging instruments, cumbersome correlation algorithms, limited accuracy and low throughput have hindered the routine adoption of cryo-FIB milling within a multimodal correlative workflow for in situ structural biology. Here, in this study, we present a workflow for 3D correlative cryo-fluorescence light microscopy-FIB-ET that streamlines fluorescence light microscopy-guided FIB milling, improving throughput while preserving both structural and contextual information. The complete integration of hardware and software described here minimizes sample contamination from cross-platform exchanges and greatly enhances the efficiency of 3D targeting in cryo-milling. We then describe procedures for implementing montage parallel array cryo-ET (MPACT), which can be easily adapted to any modern life-science transmission electron microscope. MPACT supports high-throughput cryo-ET acquisitions (10 tilt series in 1.5 h) for structure determination and comprehensive contextual understanding of macromolecules within their native surroundings. A complete session from sample preparation to MPACT data processing takes 5−7 d for an individual experienced in both cryo-EM and cryo-FIB milling.

Yang, Jie E. [Univ. of Wisconsin, Madison, WI (Uni↗

PyOED: An Extensible Suite for Data Assimilation and Model-Constrained Optimal Design of Experiments

This article describes PyOED, a highly extensible scientific package that enables developing and testing model-constrained optimal experimental design (OED) for inverse problems. Specifically, PyOED aims to be a comprehensive Python toolkit for model-constrained OED. The package targets scientists and researchers interested in understanding the details of OED formulations and approaches. It is also meant to enable researchers to experiment with standard and innovative OED technologies with a wide range of test problems (e.g., simulation models). OED, inverse problems (e.g., Bayesian inversion), and data assimilation (DA) are closely related research fields, and their formulations overlap significantly. Thus, PyOED is continuously being expanded with a plethora of Bayesian inversion, DA, and OED methods as well as new scientific simulation models, observation error models, and observation operators. These pieces are added such that they can be permuted to enable testing OED methods in various settings of varying complexities. The PyOED core is completely written in Python and utilizes the inherent object-oriented capabilities; however, the current version of PyOED is meant to be extensible rather than scalable. Specifically, PyOED is developed to “enable rapid development and benchmarking of OED methods with minimal coding effort and to maximize code reutilization.” This article provides a brief description of the PyOED layout and philosophy and provides a set of exemplary test cases and tutorials to demonstrate the potential of the package.

97 MATHEMATICS AND COMPUTING↗

Elastic Bayesian Model Calibration

Functional data are ubiquitous in scientific modeling. For instance, quantities of interest are modeled as functions of time, space, energy, density, etc. Uncertainty quantification methods for computer models with functional response have resulted in tools for emulation, sensitivity analysis, and calibration that are widely used. However, many of these tools do not perform well when the computer model’s parameters control both the amplitude variation of the functional output and its alignment (or phase variation). This paper introduces a framework for Bayesian model calibration when the model responses are misaligned functional data. The approach generates two types of data out of the misaligned functional responses: (1) aligned functions so that the amplitude variation is isolated and (2) warping functions that isolate the phase variation. These two types of data are created for the computer simulation data (both of which may be emulated) and the experimental data. The calibration approach uses both types so that it seeks to match both the amplitude and phase of the experimental data. The framework is careful to respect constraints that arise, especially when modeling phase variation, and is framed in a way that it can be done with readily available calibration software. In conclusion, we demonstrate the techniques on two simulated data examples and on two dynamic material science problems: a strength model calibration using flyer plate experiments and an equation of state model calibration using experiments performed on the Sandia National Laboratories’ Z-machine.

97 MATHEMATICS AND COMPUTING↗

Bridging Cloud and Edge Computing at NREL Using CONNECT: Cloud Optimized Networking for Next-Gen Edge Computing Technologies [Slides]

CONNECT is an innovative on-premise hardware and software solution that integrates edge and cloud computing infrastructure at NREL. Built on the AWS Greengrass middleware and leveraging the MQTT protocol, CONNECT enables real-time data streaming from IoT devices and gateways to both cloud and local services, empowering researchers to rapidly capture, analyze, and act upon edge-generated data while leveraging cloud capabilities. The platform addresses research infrastructure challenges by providing a pre-approved platform which is already configured with the correct networking and cybersecurity baselines thus eliminating procurement delays and enabling on-demand availability. CONNECT's hybrid architecture efficiently manages burstable workloads, allowing research teams to dynamically scale computational capacity, handle peak data loads, and reduce operational bottlenecks. Advanced capabilities include built-in GPU support for executing machine learning models which enables low-latency inference at the edge from models trained in the cloud. This architecture supports real-time analytics and filtering, providing a mechanism to allow only transmitting and processing high-value data. Cloud-based configuration management permits engineers to manage on-premise systems remotely, optimizing operational efficiency. By bridging edge and cloud computing, CONNECT provides NREL researchers with a flexible, scalable platform that accelerates scientific discovery while maintaining robust security and performance standards.

97 MATHEMATICS AND COMPUTING↗

Object Proxy Patterns for Accelerating Distributed Applications

Workflow and serverless frameworks have empowered new approaches to distributed application design by abstracting compute resources. However, their typically limited or one-size-fits-all support for advanced data flow patterns leaves optimization to the application programmer—optimization that becomes more difficult as data become larger. The transparent object proxy, which provides wide-area references that can resolve to data regardless of location, has been demonstrated as an effective low-level building block in such situations. Here we propose three high-level proxy-based programming patterns—distributed futures, streaming, and ownership—that make the power of the proxy pattern usable for more complex and dynamic distributed program structures. We motivate these patterns via careful review of application requirements and describe implementations of each pattern. As a result, we evaluate our implementations through a suite of benchmarks and by applying them in three meaningful scientific applications, in which we demonstrate substantial improvements in runtime, throughput, and memory usage.

Distributed Computing↗

Involving the new generations in Fermilab endeavors

Since 1984 the Italian groups of the Istituto Nazionale di Fisica Nucleare (INFN) and Italian Universities, collaborating with the DOE laboratory of Fermilab (US) have been running a two-month summer training program for Italian university students. While in the first year the program involved only four physics students of the University of Pisa, in the following years it was extended to engineering students. This extension was very successful and the engineering students have been since then extremely well accepted by the Fermilab Technical, Accelerator, and Scientific Computing Division groups. Over the many years of its existence, this program has proven to be the most effective way to engage new students in Fermilab endeavors. Many students have extended their collaboration with Fermilab with their Master’s Thesis and PhD. Since 2004 the program has been supported in part by DOE in the frame of an exchange agreement with INFN. Over its almost 40 years of history, the program has grown in scope and size and has involved more than 550 Italian students from more than 20 Italian Universities, Several Institutes of Research, including ASI and INAF in Italy, and the ISSNAF Foundation in the US, have provided additional financial support. Since the program does not exclude appropriately selected non-Italian students, a handful of students from European and non-European Universities were also accepted over the years. Each intern is supervised by a Fermilab Mentor responsible for performing the training program. Training programs spanned from Tevatron, CMS, Muon (g-2), Mu2e, and Short Baseline Neutrino Experiments and DUNE design and experimental data analysis, development of particle detectors (silicon trackers, calorimeters, drift chambers, neutrino and dark matter detectors), design of electronic and accelerator components, development of infrastructures and software for exascale data handling, research on superconductive elements and on accelerating cavities, and theory of particle accelerators. Since 2010, within an extended program supported by the Italian Space Agency and the Italian National Institute of Astrophysics, a total of 30 students in physics, astrophysics, and engineering have been hosted for two months in the summer at US space science Research Institutes and laboratories. In 2015 the University of Pisa included these programs within its educational programs. Accordingly, Summer School students are enrolled at the University of Pisa for the duration of the internship and are identified and ensured as such. At the end of the internship, the students are required to write summary reports on their achievements. After positive evaluation by a University Examining Board, interns are acknowledged credits for their Diploma Supplement. The program was canceled in 2020 and 2021 due to the pandemic but restarted successfully in 2022. We believe this program can be taken as a model and easily adopted by interested institutions.

99 GENERAL AND MISCELLANEOUS↗

LinkML: an open data modeling framework

Background Scientific research relies on well-structured, standardized data; however, much of it is stored in formats such as free-text lab notebooks, nonstandardized spreadsheets, or data repositories. This lack of structure challenges interoperability, making data integration, validation, and reuse difficult. Findings LinkML (Linked Data Modeling Language) is an open framework that simplifies the process of authoring, validating, and sharing data. LinkML can describe a range of data structures, from flat, list-based models to complex, interrelated, and normalized models that utilize polymorphism and compound inheritance. It offers an approachable syntax that is not tied to any one technical architecture and can be integrated seamlessly with many existing frameworks. The LinkML syntax provides a standard way to describe schemas, classes, and relationships, allowing modelers to build well-defined, stable, and optionally ontology-aligned data structures. Once defined, LinkML schemas may be imported into other LinkML schemas. These key features make LinkML an accessible platform for interdisciplinary collaboration and a reliable way to define and share data semantics. Conclusions LinkML helps reduce heterogeneity, complexity, and the proliferation of single-use data models while simultaneously enabling compliance with FAIR (Findable, Accessible, Interoperable, and Reusable) data standards. LinkML has seen increasing adoption in various fields, including biology, chemistry, biomedicine, microbiome research, finance, electrical engineering, transportation, and commercial software development. In short, LinkML makes implicit models explicitly computable and allows data to be standardized at their origin. LinkML documentation and code are available at https://linkml.io/.

AI-ready data↗

HDF5 in the exascale era: Delivering efficient and scalable parallel I/O for exascale applications

Accurately modeling real-world systems requires scientific applications at exascale to generate massive amounts of data and manage data storage efficiently. However, parallel input and output (I/O) faces challenges due to new application workflows and the state-of-the-art memory, interconnect, and storage architectures considered in exascale designs. The storage hierarchy has expanded with node-local persistent memory, solid-state storage, and traditional disk and tape-based storage, thus requiring efficiency at each layer and much more efficient data movement among these layers. This paper discusses how the ExaHDF5 project improved the I/O performance and data management for exascale architectures by enhancing HDF5, a widely used parallel I/O library. The team developed an Asynchronous I/O Virtual Object Layer (VOL) connector that allowed overlapping I/O with computation. They also created a Cache VOL to complement asynchronous I/O by incorporating fast storage layers, such as burst buffer and node-local storage, into the parallel I/O workflow through caching and staging data. Additionally, the team enabled data aggregation and I/O at the node level by using a Subfiling Virtual File Driver (VFD). To demonstrate superior I/O performance with HDF5 at exascale, the ExaHDF5 team collaborated with several exascale applications. In this paper, we show I/O performance improvements for three applications: Cabana (a particle-based simulation library), EQSIM (a regional earthquake simulation software), and E3SM (a climate system modeling library).

Asynchronous I/Ol↗