Search NASASearch

SEARCH · Search NASA

Results for “asynchronous processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Myna: Connecting powder bed fusion build data to simulation tools for digital twin applications

Additive manufacturing (AM), as a digital process, can generate a detailed digital thread linking a part’s design and manufacturing to its operational performance. As AM systems advance, an increasing amount of process data is stored in manufacturing databases. In principle, this data can be utilized by simulation-based digital twin approaches, such as real-time process control and asynchronous post-processing guidance. However, few tools currently exist for systematically integrating digital thread data with computational tools. Here, in this study, we propose a software package, called Myna, for connecting data from powder bed fusion processes to simulation tools. The utility of such a platform is demonstrated using build data from the Oak Ridge National Laboratory Manufacturing Demonstration Facility “Peregrine v2023-10” public dataset to automatically configure and run 54 semi-analytical 3DThesis melt pool simulations, 78 numerical Additive FOAM melt pool simulations, and 3 ExaCA microstructure simulations. The simulated, spatially registered microstructures are then compared directly with electron backscatter diffraction characterization of the corresponding as-built part locations. The resulting simulated microstructure showed variation as a function of process parameters, particularly stripe width; however, the experimental data had little variation between the microstructure texture and grain size resulting from different processing conditions. Analysis of the discrepancies suggest that it is possible a two-phase ferritic-austenitic solidification model is needed to accurately predict grain size and texture for certain stainless steel 316L feedstock compositions under powder bed fusion conditions, providing direction for future research. As illustrated here, due to the number and complexity of the simulations involved in AM process-structure–property predictions, automated methods to connect process data and simulations will remain necessary tools for testing hypotheses and implementing digital twin applications.

Knapp, Gerald L. [Oak Ridge National Laboratory (O

Benchmark Tracking System for Performance Monitoring

Benchmarking is essential for high-performance software development, particularly for monitoring performance across code iterations. This project focused on enhancing the benchmarking process for Lamellar, an asynchronous runtime for High-Performance Computing (HPC) systems developed at Pacific Northwest National Laboratory. Prior to this work, benchmark results were difficult to track and compare across code versions, presenting significant challenges in identifying performance regressions and long-term trends. The primary objective was to establish a systematic, reproducible approach for measuring performance and detecting regressions following code commits. Our methodology involved three key components: standardizing benchmark outputs, implementing data versioning, and developing analysis tools. We standardized the benchmark output format to JSON Line records containing specific fields (execution time, hardware specifications, and environmental variables). To address data management challenges, we evaluated several options and eventually chose a git repository dedicated to benchmark data. We developed a suite of Python tools that processed benchmark results, enriched them with metadata, and facilitated search in the repository. The resulting system enables more efficient filtering and comparison of performance metrics across commit histories, hardware configurations, and benchmark variants through a unified query interface. Our implementation reduces computational overhead by first checking for existing results through configuration matching before initiating new benchmark runs, thereby conserving resources. The system has been validated by Lamellar developers. It organizes results by benchmark type and build configurations for efficient retrieval. Future developments include a planned Large Language Model interface for predicting benchmark performance, incorporating the criterion package for statistical analysis, which will enable automated detection of statistically significant performance changes, and integration with continuous integration pipelines. Despite these enhancements being reserved for future work, this project has successfully provided the Lamellar development team with a framework for maintaining consistent performance standards and identifying optimization opportunities across workloads and hardware environments.

97 MATHEMATICS AND COMPUTING

rmon (Resource monitor) [SWR-24-128]

The resource monitor application provides monitoring, collection, and visualization of resource utilization in compute nodes. This package contains utilities to monitor system resource utilization (CPU, memory, disk, network). Here are the ways you can use it: -Monitor resource utilization for a compute node for a given set of resource types and process IDs. -Start a process and monitor its resource utilization. -Monitor resource utilization for a compute node asynchronously with the ability to dynamically change the resource types and process IDs being monitored. -Produce JSON reports of aggregated metrics. -Produce interactive HTML plots of the statistics.

Thom, Daniel [National Renewable Energy Laboratory

BULKI-Store v0.3.2

BULKI-Store is a distributed object storage system optimized for high-performance computing environments. Built with a Rust core and Python bindings, it efficiently manages scientific and machine learning datasets across HPC clusters. The system employs a client-server architecture with MPI integration, enabling seamless scaling on supercomputers like Perlmutter. BULKI-Store's object-oriented approach provides intuitive data organization with rich metadata support, contrasting with traditional file-based solutions. Key optimizations include selective checkpoint loading, unified checkpoint files, and object chunking for large data transfers. For machine learning workloads, BULKI-Store offers advantages through fine-grained access patterns, dynamic data sharing between training instances, and reduced memory pressure. Memory management features include strategic Python GC calls, minimized data copies, and batch processing capabilities. The system leverages Rayon's thread pool for asynchronous data prefetching and supports multiple CPU architectures (ARM64, x86, AMD, RISC-V). By combining performance optimizations with developer-friendly APIs, BULKI-Store addresses the complex data management challenges of modern HPC applications while maintaining compatibility across heterogeneous computing environments.

Zhang, Wei [Lawrence Berkeley National Laboratory

ON-OFF neuromorphic ISING machines using Fowler-Nordheim annealers

We introduce NeuroSA, a neuromorphic architecture specifically designed to ensure asymptotic convergence to the ground state of an Ising problem using a Fowler-Nordheim quantum mechanical tunneling based threshold-annealing process. The core component of NeuroSA consists of a pair of asynchronous ON-OFF neurons, which effectively map classical simulated annealing dynamics onto a network of integrate-and-fire neurons. The threshold of each ON-OFF neuron pair is adaptively adjusted by an FN annealer and the resulting spiking dynamics replicates the optimal escape mechanism and convergence of SA, particularly at low-temperatures. To validate the effectiveness of our neuromorphic Ising machine, we systematically solved benchmark combinatorial optimization problems such as MAX-CUT and Max Independent Set. Across multiple runs, NeuroSA consistently generates distribution of solutions that are concentrated around the state-of-the-art results (within 99%) or surpass the current state-of-the-art solutions for Max Independent Set benchmarks. Furthermore, NeuroSA is able to achieve these superior distributions without any graph-specific hyperparameter tuning. For practical illustration, we present results from an implementation of NeuroSA on the SpiNNaker2 platform, highlighting the feasibility of mapping our proposed architecture onto a standard neuromorphic accelerator platform.

42 ENGINEERING

Asynchronous GPU-based DEM solver embedded in commercial CFD software with polyhedral mesh support

A novel graphical processing unit-based discrete element method solver is introduced to improve stability, performance, and provide seamless integration into commercial or open-source computational fluid dynamics software. A key innovation is eliminating a need for network communication between solvers, which was previously required for cross-platform coupling. This is accomplished by a direct coupling method that employs dynamic-linked libraries. Furthermore, the solver optimizes memory usage by streamlining the particle-cell search algorithm by eliminating the cells' searching grid. This ensures the solver is compatible with a wide range of mesh types, providing high geometric flexibility. The approach simplifies the simulation process by directly incorporating computational fluid dynamics mesh information into the discrete element method solver. The performance analysis indicates about sixteen times boost in computational speed compared to benchmark central processing unit-based solvers. Finally, the solver's compatibility with polyhedral meshes, a vital advantage for complex geometries, is tested against a referenced study regarding the simulation of an immersed-tube fluidized bed.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Approach for energy efficient building design during early phase of design process

Energy consumption in the building sector is about 40% of total energy consumed globally and is trending upwards, along with its contribution to greenhouse gas (GHG) emissions. Given the adverse impacts of GHG emissions, it is crucial to integrate energy efficiency into building designs. The most significant opportunities for enhancing energy performance are present during the initial phases of building design, when there is less impact of other design constraints. Various tools exist for simulating different design options and providing feedback in terms of energy consumption and comfort parameters. These simulation outputs must then be analyzed to derive design solutions. This paper presents an innovative approach that utilizes user input parameters, processes them through cloud computing, and outputs easily understandable strategies for energy-efficient building design. The methodology employs Asynchronous Distributed Task Queues (DTQ) - a more scalable and reliable alternative to conventional speedup techniques-for conducting parametric energy simulations in the cloud. The goal of this approach is to assist design teams in identifying, visualizing, and prioritizing energy-saving design strategies from a range of possible solutions for each project. Furthermore, a tool ‘eDOT’ has been developed utilizing the discussed methodology. Unlike existing tools, eDOT leverages artificial intelligence to dynamically generate and provide design strategies during the early phases of design process. By simplifying the simulation process, eDOT enables design teams to make informed, data-driven decisions without needing to interpret complex simulation outputs. A case study simulated for two locations is provided in this paper to demonstrate the effectiveness of eDOT, further underscoring its practical impact on energy-efficient building design.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

UMap: An application-oriented user level memory mapping library

Exploiting the prominent role of complex memories in exascale node architecture, the UMap page fault handler offers new capabilities to access large memory-mapped data sets directly. UMap provides flexible configuration options to customize page handling to each application, including analysis of massive observational and simulation data sets. The high-performance design features I/O decoupling, dynamic load balancing, and application-level controls. Page faults triggered by application threads and processes accessing data mapped to a UMapp’ed region are handled via the Linux userfaultfd protocol, an asynchronous message-oriented kernel-user communication mechanism that avoids the context switch penalty of traditional signal fault handlers. UMap is fully open source. In this paper, we give an overview of the UMap library architecture, its extensible plugin architecture, and the use/performance of UMap in emerging heterogeneous memory hierarchies such as near-node Non-volatile Memory (NVM) and network attached memories. We highlight new capabilities in two pagefault management plugins, the NetworkStore and SparseStore. We demonstrate the integration between UMap and multiple ECP products including Caliper, Metall, ZFP, Mochi, and Ripples.

97 MATHEMATICS AND COMPUTING

High-frequency Data Integration for Landscape Model Calibration of Carbon Fluxes Across Diverse Tidal Marshes

Terrestrial Aquatic Interfaces (TAIs), and tidal wetlands in particular, store large amounts of carbon yet are not well represented in Earth System Models (ESMs). Predictions of carbon cycling and greenhouse gas (GHG) emissions in tidal wetlands are highly uncertain. Eddy covariance (EC) towers provide ecosystem-scale GHG flux data at a temporal resolution (every 30min) that is helpful for parameterizing and improving mechanistic realism in ESMs. We propose to use a network of eddy covariance towers and standardized ancillary data streams, along with mesocosm experiments and statistical analyses, across diverse tidal wetlands of North America to develop and improve biogeochemical modeling at the TAI. Our overarching objective is to improve understanding and process-based modeling of gross primary productivity (GPP) and CH4 emission responses, both non-linear and asynchronous, to stressors including plant inundation, disturbance, salinity and nitrogen loading.

54 ENVIRONMENTAL SCIENCES

Q-IRIS: The Evolution of the IRIS Task-Based Runtime to Enable Classical-Quantum Workflows

Extreme heterogeneity in emerging HPC systems are starting to include quantum accelerators, motivating runtimes that can coordinate between classical and quantum workloads. We present a proof-of-concept hybrid execution framework integrating the IRIS asynchronous task-based runtime with the XACC quantum programming framework via the Quantum Intermediate Representation Execution Engine (QIR-EE). IRIS orchestrates multiple programs written in the quantum intermediate representation (QIR) across heterogeneous backends (including multiple quantum simulators), enabling concurrent execution of classical and quantum tasks. Although not a performance study, we report measurable outcomes through the successful asynchronous scheduling and execution of multiple quantum workloads. To illustrate practical runtime implications, we decompose a four-qubit circuit into smaller subcircuits through a process known as quantum circuit cutting, reducing per-task quantum simulation load and demonstrating how task granularity can improve simulator throughput and reduce queueing behavior -- effects directly relevant to early quantum hardware environments. We conclude by outlining key challenges for scaling hybrid runtimes, including coordinated scheduling, classical-quantum interaction management, and support for diverse backend resources in heterogeneous systems.

Miniskar, Narasinga Rao [ORNL] (ORCID:000000018259

Classic and Quantum Task-Based Intelligent Runtime for QIRs Running on Multiple QPUs

High-performance computing systems are rapidly evolving into heterogeneous platforms that fuse quantum accelerators with traditional classical processing units (CPUs) and graphical processing units (GPUs). This convergence calls for runtimes capable of managing both classical and quantum workloads in a unified manner. We introduce an intelligent, task-based runtime that marries the Intelligent RuntIme System (IRIS) asynchronous scheduler with a quantum programming stack through the Quantum Intermediate Representation Execution Engine (QIR-EE). Our design allows programs written in the quantum intermediate representation (QIR) to be dispatched concurrently to a variety of back-ends, including multiple quantum simulators and nascent quantum processors, enabling genuine hybrid execution on a single node. To illustrate its practicality, we partition a 4-qubit and 20-qubit circuit into three sub-circuits using quantum circuit cutting via the QCut library. Each sub-circuit is simulated independently by the QIR-EE driver within IRIS, after which a classical post-processing step merges the simulation results to recover the outcome of the original full-circuit computation. This case study demonstrates how finer task granularity can enable the parallel execution and lower the simulation burden per quantum task while preserving overall accuracy, highlighting the feasibility of our hybrid approach.

Miniskar, Narasinga Rao [ORNL] (ORCID:000000018259

Direct Observations of Solute Dispersion in Rocks With Distinct Degree of Sub‐Micron Porosity

Abstract The transport of chemical species in rocks is affected by their structural heterogeneity to yield a wide spectrum of local solute concentrations. To quantify such imperfect mixing, advanced methodologies are needed that augment the traditional breakthrough curve analysis by probing solute concentration within the fluids locally. Here, we demonstrate the application of asynchronous, multimodality imaging by X‐ray computed tomography (XCT) and positron emission tomography (PET) to the study of passive tracer experiments in laboratory rock cores. The four‐dimensional concentration maps measured by PET reveal specific signatures of the transport process, which we have quantified using fundamental measures of mixing and spreading. We observe that the extent of solute spreading correlate strongly with the strength of subcore‐scale porosity heterogeneity measured by XCT, while dilution is enhanced in rocks containing substantial sub‐micron porosity. We observe that the analysis of different metrics is necessary, as they can differ in their sensitivity to the strength and forms of heterogeneity. The multimodality imaging approach is uniquely suited to probe the fundamental difference between spreading and mixing in heterogeneous media. We propose that when multi‐dimensional data is available, mixing and spreading can be independently quantified using the same metric. We also demonstrate that one‐dimensional transport models have limited predictive ability toward the internal evolution of the solute concentration, when the model is solely calibrated against the effluent breakthrough curves. The data set generated in this study can be used to build realistic digital rock models and to benchmark transport simulations that account deterministically for rock property heterogeneity.

Kurotori, Takeshi [Department of Chemical Engineer

Introduction: Neuromorphic Materials

The explosive growth in data collection and the need to process it efficiently, as well as the desire to automate increasingly complex tasks in transportation, medical care, manufacturing, security and many other fields have motivated a growing interest in neuromorphic computing. Unlike the binary, transistorbased ON/OFF logic gates and separate logic and memory functionalities employed in digital computing, neuromorphic computing is inspired by animal brains that use interconnected synapses and neurons to perform processing, storage and transmission of information at the same location, while only consuming ~20 W or less of power. Motivated by the brain’s efficiency, adaptability, self-learning and resiliency qualities, neuromorphic computing can be broadly defined as an approach to processing and storing information using hardware and algorithms inspired by models of biological neural systems. Present research in neuromorphic computing encompasses approaches that vary significantly in their degree of neuro-inspiration, from systems that only incorporate features such as asynchronous, event-driven operation or use crossbar arrays of non-volatile memory (NVM) elements to accelerate deep neural networks (DNNs), to designs that embrace the extreme parallelism, sparsity, reconfigurability, adaptability, complexity and stochasticity observed in nervous systems. The term ‘neuromorphic’ computing is often credited to Carver Mead, who in the 1980s investigated Si-based analog electronics to replicate functions of the animal retina. Earlier important advances in this field include the work of Frank Rosenblatt, who proposed the concept of the perceptron, Bernard Widrow, who used this concept to build one of the first analog neural networks, the Adaline and many other researchers (see ref. 6 for an historical perspective on neuromorphic computing). With the recent increase in the use of artificial intelligence and large language models, and rising concerns over the associated energy costs, interest in neuromorphic hardware has expanded rapidly. According to some estimates, driven largely by the drastic growth in the training use of artificial intelligence (AI) models using the current computing architectures, the energy cost of computing is projected to reach the energy supply worldwide by 2045. Furthermore, while this is not a realistic outcome, it means that, if more efficient computing technologies are not developed -- soon -- the world will soon become one where demand for energy and market constraints limit the continued increase of societal access to AI and cloud services from data centers. Data centers used for training and use of these models consume hundreds of terawatt hours of electricity, already past 4% of the US electricity demand.

Circuits

Fiber loop quantum buffer for photonic qubits

Abstract We report a fiber loop quantum buffer based on a low-loss 2 × 2 switch and a unit delay made of a fiber delay line. We characterize the device by using a two-photon polarization entangled state in which one photon of the entangled photon pair is stored and retrieved at a repetition rate up to 78 kHz. The device, which enables integer multiples of a unit delay, can store the qubit state in a unit of fiber delay line up to 5.4 km and the number of loop round-trips up to 3. Furthermore, we configure the device with other active elements to realize integer multiplier and divider of a unit delay of a qubit. The quantum state tomography is performed on the retrieved photon and its entangled photon. We obtain a state fidelity > 94 % with a maximum storage time of 52 s with an insertion loss of 5.56 dB. To further characterize the storing and retrieving processes of the device, we perform entanglement-assisted quantum process tomography on the buffered qubit state. The process fidelity of the device is > 0.98. Our result implies that the device preserves the superposition and entanglement of a qubit state from a two-photon polarization-entangled state. This is a significant step towards facilitating applications in optical asynchronous transfer mode based quantum networks.

Fook Lee, Kim (ORCID:0000000290901640)

Hybrid additive manufacturing of AISI 316L via asynchronous powder and hot-wire laser directed energy deposition

Hybrid Additive Manufacturing (AM) offers a way to leverage the advantages of different AM technologies, enabling the efficient production of sizeable parts without compromising material properties or geometric complexity capabilities. This study presents an asynchronous hybrid Directed Energy Deposition (DED) strategy employing laser powder DED and laser hot-wire DED. AISI 316L parts comprising multiple powder and wire segments were fabricated with optional machining on AISI 316L substrates to investigate how quality is impacted by (i) alternative process sequences (laser powder DED followed by laser hot-wire DED and vice versa), (ii) machined vs. as-printed interfacial conditions, and (iii) material deposition on top vs. alongside previously built segments. Optical microscopy, X-ray computed tomography, and Vickers hardness were used to characterize the morphology and microstructure of the parts, localized porosity and lack of fusion defects, bulk density, and mechanical properties. Interfacial machining was necessary for dimensional control but promoted lack of fusion voids, resulting in a 99.71 ± 0.01% dense part. As-printed interfaces resulted in a denser part (99.82 ± 0.02%) at the expense of dimensional accuracy. The hardness of the parts with as-printed and machined interfaces was 196 ± 0.37 HV and 192 ± 0.40 HV, respectively, compared to 156 ± 1.4 HV for the substrate. Depositing powder alongside or on top of wire sections resulted in interfaces with a hardness of 217 ± 2.2 HV, compared to 185 ± 3.4 HV for the wire-powder interfaces.

36 MATERIALS SCIENCE

Asynchronicity in opposed-piston RCMs: Does it matter?

Rapid Compression Machines (RCMs) are widely utilized to study combustion phenomena at engine-relevant conditions, and significant efforts are typically made to create a quiescent environment, particularly for investigations of autoignition chemistry. Opposed-piston configurations can be advantageous due to shorter compression times and reduced surface area to volume ratios. Each side must be actuated simultaneously, but this can be challenging in practice. These devices, like most RCMs, utilize hydraulics for actuation, speed control and arrestation of the piston at the end of the stroke; there is no mechanical control or linkage of the two piston trajectories. To quantify the magnitudes and effects of piston asynchronous behavior, this work employs both detailed experimental measurements and, for the first time, high-fidelity, Direct Numerical Simulation (DNS). The boundary conditions are carefully considered applying insight from high-resolution linear variable differential transformer (LVDT) measurements of the piston trajectory and a zero-dimensional kinematics model of the piston-shaft assembly. Sufficient resolution in the piston crevice region is used. The complicated fluid dynamical behavior that can evolve during piston compression and the ensuing delay processes due to offset timings from t offset = 0-10 ms is elucidated. It is found that near t offset = 6 ms and beyond, the boundary layer on the face of the first-seating piston can be sufficiently perturbed, due initially to reemergence of gas from the crevice of the firstseating piston, so that the adiabatic core can become degraded at long ignition delay times. Substantial mixing of colder gas into the interior of the reaction chamber can alter the measurements, similar to effects previously observed for improper piston crevice configuration. In conclusion, experimental techniques to mitigate asynchronous behavior are discussed and demonstrated.

33 ADVANCED PROPULSION SYSTEMS

Designing and prototyping extensions to the Message Passing Interface in MPICH

As HPC system architectures and the applications running on them continue to evolve, the MPI standard itself must evolve. The trend in current and future HPC systems toward powerful nodes with multiple CPU cores and multiple GPU accelerators makes efficient support for hybrid programming critical for applications to achieve high performance. However, the support for hybrid programming in the MPI standard has not kept up with recent trends. The MPICH implementation of MPI provides a platform for implementing and experimenting with new proposals and extensions to fill this gap and to gain valuable experience and feedback before the MPI Forum can consider them for standardization. Here, in this work, we detail six extensions implemented in MPICH to increase MPI interoperability with other runtimes, with a specific focus on heterogeneous architectures. First, the extension to MPI generalized requests lets applications integrate asynchronous tasks into MPI’s progress engine. Second, the iovec extension to datatypes lets applications use MPI datatypes as a general-purpose data layout API beyond just MPI communications. Third, a new MPI object, MPIX_Stream, can be used by applications to identify execution contexts beyond MPI processes, including threads and GPU streams. MPIX stream communicators can be created to make existing MPI functions thread-aware and GPU-aware, thus providing applications with explicit ways to achieve higher performance. Fourth, MPIX Streams are extended to support the enqueue semantics for offloading MPI communications onto a GPU stream context. Fifth, thread communicators allow MPI communicators to be constructed with individual threads, thus providing a new level of interoperability between MPI and on-node runtimes such as OpenMP. Lastly, we present an extension to invoke MPI progress, which lets users spawn progress threads with fine-grained control to adapt the communication performance to their application designs. We describe the design and implementation of these extensions, provide usage examples, and highlight their expected benefits with performance results.

97 MATHEMATICS AND COMPUTING

OpenSn: A massively parallel, open-source simulation environment for discrete ordinates radiation transport

OpenSn is an open-source, massively parallel deterministic radiation transport code for solving the discrete-ordinates ( S N ) form of the Boltzmann transport equation on unstructured, arbitrary polyhedral meshes. It supports high-fidelity simulations involving steady-state, eigenvalue, and adjoint problems for neutral particles (e.g., neutrons, photons, multi-particles), using the multigroup approximation in energy. OpenSn combines angular discretization via discrete ordinates with a discontinuous Galerkin finite element method (DGFEM) in space, enabling accurate resolution of transport physics on arbitrary polyhedral cells, included locally refined spatial grids. It includes multiple angular quadrature types, including locally refined angular quadratures. Written in modern C++ with a Python API, OpenSn runs efficiently on platforms ranging from laptops to supercomputers. The transport sweep algorithm is implemented using a task-based, directed-acyclic-graph (DAG) approach for each angle and supports asynchronous parallelism across thousands of MPI ranks. Group-set aggregation improves compute intensity, and synthetic acceleration techniques (e.g., diffusion synthetic acceleration, second-moment method) enhance solver convergence. OpenSn has been verified on reactor physics problems and demonstrated excellent weak and strong scaling performance on more than 32,768 processes, making it a versatile and robust platform for large-scale transport simulations in complex geometries.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS