Search NASA⌕ Search

SEARCH · Search NASA

Results for “execution”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

TorchBraid: High-Performance Layer-Parallel Training of Deep Neural Networks with MPI and GPU Acceleration

TorchBraid is a high-performance implementation of layer-parallel training for deep neural networks (DNNs) supporting MPI-based parallelism and GPU acceleration. Layer-parallel training has been developed to overcome the serialization inherent in forward and backward propagation of DNNs that limits utilization of computational resources in the strong scaling limit. To achieve this, TorchBraid integrates the PyTorch neural network framework with the state-of-the-art XBraid time-parallel library. Furthermore, this article presents the use and performance of TorchBraid, in addition to solutions for overcoming the algorithmic challenges inherent in combining automatic differentiation with layer-parallel. Results are presented with and without GPU acceleration for the Tiny ImageNet and MNIST image classification data sets, as well as recurrent neural networks. Overall, TorchBraid enables fast training of DNNs, both in a strong and weak scaling context. In addition to the TorchBraid software, several new advances in applying layer-parallel algorithms are detailed. Integration of layer-parallel with data-parallel algorithms is presented for the first time, showing the computational advantages of the combination. Standard deep learning techniques, like batch-normalization, are developed for layer-parallel training. Finally, a new approach combining layer-parallel with spatial coarsening in order to accelerate training for 3D image classification shows roughly a 10× speedup over serial execution.

Layer-parallel↗

The ArborX Library: Version 2.0

This article provides an overview of the 2.0 release of the ArborX library, a performance portable geometric search library based on Kokkos. We describe the major changes in ArborX 2.0 including a new interface for the library to support a wider range of user problems, new search data structures (brute force and distributed), support for user functions to be executed on the results (callbacks), and an expanded set of the supported algorithms (ray tracing and clustering).

GPU↗

The LCLStream Ecosystem for Multi-Institutional Dataset Exploration

We describe a new end-to-end experimental data streaming framework designed from the ground up to support new types of applications – AI training, extremely high-rate X-ray time-of-flight analysis, crystal structure determination with distributed processing, and custom data science applications and visualizers yet to be created. Throughout, we use design choices merging cloud microservices with traditional HPC batch execution models for security and flexibility. This project makes a unique contribution to the DOE Integrated Research Infrastructure (IRI) landscape. By creating a flexible, API-driven data request service, we address a significant need for high-speed data streaming sources for the X-ray science data analysis community. With the combination of data request API, mutual authentication web security framework, job queue system, high-rate data buffer, and complementary nature to facility infrastructure, the LCLStreamer framework has prototyped and implemented several new paradigms critical for future generation experiments.

Rogers, David [ORNL] (ORCID:0000000251871768)↗

EPOC Deep Dive Retrospective: A Brief Overview of 7 years of Science Engagement Discussions

Understanding the appropriate ways cyberinfrastructure can be designed, implemented, and executed for scientific use cases requires a deep understanding of the way that researchers and educators interact with technology, and how it may be best implemented to suit their needs. The Engagement and Performance Operations Center (EPOC) has conducted a series of scientific “Deep Dives” of use cases at partner institutions to better understand the requirements for modern scientific innovation across the United States research complex. The results of these activities have revealed gaps in the way that technology has been used to foster research activities. This gap in cyberinfrastructure support has impacts for the overall productivity and innovation possibilities for scientific users.

Zurawski, Jason↗

Scalability Analysis of Quantum Models for Stress and Emotion Detection

Stress and emotion detection from high-dimensional physiological signals is a challenging task, particularly when aiming for accurate classification across diverse behavioral states. Quantum machine learning (QML) is promising for modeling such high-dimensional data, but scalability is limited by qubit resources and the exponential cost of classical statevector simulation. This work studies the scalability of quantum support vector machines (QSVMs) for binary stress detection and three-class emotion recognition (Negative/Neutral/Positive) under varying qubit counts and angle-encoding strategies. We also present a comparison study with one-feature-per-qubit (1:1) and two-features-per-qubit (2:1) mappings. Experiments are executed on HPC infrastructure using NVIDIA CUDA-Q to evaluate performance, variance, and class-dependent separability at higher-qubit setups. Results show that larger Hilbert spaces can improve peak accuracy but may increase instability. At the same time, dense 2:1 encoding yields more consistent stress detection performance. For emotion recognition, scaling improves discrimination for classes like Negative and Positive more than Neutral. We find that effective QML scaling is task-dependent and benefits more from encoding design than simply increasing qubit count.

Onim, Md. Saif Hassan [University of Tennessee, Kn↗

Cross-Domain Reasoning for Neuromorphic Model Design

Designing performant neuromorphic models requires reasoning across neuroscience, neuromorphic computing, and machine learning, making it a natural target for cross-domain hypothesis generation. Our primary contribution is a multi-corpus knowledge graph spanning all three domains, which we show substantially increases cross-domain retrieval novelty over single-corpus baselines. We additionally introduce NeuKReAct, an agentic reasoning framework that iteratively retrieves from this graph and synthesizes design hypotheses via a step-by-step blackboard architecture, enabling structured compartmentalization of design decisions. Lastly, we introduce an execution head that translates hypotheses into structured design documents and runnable code. We evaluate novelty using a combinatorial creativity metric that measures cross-domain retrieval distance across the citation graph. Our results confirm that corpus breadth is the dominant driver of novelty. Moreover, we highlight a concrete instance of the novelty-utility tradeoff within NeuKReAct, underscoring a need for joint creativity evaluation, balancing both novelty and utility.

Ramavarapu, Vikram [ORNL] (ORCID:0009000188757213)↗

FLYING SERVING: On-the-Fly Parallelism Switching for Large Language Model Serving

Production LLM serving must simultaneously deliver high throughput, low latency, and sufficient context capacity under non-stationary traffic and mixed request requirements. Data parallelism (DP) maximizes throughput by running independent replicas, while tensor parallelism (TP) reduces per-request latency and pools memory for long-context inference. However, existing serving stacks typically commit to a static parallelism configuration at deployment; adapting to bursts, priorities, or long-context requests is often disruptive and slow. We present Flying Serving, a vLLM-based system that enables online DP-TP switching without restarting engine workers. Flying Serving makes reconfiguration practical by virtualizing the state that would otherwise force data movement: (i) a zero-copy Model Weights Manager that exposes TP shard views on demand, (ii) a KV Cache Adaptor that preserves request KV state across DP/TP layouts, (iii) an eagerly initialized Communicator Pool to amortize collective setup, and (iv) a deadlock-free scheduler that coordinates safe transitions under execution skew. Across three popular LLMs and realistic serving scenarios, Flying Serving improves performance by up to 4.79 × under high load and 3.47 × under low load while supporting latency- and memory-driven requests.

Gao, Shouwei [ORNL]↗

Scalable Federated Learning for Scientific Foundation Models on Leadership-Class Systems

Federated learning (FL) at leadership-class HPC systems remains largely unexplored, despite growing interest in deploying federated workflows on modern HPC systems. This paper provides the first system-level empirical characterization of federated fine-tuning of pretrained foundation models on an exascale supercomputer under a multi-node deployment. Using up to 96 concurrent FL clients deployed across Frontier nodes, we study the impact of client scale, model size, data heterogeneity, partial participation, and differential privacy on runtime, communication overhead, and convergence stability. Our results show that pretrained transformer models remain robust to heterogeneity, client dropout, and privacy noise, while system efficiency degrades rapidly with scale as synchronizat and orchestration dominate runtime. We further demonstrate that system-aware execution strategies, including intra-node aggregation and early aggregation, significantly reduce wall-clock time without degrading model quality. These findings establish a practical performance baseline and inform the design of communication-efficient FL systems on leadership-class HPC platforms.

Kotevska, Olivera [ORNL] (ORCID:0000000316772243)↗

Did You Win the GPU Cloud Lottery? Benchmarking from TFLOPS to Tokens/$

Cloud GPUs are commonly assumed to deliver consistent performance for a given GPU model. This assumption does not always hold: cloud providers employ diverse system configurations and virtualization mechanisms, and GPUs themselves exhibit non-negligible manufacturing variability (the silicon lottery). In this work, we present a large-scale measurement study of GPU performance variability across 11 cloud providers, covering over 3,500 physical GPUs and 6,800 benchmark runs. Our hierarchical analysis shows that while execution-level variation stays below 9%, performance varies by up to 38% across devices and providers for the same GPU model. Regression analysis indicates that driver- and OS-related software factors contribute less than 1% of the variance; instead, silicon lottery effects dominate observed performance variation, and cloud providers further amplify them through persistent, systematic second-order effects.

Slynko, Platon [Silicon Data, New York, USA] (ORCI↗

SetGo: Metadata Readiness for Scientific AI Datasets

Scientific datasets intended for AI use require both computational readiness for model training and metadata readiness for discovery, sharing, and reuse. The Readiness Engine for Data Integration (REDI) addresses computational readiness, but no corresponding tool evaluates whether a dataset’s metadata are sufficiently complete, governed, and standards-compliant for publication and agent-based consumption. Existing FAIR assessors operate only on published repository records, and no single system covers FAIR compliance, licensing, provenance, governance, reproducibility, and catalog readiness together. We present SetGo, an open-source Python toolkit that assesses and repairs metadata readiness across these six dimensions before a dataset is published or archived. Applied to four scientific corpora, SetGo surfaces deficiencies that general-purpose tools do not detect: ERA5 climate metadata scores 4% on ACDD 1.3 compliance; materials datasets fail OPTIMADE species-definition requirements; and PDB-derived proteomics data carries licensing terms incompatible with standard SPDX identifiers. Guided enrichment raises overall FAIR scores from 52–57% to 81–91%, and a single setgo publish command pushes to Hugging Face Hub, CKAN, or OpenMetadata with ML Commons Croissant 1.0 metadata sidecars. To support interactive and automated workflows, SetGo integrates with coding agents powered by large language models (LLMs) through a /setgo skill that enables natural-language execution of the full assess–enrich–publish loop, with user involvement limited to supplying missing metadata values.

Wilkinson, Sean [ORNL] (ORCID:0000000214437479)↗

Reagent-free Hyperspectral Diagnosis of SARS-CoV-2 Infection in Saliva Samples

Rapid, reagent-free pathogen-agnostic diagnostics that can be performed at the point of need are vital for preparedness against future outbreaks. Yet, many current strategies are pathogen-specific and require several reagents. We present hyperspectral sensing, using light to non-invasively measure the composition of several molecules to form a spectral signature, to overcome these barriers. To generate these spectral signatures, we present the ProSpectral TM V1, a novel, miniaturized hyperspectral platform with high spectral resolution with two mini-spectrometers. Furthermore, we developed state-of-the-art ML pipelines for near real-time analysis of spectral signatures in saliva samples. We found that we could accurately identify SARS-CoV-2 infection status in double-blinded saliva samples and demonstrate 100% accuracy on a hold out test dataset. To our knowledge, this establishes the fastest hyperspectral diagnostic platform and in a small form factor, and executable with liquid samples, without ligands or reagents, all while maintaining PCR level specificity and sensitivity.

59 BASIC BIOLOGICAL SCIENCES↗

Experimental Benchmarking for High-Reproducibility, Cross-Institutional Evaluation of Iron Redox Electrochemistry

We present a practical case study standardizing experimental protocols between collaborators with the goal of understanding ferrous iron (Fe 2+ ) chemistry and improving the iron deposition reaction for energy-efficient, electrochemical iron production. The study of iron reactions can be difficult, as aqueous iron electrolytes exhibit complex behaviors that can lead to differing interpretation of ostensibly similar experiments. The question we want to answer: are we studying the same chemistry? Our protocols address inherent challenges such as the tendency for Fe 2+ to spontaneously oxidize to ferric iron (Fe 3+ ) and the production of hydrogen at the potentials of interest. Our standardized protocol, executed by four collaborators in different labs and institutions, yields high-reproducibility results, and identified glassy carbon electrode surface quality and Fe 3+ impurities in the salt as key factors with outsized effects on cyclic voltammetry measurements. The process of developing the protocols helped to troubleshoot underlying issues that created poorer reversibility and reproducibility. This study highlights the fact that even nominally straightforward electrochemical systems can yield vastly different outcomes due to small differences in experimental preparation and serves as a useful example for creating transparent and achievable standards for the generation of reliable datasets that can be widely used and shared.

Ketter, Benjamin [Argonne National Laboratory (ANL↗

Solovay-Kitaev Algorithm and Randomized Compilation Data Availability

This zipped folder contains simulation notebooks, simulated data, and experimental data from the QSCOUT trapped-ion device that were used in the publication "Solovay-Kitaev Algorithm and Randomized Compilation" (https://doi.org/10.1103/ll6m-dbl7). The raw data is in the form of measurement outcomes of simple tomographic quantum circuits that were executed on the QSCOUT device and simulated using JAQALPAQ. These data are used to create plots within the jupyter notebooks that were included in the publication.

Quantum benchmarking↗

Fusion Neutron Generator

The proposed code, named FROG (Fusion neutron Generator) is built upon the open-source particle transport Monte Carlo toolkit Geant4. Geant4 provides C++ classes that can be leveraged to build application-specific codes dealing with the transport of particles through matter. Geant4-based codes are applied in high-energy particle physics experiments, medical applications, shielding, and space applications for example. The FROG code allows the user to define the geometry of a neutron converter device shaped as a hollow cylinder, where a neutron breeding material such as lithium deuteride (LiD) is cladded by two concentric cylinders. Such neutron converter is then placed inside a regular nuclear fission reactor, where thermal neutrons will react with the neutron breeder material (typically, Lithium 6), and through a series of reactions, will generate high-energy neutrons – neutrons whose kinetic energy are around 14 MeV. The hollowed central portion can hold a specimen that will be bombarded by high-energy neutrons created inside the neutron breeding material. Figuratively speaking, this type of device transforms neutrons from thermal (~0.625 eV) to fusion (~14 MeV) energies and is sometimes termed “fusion-to-thermal neutron converters” in the literature. The code consists of C++ source file compiled and linked to generate an executable. The user can select the dimensions of the converter (radius, length, and thickness of the breeder material), the breeder material type, the cladding material, and the specimen material that will be activated or irradiated. As input, the neutron flux for a specific location inside a reactor, for instance, positions in ATR, is required. As output, the code predicts the number of high-energy neutrons produced, the total neutron flux and fluence as well as its detailed spectrum. The physics involved in such device is very complex, as it requires modeling neutron transport, light-ion (tritons) transport, as well as fusion reactions. The Geant4 toolkit provides the required physical models.

Martin, NicholasP. [Idaho National Laboratory (INL↗

DQL (Django Query Logger) [SWR-24-96]

The Django Query Logger (DQL) is a tool that allows developers of Django web applications to stream the raw database queries that are being executed in real-time from any Django application for review, profiling, filtering, formatting, and analysis. Reference herein to any specific commercial products, process, or service by trade name, trademark, manufacturer, or otherwise, does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or Alliance for Sustainable Energy, LLC. The views and opinions of authors expressed in the available or referenced documents do not necessarily state or reflect those of the United States Government or Alliance.

Swindler, Alexander↗

Behavior, Energy, Autonomy, Mobility Modeling Framework (BEAM) v1.0

The Behavior, Energy, Autonomy, and Mobility (BEAM) model is an integrated, agent-based travel demand simulation framework. Individual agents express preferences through a utility- maximizing evolutionary algorithm that minimizes each individual’s cost and time spent traveling via diverse modal options, including the competition for scarce supply resources such as parking spaces and charging infrastructure. BEAM simulates the essential elements that compose a dynamic transportation system. From the road network, parking and charging infrastructure, to the transit system and a synthetic population with plans and preferences, the virtual system is an amalgamation of multiple spatially resolved layers that together represent an integrated transportation system. BEAM is an extension to the MATSim (Multi-Agent Transportation Simulation) model, where agents employ reinforcement learning across successive simulated days to maximize their personal utility through plan mutation (exploration) and selecting between previously executed plans (exploitation). The BEAM model shifts some of the behavioral emphasis in MATSim from across-day planning to within- day planning, where agents dynamically respond to the state of the system during the mobility simulation. In BEAM, agents can plan across all major modes of travel including driving, walking, biking, transit, and demand-responsive ride hailing. It is designed to integrate with other open source transportation models, such as ActivitySim.

Lazarus, Jessica↗

ninterp: N-dimensional interpolation Rust crate [SWR-25-25]

The ninterp crate provides multivariate interpolation over a regular, sorted, nonrepeating grid of any dimensionality. A variety of interpolation strategies are implemented, however more are likely to be added. Extrapolation beyond the range of the supplied coordinates is supported for 1-D linear interpolators, using the slope of the nearby points. There are hard-coded interpolators for lower dimensionalities (up to N = 3) for better runtime performance. All interpolation is handled through instances of the Interpolator enum, with the selected tuple variant containing relevant data. Interpolation is executed by calling Interpolator::interpolate.

Carow, Kyle↗

ATom (Acoustic Tomography Processing Suite) [SWR-24-120]

Acoustic tomography seeks the best-fit fluctuating velocity and temperature fields that explain a collection of signal travel times in a region of interest. This codebase defines an end-to-end framework for executing turbulent field retrievals from acoustic signals, acoustic signal design and processing tools for the physical array, and analysis tool that leverage virtual acoustic tomography arrays based on large-eddy simulations of the atmospheric boundary layer.

Hamilton, Nicholas↗