Search NASA⌕ Search

SEARCH · Search NASA

Results for “execution”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

Advanced Shuttle Strategies for Parallel QCCD Architectures

Trapped ions (TIs) are at the forefront of quantum computing implementation, offering unparalleled coherence, fidelity, and connectivity. However, the scalability of TI systems is hampered by the limited capacity of individual ion traps, necessitating intricate ion shuttling for advanced computational tasks. The quantum charge-coupled device (QCCD) framework has emerged as a promising solution, facilitating ion mobility for universal quantum computation. Current QCCD architectures predominantly feature a linear topology, which is increasingly recognized as inefficient for complex quantum operations. Anticipating the shift toward more efficacious designs, this article introduces an innovative quantum scheduling strategy optimized for parallel QCCD topologies. Our strategy proposes a probabilistic formula for ion movement, alongside ingenious methods for local layer generation and layer compression, yielding a significant reduction in ion shuttle times. Through simulations, we demonstrate that our strategy not only substantially outstrips the linear model but also exhibits better performance over other parallel strategies that employ greedy algorithms. This is achieved through our nuanced resolution of complexities, such as traffic blocks and trap capacity limitations. The consequent reduction in shuttle operations leads to lower energy consumption and an enhancement in the quantum computer's fidelity, ultimately accelerating program execution times.

43 PARTICLE ACCELERATORS↗

Dynamic Charging Rendezvous and Motion Planning for a Multi-AGV Team Including a Mobile Charging Host

Teams of automated battery-powered electric vehicles have the potential to execute complex mission tasks in off-road environments for agriculture, military, and other applications. Limited onboard energy reserves hinder their adoption in large-scale resource-constrained environments, where recharging is a necessity. It may be infeasible to install a network of static charging stations in off-road environments. For this reason, dedicated mobile host vehicles with charging capabilities are proposed as a means to increase range and capabilities of the multivehicle team. Here, in this study, we consider an ad hoc planning framework, where results from a high-confidence trajectory planner are leveraged to plan charging rendezvous between a host and other worker vehicles in a receding horizon fashion to provide high confidence that energy reserves will not be prematurely exhausted. The core problem is posed so as to minimize the impact of recharging on the mission in terms of task delays, overall energy utilization, and costs of fast charging. Through extensive Monte Carlo simulations of an off-road mission, we show a decrease in task delays without substantial increases in energy needs by updating the charging rendezvous plan during the mission. However, if updates are made too often, model mismatch may cause unnecessary cycling and mission failure.

Energy constraints↗

Memory-Aware External Facelist Calculation: A Data-Parallel Atomic Hash Counting Approach

Unstructured volumetric meshes serve as fundamental data representations in various scientific simulations and analyses. They play a crucial role in representing complex computational domains and are essential for important numerical techniques, such as finite element analysis. Whenever such a mesh is read from a file, streamed in-situ, or generated by algorithms, scientific visualization libraries rely on calculating the external surface of a geometry, named “external facelist”, to produce a polygonal mesh for rendering. Consequently, external facelist calculation has become one of the most widely used algorithms in the scientific visualization domain, necessitating optimal performance. In this paper, we explore relevant work on external facelist calculation algorithms in two common visualization libraries, VTK and Viskores, assess their performance and memory constraints, and introduce a novel memory-aware external facelist calculation algorithm employing an atomic hash counting approach. This algorithm fully leverages Viskores' data-parallel primitive operations, facilitating its execution across diverse many-core architectures. Our algorithm features the lowest memory footprint on the GPU and the second-lowest on the CPU among all evaluated methods, and it also delivers the fastest performance on both CPU and GPU. It has been made available under an open-source license in the VTK and Viskores visualization systems.

Tsalikis, Spiros [Kitware] (ORCID:0000000151137195↗

On the bulk compaction of brittle granular materials, Part III: Brittle‐to‐ductile transition and yield strength

A new and simple method is presented that enables the estimation of the yield strength (σ y ) of brittle materials (e.g., ceramics, glasses). It results from the combination of sufficiently high-stress compaction of their granular form, postmortem analysis of the crushed particles to identify the critical particle size corresponding to their brittle-to-ductile transition, and the use of a developed and simple analytical expression. Here, this method was an outcome from Part I of this three-paper series. To execute it, a granular brittle material is compacted to a sufficiently high stress, whereby the acting comminution produces both a fraction of particles having a sufficiently small size formed by ductile or plastic-like deformation and a remaining fraction of larger particles formed from brittle fracture. Postmortem microscopy is then used to identify the smallest particle size whose morphology indicates it formed from brittle fracture (d B2D ). The brittle material's σ y can then be estimated using a combination of the d B2D , Kendall's and Griffith's theories, a priori knowledge of the material's fracture toughness (K Ic ), and a fracture mechanics shape factor constant (Y) using σ y = √((32 π K Ic 2 )/(3 Y 2 d B2D )). The method's development and its use to estimate σ y for several vitreous silicates, α-quartzes, and NaCl are provided.

Compaction↗

Investigating In-Situ Fracture Behaviors of Polymer Pipeline Materials in Hydrogen and Hydrogen-Methane Blended Gas Environments

To reduce carbon emissions, the US natural gas infrastructure is seen as a primary solution for efficiently transporting hydrogen gas. Blending hydrogen gas with natural gas and transporting it across a national infrastructure could save significant infrastructure costs. To properly operate the infrastructure under the new gas system, it is critical to understand material compatibility with hydrogen under various conditions. The Blended Gas CRADA, a Hyblend project, is established to determine the material compatibility of existing natural gas pipes with hydrogen gas. In this study, we investigate the in-plane fracture behaviors of MDPEMarlex and HDPEGDB exposed to hydrogen and hydrogen-methane blended gas. Single-edge notch bending geometry is used. All tests are executed in-situ with the gas environment. The experimental results show a significant effect of the gas environment on HDPEGDB specimens, reducing 5% (H2) to 42% (Blended gas) of specific fracture energy compared to non-aged specimens. For the MDPEMarlex, the effects of the gas environment have increased the specific fracture energy by 10% (H2) to 15% (Blended gas). Fracture surfaces of the tested samples are observed using an electronic microscope. The in-plane fracture surface of HDPEGDB shows a pronounced dimple fracture pattern after exposure to hydrogen and blended gas. The expanded fracture pattern contributes to lower the specific fracture energy. These observations provide critical information for validating polymer pipeline materials when interact with hydrogen and hydrogen-blend gas.

Ko, Seunghyun↗

Fully quantum algorithm for mesoscale fluid simulations with application to partial differential equations

Fluid flow simulations marshal our most powerful computational resources. In many cases, even this is not enough. Quantum computers provide an opportunity to speed up traditional algorithms for flow simulations. We show that lattice-based mesoscale numerical methods can be executed as efficient quantum algorithms due to their statistical features. This approach revises a quantum algorithm for lattice gas automata to reduce classical computations and state preparation at every time step. For this, the algorithm approximates the qubit relative phases and subtracts them at the end of each time step. Phases are evaluated using the iterative phase estimation algorithm and subtracted using single-qubit rotation phase gates. Further, this method optimizes the quantum resource required and makes it more appropriate for near-term quantum hardware. We also demonstrate how the checkerboard deficiency that the D1Q2 scheme presents can be resolved using the D1Q3 scheme. The algorithm is validated by simulating two canonical partial differential equations: the diffusion and Burgers' equations on different quantum simulators. We find good agreement between quantum simulations and classical solutions for the presented algorithm.

97 MATHEMATICS AND COMPUTING↗

Machine-learning-enabled on-the-fly analysis of RHEED patterns during thin film deposition by molecular beam epitaxy

Thin film deposition is a fundamental technology for the discovery, optimization, and manufacturing of functional materials. Deposition by molecular beam epitaxy (MBE) typically employs reflection high-energy electron diffraction (RHEED) as a real-time in situ probe of the growing film. However, the state-of-the-art for RHEED analysis during deposition requires human observation. Here, we present an approach using machine learning (ML) methods to monitor, analyze, and interpret RHEED images on-the-fly during thin film deposition. In the analysis workflow, RHEED pattern images are collected at one frame per second and featurized using a pretrained deep convolutional neural network. The feature vectors are then statistically analyzed to identify changepoints; these changepoints can be related to changes in the deposition mode from initial film nucleation to a transition regime, smooth film deposition, and in some cases, an additional transition to a rough, islanded deposition regime. The feature vectors are additionally analyzed via graph analysis and community classification. The graph is quantified as a stabilization plot, and we show that inflection points in the stabilization plot correspond to changes in the growth regime. The full RHEED analysis workflow is termed RHAAPsody and includes data transfer and output to a visual dashboard. We demonstrate the functionality of RHAAPsody by analyzing the precaptured RHEED images from epitaxial depositions of anatase TiO2 on SrTiO3(001) and show that the analysis workflow can be executed in less than 1 s. Our approach shows promise as one component of ML-enabled real-time feedback control of the MBE deposition process.

36 MATERIALS SCIENCE↗

Evidence of scaling advantage for the quantum approximate optimization algorithm on a classically intractable problem

The quantum approximate optimization algorithm (QAOA) is a leading candidate algorithm for solving optimization problems on quantum computers. However, the potential of QAOA to tackle classically intractable problems remains unclear. Here, we perform an extensive numerical investigation of QAOA on the low autocorrelation binary sequences (LABS) problem, which is classically intractable even for moderately sized instances. We perform noiseless simulations with up to 40 qubits and observe that the runtime of QAOA with fixed parameters scales better than branch-and-bound solvers, which are the state-of-the-art exact solvers for LABS. The combination of QAOA with quantum minimum finding gives the best empirical scaling of any algorithm for the LABS problem. We demonstrate experimental progress in executing QAOA for the LABS problem using an algorithm-specific error detection scheme on Quantinuum trapped-ion processors. Our results provide evidence for the utility of QAOA as an algorithmic component that enables quantum speedups.

97 MATHEMATICS AND COMPUTING↗

Dynamic STEM-EELS for single-atom and defect measurement during electron beam transformations

This study introduces the integration of dynamic computer vision–enabled imaging with electron energy loss spectroscopy (EELS) in scanning transmission electron microscopy (STEM). This approach involves real-time discovery and analysis of atomic structures as they form, allowing us to observe the evolution of material properties at the atomic level, capturing transient states traditional techniques often miss. Rapid object detection and action system enhances the efficiency and accuracy of STEM-EELS by autonomously identifying and targeting only areas of interest. This machine learning (ML)–based approach differs from classical ML in that it must be executed on the fly, not using static data. We apply this technology to V-doped MoS 2 , uncovering insights into defect formation and evolution under electron beam exposure. This approach opens uncharted avenues for exploring and characterizing materials in dynamic states, offering a pathway to increase our understanding of dynamic phenomena in materials under thermal, chemical, and beam stimuli.

47 OTHER INSTRUMENTATION↗

Discovery of GuaB inhibitors with efficacy against Acinetobacter baumannii infection

ABSTRACT Guanine nucleotides are required for growth and viability of cells due to their structural role in DNA and RNA, and their regulatory roles in translation, signal transduction, and cell division. The natural antibiotic mycophenolic acid (MPA) targets the rate-limiting step inde novoguanine nucleotide biosynthesis executed by inosine-5´-monophosphate dehydrogenase (IMPDH). MPA is used clinically as an immunosuppressant, but whetherin vivoinhibition of bacterial IMPDH (GuaB) is a valid antibacterial strategy is controversial. Here, we describe the discovery of extremely potent small molecule GuaB inhibitors (GuaBi) specific to pathogenic bacteria with a low frequency of on-target spontaneous resistance and bactericidal efficacyin vivoagainstAcinetobacter baumanniimouse models of infection. The spectrum of GuaBi activity includes multidrug-resistant pathogens that are a critical priority of new antibiotic development. Co-crystal structures ofA. baumannii, Staphylococcus aureus, andEscherichia coliGuaB proteins bound to inhibitors show comparable binding modes of GuaBi across species and identifies key binding site residues that are predictive of whole-cell activity across both Gram-positive and Gram-negative clades of Bacteria. The clearin vivoefficacy of these small molecule GuaB inhibitors in a model ofA. baumanniiinfection validates GuaB as an essential antibiotic target. IMPORTANCE The emergence of multidrug-resistant bacteria worldwide has renewed interest in discovering antibiotics with novel mechanism of action. For the first time ever, we demonstrate that pharmacological inhibition ofde novoguanine biosynthesis is bactericidal in a mouse model ofAcinetobacter baumanniiinfection. Structural analyses of novel inhibitors explain differences in biochemical and whole-cell activity across bacterial clades and underscore why this discovery may have broad translational impact on treatment of the most recalcitrant bacterial infections.

Microbiology↗

MVP: a modular viromics pipeline to identify, filter, cluster, annotate, and bin viruses from metagenomes

While numerous computational frameworks and workflows are available for recovering prokaryote and eukaryote genomes from metagenome data, only a limited number of pipelines are designed specifically for viromics analysis. With many viromics tools developed in the last few years alone, it can be challenging for scientists with limited bioinformatics experience to easily recover, evaluate quality, annotate genes, dereplicate, assign taxonomy, and calculate relative abundance and coverage of viral genomes using state-of-the-art methods and standards. Here, we describe Modular Viromics Pipeline (MVP) v.1.0, a user-friendly pipeline written in Python and providing a simple framework to perform standard viromics analyses. MVP combines multiple tools to enable viral genome identification, characterization of genome quality, filtering, clustering, taxonomic and functional annotation, genome binning, and comprehensive summaries of results that can be used for downstream ecological analyses. Overall, MVP provides a standardized and reproducible pipeline for both extensive and robust characterization of viruses from large-scale sequencing data including metagenomes, metatranscriptomes, viromes, and isolate genomes. As a typical use case, we show how the entire MVP pipeline can be applied to a set of 20 metagenomes from wetland sediments using only 10 modules executed via command lines, leading to the identification of 11,656 viral contigs and 8,145 viral operational taxonomic units (vOTUs) displaying a clear beta-diversity pattern. Further, acting as a dynamic wrapper, MVP is designed to continuously incorporate updates and integrate new tools, ensuring its ongoing relevance in the rapidly evolving field of viromics. MVP is available at https://gitlab.com/ccoclet/mvp and as versioned packages in PyPi and Conda.

59 BASIC BIOLOGICAL SCIENCES↗

Fast and Scalable FFT-Based GPU-Accelerated Algorithms for Block-Triangular Toeplitz Matrices with Application to Linear Inverse Problems Governed by Autonomous Dynamical Systems

In this work, we present an efficient and scalable algorithm for performing matrix-vector multiplications (matvecs) for block Toeplitz matrices. Such matrices, which are shift-invariant with respect to their blocks, arise in the context of solving inverse problems governed by autonomous systems, and time-invariant systems in particular. In this article, we consider inverse problems that infer unknown parameters from observational data of a linear time-invariant dynamical system given in the form of partial differential equations (PDEs). Matrix-free Newton-conjugate-gradient methods are often the gold standard for solving these inverse problems, but they require numerous actions of the Hessian on a vector. Matrix-free adjoint-based Hessian matvecs require solution of a pair of linearized forward/adjoint PDE solves per Hessian action, which may be prohibitive for large-scale inverse problems. Time invariance of the forward PDE problem leads to a block Toeplitz structure of the discretized parameter-to-observable (p2o) map defining the mapping from inputs (parameters) to outputs (observables) of the PDEs. This block Toeplitz structure enables us to exploit two key properties: (1) compact storage of the p2o map and its adjoint, and (2) efficient fast Fourier transform–based Hessian matvecs. The proposed algorithm is mapped onto large multi-GPU clusters and achieves more than 80% of peak bandwidth on NVIDIA A100 GPUs. Excellent weak scaling is shown for up to 48 A100 GPUs. For the targeted problems, the implementation executes Hessian matvecs within fractions of a second, which is orders of magnitude faster than can be achieved by conventional matrix-free Hessian matvecs via forward/adjoint PDE solves.

97 MATHEMATICS AND COMPUTING↗

Scientific program for the Forward Physics Facility

The recent direct detection of neutrinos at the LHC has opened a new window on high-energy particle physics and highlighted the potential of forward physics for groundbreaking discoveries. In the last year, the physics case for forward physics has continued to grow, and there has been extensive work on defining the Forward Physics Facility and its experiments to realize this physics potential in a timely and cost-effective manner. Following a 2-page Executive Summary, we first present the status of the FPF, beginning with the FPF’s unique potential to shed light on dark matter, new particles, neutrino physics, QCD, and astroparticle physics. We then summarize the current designs for the Facility and its experiments, FASER2, FASER 2, FORMOSA, and FLArE.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

iDDS: intelligent distributed dispatch and scheduling for workflow orchestration

The intelligent distributed dispatch and scheduling (iDDS) service is a versatile workflow orchestration system designed for large-scale, distributed scientific computing. iDDS extends traditional workload and data management by integrating data-aware execution, conditional logic, and programmable workflows, enabling automation of complex and dynamic processing pipelines. Originally developed for the ATLAS experiment at the large hadron collider, iDDS has evolved into an experiment-agnostic platform that supports both template-driven workflows and a Function-as-a-Task model for Python-based orchestration. This paper presents the architecture and core components of iDDS, highlighting its scalability, modular message-driven design, and integration with systems such as PanDA and Rucio. We demonstrate its versatility through real-world use cases: fine-grained tape resource optimization for ATLAS, orchestration of large Directed Acyclic Graph (DAG) workflows for the Rubin Observatory, distributed hyperparameter optimization for machine learning applications, active learning for physics analyses, and AI-assisted detector design at the electron–ion collider. By unifying workload scheduling, data movement, and adaptive decision-making, iDDS reduces operational overhead and enables reproducible, high-throughput workflows across heterogeneous infrastructures. We conclude with current challenges and future directions, including interactive, cloud-native, and serverless workflow support.

97 MATHEMATICS AND COMPUTING↗

FuseIM: Fusing Probabilistic Traversals for Influence Maximization on Exascale Systems

Probabilistic breadth-first traversals (BPTs) are used in many network science and graph machine learning applications. In this paper, we are motivated by the application of BPTs in stochastic diffusion-based graph problems such as influence maximization. These applications heavily rely on BPTs to implement a Monte-Carlo sampling step for their approximations. Given the large sampling complexity, stochasticity of the diffusion process, and the inherent irregularity in real-world graph topologies, efficiently parallelizing these BPTs remains significantly challenging. In this paper, we present a new algorithm to fuse massive number of concurrently executing BPTs with random starts on the input graph. Our algorithm is designed to fuse BPTs by combining separate traversals into a unified frontier on distributed multi-GPU systems. To show the general applicability of the fused BPT technique, we have incorporated it into two state-of-the-art influence maximization parallel implementations (gIM and Ripples). Our experiments on up to 4K nodes of the OLCF Frontier supercomputer (32,768 GPUs and 196K CPU cores) show strong scaling behavior, and that fused BPTs can improve the performance of these implementations up to 34x (for gIM) and ~360x (for Ripples).

Neff, Reece W.↗

Real-time High-resolution X-Ray Computed Tomography

Computed Tomography (CT) serves as a key imaging technology that relies on computationally intensive filtering and back-projection algorithms for 3D image reconstruction. While conventional high-resolution image reconstruction (> 2K3) solutions provide quick results, they typically treat reconstruction as an offline workload to be performed remotely on large-scale HPC systems. The growing demand for post-construction AI-driven analytics and the need for real-time adjustments call for high-resolution reconstruction solutions that are feasible on local computing resources, i.e. a multi-GPU server at most. In this paper, we propose a novel approach that utilizes Tensor Cores to optimize image reconstruction without sacrificing precision. We also introduce a framework designed to enable real-time execution of end-to-end distributed image reconstruction in a multi-GPU environment. Evaluations conducted on a single Nvidia A100 and H100 GPU show performance improvements of 1.91 × and 2.15 × compared to highly optimized production libraries. Furthermore, our framework, when deployed on 8-card Nvidia A100 GPU system, demonstrates the ability to reconstruct real-world datasets into 20483 volumes (32 GB) in slightly more than one minute and 40963 volumes (256 GB) in 7 minutes.

Wu, Du↗

High-Throughput Computing: Case Study of Medical Image Processing Applications

HPC is designed for large-scale simulations using monolithic codes of tightly coupled processes highly optimized to deliver decreased time to solution. Medical image processing is not a traditional field of HPC. Similar to AI applications, medical image processing parses large datasets, typically multiple times, to support a variety of studies for classification, diagnosis or monitoring purposes. The convergence of AI, HPC and Big Data encouraged more fields using image processing to transition to HPC. However, not all applications benefit from the same optimizations. In this paper we focus on high throughput medical image processing applications that analyze a huge dataset of small MRI images and that require HPC systems to decrease the time of parsing the entire dataset and not individual MRIs. We show in this research the performance of running SLANT, an image processing application for a whole brain segmentation, on large-scale systems and highlight performance limitations. We present optimizations prioritizing throughput that exhibit a 3.5x speed-up on the Summit Supercomputer that can be used as a baseline for building a high-throughput execution framework for other HPC systems.

Predescu, Maria↗

Quantum Circuit Cutting for Classical Shadows

Classical shadow tomography is a sample-efficient technique for characterizing quantum systems and predicting many of their properties. Circuit cutting is a technique for dividing large quantum circuits into smaller fragments that can be executed more robustly using fewer quantum resources. We introduce a divide-and-conquer circuit cutting method for estimating the expectation values of observables using classical shadows. We derive a general formula for making predictions using the classical shadows of circuit fragments from arbitrarily cut circuits and provide the sample complexity analysis for the case when observables factorize across fragments. Then, we numerically show that our divide-and-conquer method outperforms traditional uncut shadow tomography when estimating high-weight observables that act non-trivially on many qubits and discuss the mechanisms for this advantage.

97 MATHEMATICS AND COMPUTING↗