Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallel”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 955 records · Page 53

High energy density picoliter-scale zinc-air microbatteries for colloidal robotics

The recent interest in microscopic autonomous systems, including microrobots, colloidal state machines, and smart dust, has created a need for microscale energy storage and harvesting. However, macroscopic materials for energy storage have noted incompatibilities with microfabrication techniques, creating substantial challenges to realizing microscale energy systems. Here, we photolithographically patterned a microscale zinc/platinum/SU-8 system to generate the highest energy density microbattery at the picoliter (10 −12 liter) scale. The device scavenges ambient or solution-dissolved oxygen for a zinc oxidation reaction, achieving an energy density ranging from 760 to 1070 watt-hours per liter at scales below 100 micrometers lateral and 2 micrometers thickness in size. The parallel nature of photolithography processes allows 10,000 devices per wafer to be released into solution as colloids with energy stored on board. Within a volume of only 2 picoliters each, these primary microbatteries can deliver open circuit voltages of 1.05 ± 0.12 volts, with total energies ranging from 5.5 ± 0.3 to 7.7 ± 1.0 microjoules and a maximum power near 2.7 nanowatts. We demonstrated that such systems can reliably power a micrometer-sized memristor circuit, providing access to nonvolatile memory. We also cycled power to drive the reversible bending of microscale bimorph actuators at 0.05 hertz for mechanical functions of colloidal robots. Additional capabilities, such as powering two distinct nanosensor types and a clock circuit, were also demonstrated. The high energy density, low volume, and simple configuration promise the mass fabrication and adoption of such picoliter zinc-air batteries for micrometer-scale, colloidal robotics with autonomous functions.

Robotics↗

Laboratory evolution in Novosphingobium aromaticivorans enables rapid catabolism of a model lignin-derived aromatic dimer

Lignin contains a variety of interunit linkages, leading to a range of potential decomposition products that can be used as carbon and energy sources by microbes. β-O-4 linkages are the most common in native lignin, and associated catabolic pathways have been well characterized. However, the fate of the mono-aromatic intermediates that result from β-O-4 dimer cleavage has not been fully elucidated. Here, we used experimental evolution to identify mutant strains of Novosphingobium aromaticivorans with improved catabolism of a model aromatic dimer containing a β-O-4 linkage, guaiacylglycerol-β-guaiacyl ether (GGE). We identified several parallel causal mutations, including a single nucleotide polymorphism in the promoter of an uncharacterized gene that roughly doubled the growth yield with GGE. We characterized the associated enzyme and demonstrated that it oxidizes an intermediate in GGE catabolism, β-hydroxypropiovanillone, to vanilloyl acetaldehyde. Identification of this enzyme and its key role in GGE catabolism furthers our understanding of catabolic pathways for lignin-derived aromatic compounds.

59 BASIC BIOLOGICAL SCIENCES↗

Modeling of hepatitis B virus infection spread in primary human hepatocytes

ABSTRACT Chronic hepatitis B virus (HBV) infection poses a significant global health threat, causing severe liver diseases including cirrhosis and hepatocellular carcinoma. We characterized HBV DNA kinetics in primary human hepatocytes (PHHs) over 32 days post-inoculation (p.i.) and modified ourin-vivoagent-based modeling (ABM) to gain insights into the HBV lifecycle and spreadin vitro. Parallel PHH cultures were mock-treated or treated with HBV entry inhibitor Myr-preS1 (6.25 µg/mL) was initiated 24 h p.i. In untreated PHH, three viral DNA kinetic patterns were identified: (i) an initial decline, followed by (ii) rapid amplification and (iii) slower amplification/accumulation. In the presence of Myr-preS1, viral DNA and infected cell numbers in phase 3 were effectively blocked, with minimal to no increase. This suggests that phase 2 represents viral amplification in initially infected cells, while phase 3 corresponds to viral spread to naïve cells. The ABM reproduced well the HBV kinetic patterns observed and predicted that the viral eclipse phase lasts between 18 and 38 h. After the eclipse phase, the viral production rate increased over time, starting with a slow production cycle of 1 virion per day, which gradually accelerated to 1 virion per hour after 3 days. Approximately 4 days later, virion production reached a steady state production rate of 4 virions/h. The estimated median efficacy of Myr-preS1 in blocking HBV spread was 91% (range: 90–92%). The HBV kinetics and the predicted estimates of the HBV eclipse phase duration and HBV production cycles in PHH are similar to those predicted in uPA/SCID mice with human livers. IMPORTANCE While primary human hepatocytes (PHHs) are the most physiologically relevant culture system for studying HBV infectionin vitro, a comprehensive understanding of HBV infection kinetics and spread in PHH is lacking. In this study, we characterize HBV viral kinetics and modify ourin vivoagent-based modeling (ABM) to provide quantitative insights into the HBV production cycle and viral spread in PHH. The ABM provides an estimate of the HBV eclipse phase duration, HBV production cycles, and Myr-preS1 efficacy in blocking HBV spread in PHH. The results resemble those predicted in uPA/SCID mice with human livers, demonstrating that estimated HBV infection kinetic parameters in PHHin vitromirror those observed in thein vivoHBV infection chimeric mouse model.

Virology↗

Orthogonal chemical genomics approaches reveal genomic targets for increasing anaerobic chemical tolerance in Zymomonas mobilis

Genetically engineered microbes have the potential to increase efficiency in the bioeconomy by overcoming growth-limiting production stress. Screens of gene perturbation libraries against production stressors can identify high-value engineering targets, but follow-up experiments needed to guard against false positives are slow and resource-intensive. In principle, the use of orthogonal gene perturbation approaches could increase recovery of true positives over false positives because the strengths of one technique compensate for the weaknesses of the other, but, in practice, two parallel screens are rarely performed at the genome scale. Here, we screen genome-scale CRISPRi (CRISPR interference) knockdown and transposon insertion libraries of the bioenergy-relevant Alphaproteobacterium, Zymomonas mobilis, against growth inhibitors commonly found in deconstructed plant material. Integrating data from the two gene perturbation techniques, we established an approach for defining engineering targets with high specificity. This allowed us to identify all known genes in the cytochrome bc1 and cytochrome c synthesis pathway as potential targets for engineering resistance to phenolic acids under anaerobic conditions, a subset of which we validated using precise gene deletions. Strikingly, this finding is specific to the cytochrome bc1 and cytochrome c pathway and does not extend to other branches of the electron transport chain. We further show that exposure of Z. mobilis to ferulic acid causes substantial remodeling of the cell envelope proteome, as well as the downregulation of TonB-dependent transporters. Our work provides a generalizable strategy for identifying high-value engineering targets from gene perturbation screens that is broadly applicable.

CRISPRi↗

Consistent Second Moment Methods with Scalable Linear Solvers for Radiation Transport

Second moment methods (SMMs) are developed that are consistent with the discontinuous Galerkin spatial discretization of the discrete ordinates (or S\(_N\)) transport equations. The low-order (LO) diffusion system of equations is discretized with fully consistent P\(_1\), local discontinuous Galerkin (LDG), and interior penalty (IP) methods. A discrete residual approach is used to derive SMM correction terms that make each of the LO systems consistent with the high-order discretization. We show that the consistent methods are more accurate and have better solution quality than independently discretized LO systems, that they preserve the diffusion limit, and that the LDG and IP consistent SMMs can be scalably solved in parallel on a challenging, multimaterial benchmark problem.

97 MATHEMATICS AND COMPUTING↗

Fast and Accurate Intersections on a Sphere

We introduce a fast, high-precision algorithm for calculating intersections between great circle arcs and lines of constant latitude on the unit sphere. We first propose a simplified intersection point formula with improved speed and numerical robustness over the ones traditionally implemented in geoscience software. We then show how algorithms based on the concept of error-free transformations (EFT) can be applied to evaluate this formula within a relative error bound that is on the order of machine precision. Here, we demonstrate that, with a vectorized and parallelized implementation, this enhanced accuracy is achieved with no compute time overhead compared to a direct calculation in hardware floating point, making our algorithm suitable for performance-sensitive applications like regridding of high-resolution climate data. In contrast, evaluating our formula using high-precision data types like quadruple precision and arbitrary precision, or using the robust intersection computation routines from the Computational Geometry Algorithms Library, leads to significant computational overhead, especially since these alternatives inhibit vectorization. More generally, our work demonstrates how EFT techniques can be combined and extended to implement nontrivial geometric calculations with high accuracy and speed.

Environmental sciences↗

Operational experience and R&D results using the Google Cloud for High-Energy Physics in the ATLAS experiment

The ATLAS experiment at CERN relies on a Worldwide Distributed Computing Grid infrastructure to support its physics program at the Large Hadron Collider. ATLAS has integrated cloud computing resources to complement its Grid infrastructure and conducted an R&D program on Google Cloud Platform. These initiatives leverage key features of commercial cloud providers: lightweight configuration and operation, elasticity and availability of diverse infrastructures. Here this paper examines the seamless integration of cloud computing services as a conventional Grid site within the ATLAS workflow management and data management systems, while also offering new setups for interactive, parallel analysis. It underscores pivotal results that enhance the on-site computing model and outlines several R&D projects that have benefited from large-scale, elastic resource provisioning models. Furthermore, this study discusses the impact of cloud-enabled R&D projects in three domains: accelerators and AI/ML, ARM CPUs and columnar data analysis techniques.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Synthesis of (oxo)chlorin dimers chelated with thallium(III)

Two target dimers have been prepared for fundamental studies of hole/electron transfer. Metalation with thallium(III) enables clocking of the rate of hole/electron transfer between the two macrocycles. Each dimer contains a diphenylethyne linker joining two identical hydroporphyrins (chlorin or oxochlorin). The linker is substituted at the 4,4′-positions whereas each (oxo)chlorin is joined at the 10-position. Each (oxo)chlorin is equipped with a gem-dimethyl group at the 18-position to stabilize the hydroporphyrin chromophore toward adventitious dehydrogenation and a 3,5-di-tert-butyl group at the 5-position to achieve increased solubilization in organic media. The dimers parallel a prior set of diphenylethyne-linked (oxo)chlorin constructs containing zinc-free base, zinc-zinc, and copper-copper metalation states that have been examined in studies of electronic communication. The building block (oxo)chlorins for preparing the thallium-containing dimers have been prepared in quantities of 32–404 mg, a scale up to 14-fold larger than previously. Thallation of the free base (oxo)chlorin dimers was achieved with excess TlCl 3 ⋅4H 2 O in CH 2 Cl 2 /CH 3 OH (3–4:1) upon overnight reaction at room temperature. The long-wavelength (Q[Formula: see text] absorption band of (oxo)chlorins lies between that of the zinc(II) and free base counterparts. Absorption spectral comparisons are provided of the thallium(III) and free base (oxo)chlorin monomers and dimers.

Chemistry↗

Measuring Local Turbulence Along the Optical Path: Multi-Beam Optical Seeing Sensor

Deflection of light along the optical path is a major source of image degradation for ground-based telescopes. Methods have been developed to measure upper atmospheric seeing based on models of the turbulence in the atmosphere, but due to boundary conditions, transmission within telescope enclosures is more complex. The Multi-beam Optical Seeing Sensor (MOSS) directly measures the component of the image quality degradation from inhomogeneity of the index of refraction within the telescope dome. MOSS outputs four near-parallel beams of light that travel along the optical path and are imaged by the telescope’s detector, landing like starlight on the telescope’s focal plane. By using a strobed light source, we can ‘freeze’ the instantaneous index variations transverse to the optical path. This system captures both ‘dome’ and ‘mirror’ seeing. Through plotting the standard deviation of differential motion between pairs of beams, MOSS enables characterization of the length scale of turbulence within the dome. The temporal coherence of temperature gradients can be probed with different pulse lengths, and the spatial coherence by comparing pairs at different separations across the aperture of the telescope. Optical path turbulence measurements, alongside other telemetry metrics, will guide thermal and airflow management to optimize image quality. A MOSS prototype was installed in the 1.2[Formula: see text]m Auxiliary Telescope (AuxTel) at the Vera C. Rubin Observatory in Chile, and preliminary data constrain the optical path turbulence with a lower bound of 1.4 arcsec. The optical path turbulence varied throughout the night of observing.

Astronomical seeing↗

AGS-GNN: Attribute-guided Sampling for Graph Neural Networks

We propose AGS-GNN, a novel attribute-guided sampling algorithm for Graph Neural Networks (GNNs) that exploits node features and connectivity structure of a graph while simultaneously adapting for both homophily and heterophily in graphs. (In homophilic graphs vertices of the same class are more likely to be connected, and vertices of different classes tend to be linked in heterophilic graphs.) While GNNs have been successfully applied to homophilic graphs, their application to heterophilic graphs remains challenging. The best-performing GNNs for heterophilic graphs do not fit the sampling paradigm, suffer high computational costs, and are not inductive. We employ samplers based on feature-similarity and feature-diversity to select subsets of neighbors for a node, and adaptively capture information from homophilic and heterophilic neighborhoods using dual channels. Currently, AGS-GNN is the only algorithm that we know of that explicitly controls homophily in the sampled subgraph through similar and diverse neighborhood samples. For diverse neighborhood sampling, we employ submodularity, which was not used in this context prior to our work. The sampling distribution is pre-computed and highly parallel, achieving the desired scalability. Using an extensive dataset consisting of 35 small (<=100K nodes) and large (>100K nodes) homophilic and heterophilic graphs, we demonstrate the superiority of AGS-GNN compare to the current approaches in the literature. AGS-GNN achieves comparable test accuracy to the best-performing heterophilic GNNs, even outperforming methods using the entire graph for node classification. AGS-GNN also converges faster compared to methods that sample neighborhoods randomly, and can be incorporated into existing GNN models that employ node or graph sampling.

artificial intelligence↗

Realistic Cost to Execute Practical Quantum Circuits using Direct Clifford+T Lattice Surgery Compilation

We report a resource estimation pipeline that explicitly compiles quantum circuits expressed using the Clifford+T gate set into a surface code lattice surgery instruction set. The cadence of magic state requests from the compiled circuit enables the optimization of magic state distillation and storage requirements in a post-hoc analysis. To compile logical circuits into lattice surgery operations, we build upon the open-source Lattice Surgery Compiler. The revised compiler operates in two stages: the first translates logical gates into an abstract, layout-independent instruction set; the second compiles these into local lattice surgery instructions that are allocated to hardware tiles according to a specified resource layout. The second stage retains logical parallelism while avoiding resource contention in the fault-tolerant layer, aiding realism. Additionally, users can specify dedicated tiles at which magic states are replenished, enabling resource costs from the logical computation to be considered independently from magic state distillation and storage. We demonstrate the applicability of our pipeline to large practical quantum circuits by providing resource estimates for the ground state estimation of molecules. Finally, we find that variable magic state consumption rates in real circuits can cause the resource costs of magic state storage to dominate unless production is varied to suit.

97 MATHEMATICS AND COMPUTING↗

TrioSim: A Lightweight Simulator for Large-Scale DNN Workloads on Multi-GPU Systems

Deep Neural Networks (DNNs) have become increasingly capable of performing tasks ranging from image recognition to content generation. The training and inference of DNNs heavily rely on GPUs, as GPUs' massively parallel architecture delivers extremely high computing capability. With the growing complexity of DNNs and the size of training datasets, training DNNs with a large number of GPUs is becoming a prevalent strategy. Researchers have been exploring how to design software and hardware systems for GPU farms to achieve the best utilization, efficiency, and DNN accuracy during training or inference. However, when designing and deploying such systems, designers usually rely on testing on physical hardware platforms equipped with many GPUs, incurring high costs that are almost prohibitive for system designers to test different configurations and designs, even for highly resourceful companies. While an alternative solution is to test on GPU simulators, they are often too slow for these l

Li, Ying [William & Mary, Williamsburg, VA, USA] (↗

Studying CPU and memory utilization of applications on Fujitsu A64FX and Nvidia Grace Superchip

ARM-based manycore CPU architectures are well-positioned to provide the rising memory throughput requirements of modern data intensive scientific applications in High Performance Computing (HPC). The Fujitsu A64FX CPU platform is based on the ARM v8.2A architecture, and is the processor of the flagship Japanese supercomputer - "Fugaku", which was previously ranked as the #1 supercomputer in the world according to the Top500 list. The Nvidia Grace superchip features 144 Neoverse V2 cores based on the ARMv9 architecture with 4x128b SVE2, providing exceptional computational power. The chip supports up to 480GB of memory, making it ideal for AI, machine learning, and scientific computing workloads. In this paper, we conduct a thorough performance exploration of a variety of parallel bandwidth-sensitive benchmarks and applications compiled with the native Fujitsu compiler on a Fugaku A64FX compute node and ARM (LLVM) Compiler on an NVIDIA Grace superchip compute node, engaging all the computational cores per cluster using OpenMP multithreading (assuming the cores can drive the available bandwidth). Our ultimate goals are to study the resource utilization of scientific applications and benchmarks on A64FX and Grace superchip, considering graph application scenarios ( GAP Benchmark suite) and eleven appli- cation proxies from the Rodinia heterogeneous benchmark suite (considering domains such as Data Mining, Bioinformatics, Fluid Dynamics, Pattern Recognition, etc.). Through exhaustive performance monitoring, we quantify the resource utilization of diverse OpenMP-based HPC applications on both the Fujitsu A64FX and the Nvidia Grace Superchip platforms.

benchmarking, Performance Analysis, High performan↗

ChatHPC: Building the Foundations for a Productive and Trustworthy AI-Assisted HPC Ecosystem

ChatHPC democratizes large language models for the high-performance computing (HPC) community by providing the infrastructure, ecosystem, and knowledge needed to apply modern generative AI technologies to rapidly create specific capabilities for critical HPC components while using relatively modest computational resources. Our divide-and-conquer approach focuses on creating a collection of reliable, highly specialized, and optimized AI assistants for HPC based on the cost-effective and fast Code Llama fine-tuning processes and expert supervision. We target major components of the HPC software stack, including programming models, runtimes, I/O, tooling, and math libraries. Thanks to AI, ChatHPC provides a more productive HPC ecosystem by boosting important tasks related to portability, parallelization, optimization, scalability, and instrumentation, among others. With relatively small datasets (on the order of KB), the AI assistants, which are created in a few minutes by using one node with two NVIDIA H100 GPUs and the ChatHPC library, can create new capabilities with Meta’s 7-billion parameter Code Llama base model to produce high-quality software with a level of trustworthiness of up to 90% higher than the 1.8-trillion parameter OpenAI ChatGPT-4o model for critical programming tasks in the HPC software stack.

Young, Aaron [ORNL] (ORCID:0000000254484667)↗

Stability-preserving Lossy Compression for Large-scale Partial Differential Equations

Checkpoint/Restart (C/R) strategies are vital for fault tolerance in PDE-based scientific simulations, yet traditional checkpointing incurs significant I/O overhead. Lossy compression offers a scalable solution by reducing checkpoint data size, but conventional methods often lack control over physical invariants (e.g., energy), leading to instability such as oscillations or divergence in Partial Differential Equations (PDE) systems. This paper introduces a stability-preserving compression approach tailored for PDE simulations by explicitly controlling kinetic and potential energy perturbations to ensure stable restarts. Extensive experiments conducted across diverse PDE configurations demonstrate that our method maintains numerical stability with minimal error magnification—even across multiple checkpoint-restart cycles—outperforming state-of-the-art lossy compressors. Parallel evaluations on the Frontier supercomputer show up to 8.4× improvement in checkpoint write performance and 6.3× in read performance, while maintaining relative L2 errors ∼ 2e-6 throughout continued simulation. These results provide practical guidance for balancing compression accuracy, stability, and computational efficiency in large-scale PDE applications.

Gong, Qian [ORNL] (ORCID:0000000235704142)↗

ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling

Sparse observations and coarse-resolution climate models limit effective regional decision-making, underscoring the need for robust downscaling. However, existing AI methods struggle with generalization across variables and geographies and are constrained by the quadratic complexity of Vision Transformer (ViT) self-attention. We introduce ORBIT-2, a scalable foundation model for global, hyper-resolution climate downscaling. ORBIT-2 incorporates two key innovations: (1) Residual Slim ViT (Reslim), a lightweight architecture with residual learning and Bayesian regularization for efficient, robust prediction; and (2) TILES, a tile-wise sequence scaling algorithm that reduces self-attention complexity from quadratic to linear, enabling long-sequence processing and massive parallelism. ORBIT-2 scales to 10 billion parameters across 65,536 GPUs, achieving up to 4.1 ExaFLOPS sustained throughput and 74–98% strong scaling efficiency. It supports downscaling to 0.9 km global resolution and processes sequences up to 4.2 billion tokens. On 7 km resolution benchmarks, ORBIT-2 achieves high accuracy with R2 scores in range of 0.98–0.99 against observation data.

Wang, Xiao [ORNL] (ORCID:0000000165451943)↗

DIMPLES: Distributed Influence Maximization for Pandemic pLanning on Exascale Systems

We study exascale parallel algorithms for the selection of intervention or monitoring strategies in massive realistic socio-technical networks through scalable Influence Maximization (InfMax) algorithms. We employ novel techniques to enable efficient scaling on up to 8k nodes of OLCF Frontier, with 65k AMD GPUs and 458k AMD CPU cores. Current state-of-the-art InfMax tools are limited to networks with only a few million actors (vertices) and a few hundred million interactions (edges). By overcoming these limitations, we show that our approach is capable of processing a realistic social contact network of the United States with 285 million nodes and about 8 billion edges. This two orders-of-magnitude improvement over the previous state-of-the-art is obtained by leveraging algorithmic advancements for the InfMax problem and designing several problem-specific approaches to overlap communication with computation, improve GPU efficiency, and lower the application’s memory requirements. We evaluate strong scaling for computing 10k most influential seeds using up to 8k nodes of an exascale system, and weak scaling from 128 to 8k system nodes for seed sets ranging from 625 to 40k seeds. We achieve the fastest-known runtime of 25 minutes while performing 48 million diffusion simulations totaling 2.31 petabytes to identify 40k influential seeds using 8k nodes, and take 5.75 minutes to identify 10k seeds while using 4k nodes.

Minutoli, Marco [Pacific Northwest National Labora↗

I/O in Machine Learning Applications on HPC Systems: A 360-degree Survey

Growing interest in Artificial Intelligence (AI) has resulted in a surge in demand for faster methods of Machine Learning (ML) model training and inference. This demand for speed has prompted the use of high performance computing (HPC) systems that excel in managing distributed workloads. Because data is the main fuel for AI applications, the performance of the storage and I/O subsystem of HPC systems is critical. In the past, HPC applications accessed large portions of data written by simulations or experiments or ingested data for visualizations or analysis tasks. ML workloads perform small reads spread across a large number of random files. This shift of I/O access patterns poses several challenges to modern parallel storage systems. In this paper, we survey I/O in ML applications on HPC systems, and target literature within a 6-year time window from 2019 to 2024. We define the scope of the survey, provide an overview of the common phases of ML, review available profilers and benchmarks, examine the I/O patterns encountered during offline data preparation, training, and inference, and explore I/O optimizations utilized in modern ML frameworks and proposed in recent literature. Lastly, we seek to expose research gaps that could spawn further R&D.

97 MATHEMATICS AND COMPUTING↗