Search NASASearch

SEARCH · Search NASA

Results for “parallel cluster”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Characterization of DESI fiber assignment incompleteness effect on 2-point clustering and mitigation methods for DR1 analysis

We present an in-depth analysis of the fiber assignment incompleteness in the Dark Energy Spectroscopic Instrument (DESI) Data Release 1 (DR1). This incompleteness is caused by the restricted mobility of the robotic fiber positioner in the DESI focal plane, which limits the number of galaxies that can be observed at the same time, especially at small angular separations. As a result, the observed clustering amplitude is suppressed in a scale-dependent manner, which, if not addressed, can severely impact the inference of cosmological parameters. We discuss the methods adopted for simulating fiber assignment on mocks and data. In particular, we introduce the fast fiber assignment (FFA) emulator, which was employed to obtain the power spectrum covariance adopted for the DR1 full-shape analysis. We present the mitigation techniques, organised in two classes: measurement stage and model stage. We then use high fidelity mocks as a reference to quantify both the accuracy of the FFA emulator and the effectiveness of the different measurement-stage mitigation techniques. This complements the studies conducted in a parallel paper for the model-stage techniques, namely the θ-cut approach. We find that pairwise inverse probability (PIP) weights with angular upweighting recover the “true” clustering in all the cases considered, in both Fourier and configuration space. Notably, we present the first ever power spectrum measurement with PIP weights from real data.

cosmological simulations

HARD: A performance portable radiation hydrodynamics code based on FleCSI framework

Hydrodynamics And Radiation Diffusion (HARD) is an open-source application for high-performance simulations of compressible hydrodynamics with radiation-diffusion coupling. Built on the FleCSI (Bergen et al., 2021 [1]) (Flexible Computational Science Infrastructure) framework, HARD expresses its computational units as tasks whose execution can be orchestrated by multiple back-end runtimes, including Legion (Bauer et al., 2012 [2]), MPI (Forum, 1994 [3]), and HPX (Kaiser et al., 2020 [4]). Node-level parallelism is handled through Kokkos (Edwards et al., 2014 [5]), providing a single-source, portable code base that runs efficiently on laptops, small homogeneous clusters, and the largest heterogeneous supercomputers currently available. To ensure scientific reliability, HARD includes a regression test suite that automatically reproduces canonical verification problems such as the Sod and LeBlanc shock tubes, and the Sedov blast wave, comparing numerical solutions against known analytical results. The project is distributed under an OSI-approved license, hosted on GitHub, and accompanied by reproducible build scripts and continuous integration workflows. This combination of performance portability, verification infrastructure, and community-focused development makes HARD a sustainable platform for advancing radiation hydrodynamics research across multiple domains.

97 MATHEMATICS AND COMPUTING

Thermodynamics and collisionality in firehose-susceptible high- β plasmas

We study the evolution of collisionless plasmas that, due to their macroscopic evolution, are susceptible to the firehose instability, using both analytic theory and hybrid-kinetic particle-in-cell simulations. We establish that, depending on the relative magnitude of the plasma β, the characteristic time scale of macroscopic evolution and the ion-Larmor frequency, the saturation of the firehose instability in high-β plasmas can result in three qualitatively distinct thermodynamic (and electromagnetic) states. By contrast with the previously identified ‘ultra-high-beta’ and ‘Alfvén-inhibiting’ states, the newly identified ‘Alfvén-enabling’ state, which is realised when the macroscopic evolution time τ exceeds the ion-Larmor frequency by a β-dependent critical parameter, can support linear Alfvén waves and Alfvénic turbulence because the magnetic tension associated with the plasma’s macroscopic magnetic field is never completely negated by anisotropic pressure forces. We characterise these states in detail, including their saturated magnetic-energy spectra. The effective collision operator associated with the firehose fluctuations is also described; we find it to be well approximated in the Alfvén-enabling state by a simple quasi-linear pitch-angle scattering operator. The box-averaged collision frequency is ν eff ∼ β/τ, in agreement with previous results, but certain subpopulations of particles scatter at a much larger (or smaller) rate depending on their velocity in the direction parallel to the magnetic field. Our findings are essential for understanding low-collisionality astrophysical plasmas including the solar wind, the intracluster medium of galaxy clusters and black hole accretion flows. We show that all three of these plasmas are in the Alfvén-enabling regime of firehose saturation and discuss the implications of this result.

astrophysical plasmas

Function, Structure, and Regulation of Nitrogen Fixation-like Metalloproteins for Nitrogen, Energy, Carbon, and Sulfur Metabolism

Nitrogenases (N 2 ases) and nitrogen fixation-like (NFL) systems play distinct roles in nitrogen, carbon, sulfur, and energy metabolism based on their fundamental differences in structure and metallocofactor identity. As new NFL systems have recently been identified and characterized, striking parallels and differences compared to N 2 ase structure, catalysis, and regulation have emerged. NFL systems use metallocofactors that span from simple [4Fe-4S] clusters to complex clusters akin to FeMo-co, previously only thought to occur in N 2 ase. This review describes the present state of knowledge on the function, structure, catalytic mechanisms, and regulation of NFL systems that perform distinct biological roles across all three domains of life. Recent advancements in N 2 ase spectroscopic techniques for probing metallocofactor structure and electronic states guide current and future work on how each NFL system catalyzes its specific biological reaction(s). Key knowledge gaps and needed areas of research for uncovering the specific metallocofactors and structural motifs that are at the heart of NFL system reaction specificity, along with how these systems are regulated, are discussed.

Bacteria

A Versatile Simulated Data Transport Layer for in Situ Workflows Performance Evaluation

In situ processing does not only allow scientific applications to face the explosion in data volume and velocity but also to address the time constraints of many simulation-analysis workflows by providing scientists with early insights about their applications at runtime. Multiple frameworks implement the concept of a data transport layer (DTL) to enable such in situ workflows. These tools are very versatile, directly or indirectly access the data generated on the same node, another node of the same compute cluster, or a completely distinct node, and allow data publishers and subscribers to run on the same computing resources or not. This versatility puts on researchers the onus of taking key decisions related to resource allocation and how to transport data to ensure the most efficient execution of their in situ workflows. However, domain scientists and workflow practitioners lack the appropriate tools to assess the respective performance of particular design and deployment options. In this paper we introduce a versatile simulated DTL designed to provide researchers with insights on the respective performance of different execution scenarios of in situ workflows. This open-source, standalone library builds on the SimGrid toolkit and can be linked to any SimGrid-based simulator. It facilitates the evaluation of the performance behavior, at scale, of different data transport configurations and the study of the effects of resource allocation strategies. We demonstrate the scalability, versatility, and accuracy of this simulated DTL by reproducing the execution of two synthetic benchmarks and of a real-world in situ workflow composed of an MPI application and a parallel data analysis. Results of simulations run on a single core show that the proposed library can simulate the interactions of tens of thousands of simulated processes deployed on two interconnected commodity clusters in a few seconds, and the execution by a thousand simulated processes of an in situ workflow in less than three minutes.

Suter, Fred [ORNL] (ORCID:0000000319021955)

Quantum Electrodynamics Coupled-Cluster at Scale: High-Performance Implementation for Complex Systems

Coupled-cluster theory (CC) is a highly accurate and versatile method for simulating complex interactions within quantum systems. The extension of CC theory to model mixed electron-photon processes with quantum electrodynamics (QED) has improved our capability to predict cavity-modified chemistry, a field where photons are used as cost-effective and eco-friendly alternatives to catalyze/inhibit chemical reactions. However, calculations with CC methods, even without incorporating QED effects, are often prohibitively expensive. Simulations of larger systems require scalable infrastructures that exist for traditional CC methods but not for QED-CC methods. As such, we present a GPU-enabled, high-performance, open-source implementation of the quantum electrodynamics coupled-cluster method with single and double excitations (QED-CCSD) within the ExaChem quantum chemistry software package. ExaChem relies on the Tensor Algebra for Many-body Methods (TAMM) infrastructure: a parallel heterogeneous tensor library designed to achieve scalable performance on modern heterogeneous supercomputing platforms. Furthermore, we discuss theoretical foundations, algorithmic details, and numerical benchmarks to showcase the larger systems that ExaChem can simulate and how the integration of photonic degrees-of-freedom alters their ground-state properties.

Basis sets

Exploring Anomalous Photoelectron Angular Distributions in the Photoelectron Spectra of Gd 3 O 3 – : Study of Gd 3 O 2 – and Gd 3 O 3 – Using Photoelectron Spectroscopy and Density Functional Theory Calculations

Anion photoelectron (PE) spectra of lanthanide oxide clusters obtained previously have exhibited anomalous photoelectron angular distributions which were attributed to strong PE–valence electron (PEVE) interactions. Here, to further explore this effect, we have obtained the PE spectra of Gd 3 O 2 – and Gd 3 O 3 – , two clusters that have similarly complex electronic structures but contrasting symmetries. The spectra exhibit manifolds of detachment transitions at similar binding energies in a 0.5 eV window of energy. The electron affinity of Gd 3 O 2 is measured to be 1.29 ± 0.05 eV, and that of Gd 3 O 3 is 1.31 ± 0.05 eV. As seen in previous studies on lanthanide oxide cluster anions in lower than conventional oxidation states, transitions in spectra obtained lower photon energies are more congested than those obtained with higher photon energy, a signature of strong PEVE interactions. While the detachment transitions have predominantly parallel photoelectron angular distributions (PAD), the PAD varies across the manifold of transitions in the PE spectrum of Gd 3 O 3 – in a way that suggests four different subgroups of transitions. Results of calculations on Gd 3 O 2 – suggest kite or V-shape structures with antiferromagnetic coupling between one of the 4f 7 subshells with the two others. Calculations on Gd 3 O 3 – more definitively point to ring structures with a nearly isoenergetic ferromagnetically coupled high spin (24-tet) state and a dectet state in which one of the 4f 7 subshells is antiferromagnetically coupled with the other two. Taking these results as qualitative, we propose that strong mixing between the unperturbed states predicted computationally leads to overlapping transitions with different PADs.

anions

A biosynthetic gene cluster for three post-chorismate pathways in Arabidopsis

Chorismate is a branch-point metabolite in the biosynthesis of aromatic amino acids, vitamins, antibiotics and various other aromatic products in bacteria, fungi and plants. Although 13 chorismate-utilizing enzymes have been identified in bacteria, only 6 have been described in plants, where an estimated 30% of all photosynthetically fixed carbon passes through chorismate. Here, in this study, we describe a biosynthetic gene cluster (BGC) consisting of five core genes, including two reductases, two methyltransferases and one glucosyltransferase. Genetic and biochemical evidence shows that these five enzymes collectively give rise to three biosynthetic pathways, each originating from chorismate: two parallel pathways produce a class of non-aromatic, isomeric compounds abundant in the roots of Arabidopsis thaliana, whereas the third pathway produces methylated and glucosylated chorismate derivatives that subsequently react non-enzymatically with glutathione. Genome analysis revealed that variants of this BGC are present in some but not all species in the Brassicaceae family. Taken together, our study uncovered a BGC, containing three chorismate-utilizing enzymes, that controls three distinct post-chorismate pathways in A. thaliana. This work not only advances our understanding of carbon flow in this model plant but also highlights that the biochemical complexity encoded by plant BGCs is greater than previously appreciated.

Peng, Meng [Ghent Univ. (Belgium); Flemish Institu

Studying CPU and memory utilization of applications on Fujitsu A64FX and Nvidia Grace Superchip

ARM-based manycore CPU architectures are well-positioned to provide the rising memory throughput requirements of modern data intensive scientific applications in High Performance Computing (HPC). The Fujitsu A64FX CPU platform is based on the ARM v8.2A architecture, and is the processor of the flagship Japanese supercomputer - "Fugaku", which was previously ranked as the #1 supercomputer in the world according to the Top500 list. The Nvidia Grace superchip features 144 Neoverse V2 cores based on the ARMv9 architecture with 4x128b SVE2, providing exceptional computational power. The chip supports up to 480GB of memory, making it ideal for AI, machine learning, and scientific computing workloads. In this paper, we conduct a thorough performance exploration of a variety of parallel bandwidth-sensitive benchmarks and applications compiled with the native Fujitsu compiler on a Fugaku A64FX compute node and ARM (LLVM) Compiler on an NVIDIA Grace superchip compute node, engaging all the computational cores per cluster using OpenMP multithreading (assuming the cores can drive the available bandwidth). Our ultimate goals are to study the resource utilization of scientific applications and benchmarks on A64FX and Grace superchip, considering graph application scenarios ( GAP Benchmark suite) and eleven appli- cation proxies from the Rodinia heterogeneous benchmark suite (considering domains such as Data Mining, Bioinformatics, Fluid Dynamics, Pattern Recognition, etc.). Through exhaustive performance monitoring, we quantify the resource utilization of diverse OpenMP-based HPC applications on both the Fujitsu A64FX and the Nvidia Grace Superchip platforms.

benchmarking, Performance Analysis, High performan

A Melanoma Brain Metastasis CTC Signature and CTC:B-cell Clusters Associate with Secondary Liver Metastasis: A Melanoma Brain–Liver Metastasis Axis

Melanoma brain metastasis is linked to dismal prognosis and low overall survival and is detected in up to 80% of patients at autopsy. Circulating tumor cells (CTC) are the smallest functional units of cancer and precursors of fatal metastasis. We previously used an unbiased multilevel approach to discover a unique ribosomal protein large/small subunit (RPL/RPS) CTC gene signature associated with melanoma brain metastasis. In this study, we hypothesized that CTC-driven melanoma brain metastasis secondary metastasis (“metastasis of metastasis” per clinical scenarios) has targeted organ specificity for the liver. We injected parallel cohorts of immunodeficient and newly developed humanized NBSGW (huNBSGW) mice with cells from CTC-derived melanoma brain metastasis to identify secondary metastatic patterns. We found the presence of a melanoma brain–liver metastasis axis in huNBSGW mice. Furthermore, RNA sequencing analysis of tissues showed a significant upregulation of the RPL/RPS CTC gene signature linked to metastatic spread to the liver. Additional RNA sequencing of CTCs from huNBSGW blood revealed extensive CTC clustering with human B cells in these mice. CTC:B-cell clusters were also upregulated in the blood of patients with primary melanoma and maintained either in CTC-driven melanoma brain metastasis or melanoma brain metastasis CTC–derived cells promoting liver metastasis. CTC-generated tumor tissues were interrogated at single-cell gene and protein expression levels (10x Genomics Xenium and HALO spatial biology platforms, respectively). Collectively, our findings suggest that heterotypic CTC:B-cell interactions can be critical at multiple stages of metastasis.

60 APPLIED LIFE SCIENCES

Architecture, catalysis and regulation of methylthio-alkane reductase for bacterial sulfur acquisition from volatile organic compounds

Bacteria utilize methylthio-alkane reductase (MAR) to acquire sulfur from volatile organic sulfur compounds. Reductive cleavage of methylthio-ethanol and dimethylsulfide liberates methanethiol for methionine synthesis and concomitantly releases ethylene and methane, respectively. Here we show that the native MAR of Rhodospirillum rubrum is a two-component system composed of a MarH ATP-dependent reductase and a MarDK catalytic core, whose architecture parallels nitrogenase. MarS complexes with MarDK to downregulate MAR activity during cellular sulfate influx, based on chromatographic and activity analyses. MarDK possesses complex metallocofactors resembling, but not identical to, nitrogenase P- and iron-only M-clusters, designated as mar1 and mar2 clusters based on metal, spectroscopic and mutagenesis analyses. They exhibit electronic features similar to the iron-only nitrogenase under turnover and, remarkably, are matured by MarB or nitrogenase NifB, resulting in maturase-dependent activity profiles. Altogether, this suggests a broader scope of reactivity, mechanisms and regulation in microbial metabolism for the nitrogenase-like family of enzymes than previously considered.

09 BIOMASS FUELS

Scale-up Unlearnable Examples Learning with High-performance Computing

Recent advancements in AI models, like ChatGPT, are structured to retain user interactions, which could inadvertently include sensitive healthcare data. In the healthcare field, particularly when radiologists use AI-driven diagnostic tools hosted on online platforms, there is a risk that medical imaging data may be repurposed for future AI training without explicit consent, spotlighting critical privacy and intellectual property concerns around healthcare data usage. Addressing these privacy challenges, a novel approach known as Unlearnable Examples (UEs) has been introduced, aiming to make data unlearnable to deep learning models. A prominent method within this area, called Unlearnable Clustering (UC), has shown improved UE performance with larger batch sizes but was previously limited by computational resources (e.g., a single workstation). To push the boundaries of UE performance with theoretically unlimited resources, we scaled up UC learning across various datasets using Distributed Data Parallel (DDP) training on the Summit supercomputer. Our goal was to examine UE efficacy at high-performance computing (HPC) levels to prevent unauthorized learning and enhance data security, particularly exploring the impact of batch size on UE’s unlearnability. Utilizing the robust computational capabilities of the Summit, extensive experiments were conducted on diverse datasets such as Pets, MedMNist, Flowers, and Flowers102. Our findings reveal that both overly large and overly small batch sizes can lead to performance instability and affect accuracy. However, the relationship between batch size and unlearnability varied across datasets, highlighting the necessity for tailored batch size strategies to achieve optimal data protection. The use of Summit’s high-performance GPUs, along with the efficiency of the DDP framework, facilitated rapid updates of model parameters and consistent training across nodes. Our results underscore the critical role of selecting appropriate batch sizes based on the specific characteristics of each dataset to prevent learning and ensure data security in deep learning applications. The source code is publicly available at https: // github. com/ hrlblab/ UE_ HPC .

Zhu, Yanfan [Vanderbilt University, Nashville, TN,

Spatially Accelerated Winding Numbers for Curved Geometry

The generalized winding number (GWN) is a scalar field that supports robust containment queries on curved geometry, including non-watertight, overlapping, and nested boundary representations. While queries can be easily parallelized over samples, direct evaluation on parametric curves and surfaces remains costly for large and complex models. Fast, state-of-the-art GWN approaches leverage a spatial index to approximate the GWN, typically coupled with a Taylor expansion which approximates the GWN contribution for far clusters of geometric primitives. However, such methods operate only on discrete inputs such as triangle meshes and point clouds, and would introduce containment errors near boundaries if applied to curved input. We extend support for fast GWN evaluation over arbitrary collections of NURBS curves in 2D and trimmed NURBS patches in 3D via a Bounding Volume Hierarchy that stores efficiently precomputed moment data in the hierarchy nodes. When querying the hierarchy, approximations for far clusters are used alongside direct evaluation for nearby NURBS primitives, achieving sub-linear complexity while preserving the geometric features in the vicinity of the query point. Central to our performance improvements is an adaptive subdivision strategy for NURBS primitives during a preprocessing phase, creating better spatial partitions while retaining the same accuracy for containment decisions as a direct evaluation. We demonstrate the performance and accuracy of our approach across a large collection of 2D and 3D datasets.

Computer science

MBX V1.2: Accelerating Data-Driven Many-Body Molecular Dynamics Simulations

The MBX software provides an advanced platform for molecular dynamics simulations, leveraging state-of-the-art MB-pol and MB-nrg data-driven many-body potential energy functions. Developed over the past decade, these potential energy functions integrate physics-based and machine-learned many-body terms trained on electronic structure data calculated at the "gold standard" coupled-cluster level of theory. Recent advancements in MBX have focused on optimizing its performance, resulting in the release of MBX v1.2. While the inherently many-body nature of MB-pol and MB-nrg ensures high accuracy, it poses computational challenges. MBX v1.2 addresses these challenges with significant performance improvements, including enhanced parallelism that fully harnesses the power of modern multicore CPUs. In conclusion, these advancements enable simulations on nanosecond time scales for condensed-phase systems, significantly expanding the scope of high-accuracy, predictive simulations of complex molecular systems powered by data-driven many-body potential energy functions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Noise-aware optimization in nominally identical manufacturing and measuring systems for high-throughput parallel workflows

Device-to-device variability in experimental noise critically impacts reproducibility, especially in automated, high-throughput systems like additive manufacturing farms. While manageable in small labs, such variability can escalate into serious risks at larger scales, such as architectural 3D printing, where noise may cause structural or economic failures. This contribution presents a noise-aware decision-making algorithm that quantifies and models device-specific noise profiles to manage variability adaptively. It uses distributional analysis and pairwise divergence metrics with clustering to choose between single-device and robust multi-device Bayesian optimization strategies. Unlike conventional methods that assume homogeneous devices or enforce generic robustness, the proposed framework explicitly determines whether shared optimization across devices is appropriate based on the degree of inter-device noise heterogeneity. This enables improved performance, reproducibility, and efficiency. An experimental case study involving three nominally identical 3D printers (same brand, model, and close serial numbers) demonstrates reduced redundancy, lower resource usage, and improved reliability, along with improved convergence stability and solution quality through the selection of the appropriate optimization strategy based on the degree of inter-device noise heterogeneity. Overall, this framework establishes a general approach for precision- and resource-aware optimization in scalable, automated experimental platforms, demonstrated here on a representative multi-device 3D printing case study.

Schenk, Christina

Stochastic GW -GPU: Rapid Quasi-Particle Energies for Molecules beyond 10,000 Atoms

StochasticGW is a code for computing accurate quasi-particle (QP) energies of molecules and material systems in the GW approximation. StochasticGW utilizes the stochastic Resolution of the Identity (sROI) technique to enable a massively parallel implementation with computational costs that scale semilinearly with system size, allowing the method to access systems with tens of thousands of electrons. Here, we introduce a new implementation, StochasticGW-GPU, for which the main bottleneck steps have been ported to GPUs and give substantial performance improvements over previous versions of the code. We showcase the new code by computing band gaps of hydrogenated silicon clusters (Si x H y ) containing up to 10,001 atoms and 35,144 electrons, and we obtain individual QP energies with a statistical precision of better than ±0.03 eV with times-to-solution of less than 1 h.

Thomas, Phillip S. [Lawrence Berkeley National Lab

Using Apptainer in a Pilot-based Distributed Workload

GlideinWMS is a pilot and pressure-based workload manager for distributed scientific computing. Many experiments like CMS and Fermilab’s Neutrino experiments use it to provision elastic clusters for their analysis and simulations, split into close to a million concurrent jobs. Most user jobs require containers, and the pilots use Apptainer to set up the desired platform. For the pilots that run as regular batch jobs, Apptainer is safer, lighter, and easier to use than other containerization solutions. Many images used by the pilots are expanded SIF images distributed via the CernVM-FS: this combination is very efficient. At Fermilab, for example, we store on GitHub Dockerfiles that mimic the platform in the worker nodes of local clusters. GitHub workflows build and push the images to Docker Hub, and a service periodically pulls and converts them to the expanded SIF images in the CernVM-FS, so the scientists can find a familiar environment everywhere. Apptainer has also been used to run services inside the pilot jobs, like benchmarks that characterize the worker node being used, or a Triton Inference Server that allows sharing a GPU with all the jobs that run in parallel on a node.

Mambelli, Marco [Fermilab] (ORCID:0000000294892681

Multi-GPU porting of a phase-change cascaded lattice Boltzmann method for three-dimensional pool boiling simulations

The Lattice Boltzmann method (LBM) has proven effective in simulating phase-change phenomena, such as melting, solidification, evaporation, and boiling. In this work, we develop a highly parallelized multi-GPU implementation of LBM for three-dimensional pool boiling simulations. The code is based on the OpenACC programming model, which enables the code to be deployed efficiently on multi-core CPUs, GPUs, and potentially other accelerators, without the need for architecture-specific rewrites. To support large-scale simulations, the domain is decomposed and distributed across multiple compute nodes using MPI. We demonstrate that the code exhibits excellent scaling properties, with ideal strong-scaling running with up to 256 GPUs on the MareNostrum5 cluster.

97 MATHEMATICS AND COMPUTING