Search NASASearch

SEARCH · Search NASA

Results for “Parallel Performance Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Coupling Metabolic Source Isotopic Pair Labeling and Genome Wide Association for Metabolite and Gene Annotation in Plants (Final Technical Report)

In this project, we applied our labeling pipeline to Arabidopsis and sorghum by feeding tissues with isotopically labeled versions of commercially available amino acids to identify all metabolite features that incorporate the label. In sorghum, we fed five accessions, sampled across the diversity of sorghum, to identify the precursor-of-origin for metabolites that vary between accessions as well as those that may be missing from a single reference genotype. This provided us with precursor-of-origin annotation for thousands of unknown metabolites. We then used GWA to map genes responsible for the synthesis of precursor-of-origin classified metabolites. For sorghum leaf and root ducible metabolites, we performed untargeted metabolomics on leaf and root tissues from 300 diverse genotyped sorghum inbred lines. The amino acid precursor-of-origin metabolite library were then used to identify the corresponding metabolites in the GWA data sets and to identify novel gene-metabolite associations. Finally, we utilized existing and newly generated sequenced EMS mutants of sorghum to validate the predicted gene-metabolite relationships that our labelling analysis identified. In parallel, we conducted similar feeding experiments in Arabidopsis to categorize metabolites based on precursor-of-origin, identify those that vary across our existing Arabidopsis metabolite GWA dataset, and identify genes required for the synthesis of each metabolite. To provide an independent test of gene annotation and pathway involvement, we tested the GWA gene-metabolite associations in Arabidopsis by analyzing the metabolic phenotypes of gene knockouts. Genes of particular interest from both sorghum and Arabidopsis were studied in detail by directly measuring the activity of the corresponding enzymes following heterologous expression. In summary, this work classified as-yet-unknown amino acid-derived metabolites and identified genes involved in their production generated through “omics” technologies. This information was used to validate gene function and identify new metabolism in Arabidopsis and sorghum.

09 BIOMASS FUELS

Studies of Quark Transport and Hadronization in Nuclei

In this project, we conducted the first measurement of di‑hadron azimuthal correlations in deep inelastic scattering (DIS) off nuclei using the CLAS detector at Jefferson Lab. Using 5 GeV electron‑beam data collected on deuterium, carbon, iron, and lead targets, we extracted di‑pion correlation functions over a broad kinematic range. The results show a monotonic broadening of the correlation peak with increasing nuclear mass, along with pronounced dependencies on the pions’ kinematics. Separately, we implemented an algorithm based on the Kalman filter that achieved the first complete alignment of the CLAS12 central tracking system. In parallel, we developed simulations, algorithms, and performance studies that informed the conceptual designs of the forward hadronic calorimeter Insert and the Zero Degree Calorimeter, both of which are now included in the ePIC detector baseline for the forthcoming Electron Ion Collider.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Randomized Preconditioned Solvers for Strong Constraint 4D-Var Data Assimilation

The Strong Constraint 4D Variational (SC-4DVAR) data assimilation method is widely used in climate and weather applications. SC-4DVAR involves solving a minimization problem to compute the maximum a posteriori estimate, which we tackle using the Gauss-Newton method. The computation of the descent direction is expensive since it involves the solution of a large-scale and potentially ill-conditioned linear system, solved using the preconditioned conjugate gradient (PCG) method. Here, to address this cost, we efficiently construct scalable preconditioners using three different randomization techniques, which all rely on a certain low-rank structure involving the Gauss-Newton Hessian. The proposed techniques come with theoretical guarantees on the condition number, and at the same time, are amenable to parallelization. We also develop an adaptive approach to estimate the sketch size and choose between the reuse or recomputation of the preconditioner. We demonstrate the performance and effectiveness of our methodology on two representative model problems—the Burgers and barotropic vorticity equation—showing a drastic reduction in both the number of PCG iterations and the number of Gauss-Newton Hessian products after including the preconditioner construction cost.

Gauss-Newton

Conceptual Designs for Irradiation Creep Testing of SiC in HFIR

Understanding irradiation creep of nuclear fuel cladding is important to properly size the initial fuel-cladding gap and understand when pellet-cladding contact is expected to occur due to a combination of fuel swelling and cladding creep-down. Irradiation creep also plays a role in relaxing stresses that develop in-pile. Silicon carbide fiber–reinforced silicon carbide matrix (SiC/SiC) composites are the leading long-term accident-tolerant fuel cladding concept for light-water reactors (LWRs). Although some limited data are available regarding irradiation creep of the individual constituents (fibers, matrix), data regarding irradiation creep of SiC/SiC composites are currently insufficient. Additional data regarding irradiation creep compliance and the rupture lifetime (combination of creep and slow crack growth) are needed to understand material limitations. This work describes the design and development of two irradiation vehicles that are being pursued for testing SiC/SiC concepts in the High Flux Isotope Reactor (HFIR). The first is a passive experiment, referred to as the PRECISE experiment, that leverages the constant coolant pressure of HFIR to compress a metallic bellows and provide a well-characterized load to drive creep in a SiC/SiC dog bone specimen. The total creep strain would be quantified post-irradiation by measuring dimensional changes of the specimen length as well as local dimensional changes within the gauge region. Non-stressed specimens would also be irradiated under the same conditions to provide an indication of dimensional changes due to radiation-induced swelling in the absence of creep. A second, more complex experiment, referred to as the INSITE experiment, is being designed in parallel that would use pneumatics to pressurize a metal bellows and linear variable differential transformers (LVDTs) to measure the specimen displacement in situ during irradiation. Such an experiment would provide significantly more data regarding the evolution of the creep compliance as a function of dose and applied stress within a single experiment but would require significantly more development time and cost to execute. The primary concern with the INSITE experiment is the accuracy, reliability, and expected lifetime of the LVDTs during irradiation at elevated temperatures. Efforts are being made to adjust the experiment design and operating procedure to limit LVDT temperatures and mitigate or otherwise compensate for uncertainties due to factors such as temperature fluctuations, creep in the surrounding structural materials, and drift of the LVDTs. This work describes the experiment designs, thermal and structural analysis that were performed to ensure that the desired temperature and stress conditions can be achieved, some initial sensitivity analyses to predict the evolution of the radiation-induced specimen displacements, and potential sources of uncertainty in the measurements. Out-of-pile testing is being performed in parallel to confirm that the test trains achieve the expected stress states in the specimens and do not result in prohibitive stress concentrators (e.g., in the grip regions) that might risk pre-mature failure. The PRECISE experiments are proceeding toward fabrication and assembly with HFIR insertion planned during fiscal year 2026. The INSITE experiment is progressing toward out-of-pile demonstrations, which will provide more conclusive evidence regarding the feasibility of executing these tests in HFIR or whether alternative displacement monitoring techniques may need to be considered.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS

A Survey on Error-Bounded Lossy Compression for Scientific Datasets

Error-bounded lossy compression has been effective in significantly reducing the data storage/transfer burden while preserving the reconstructed data fidelity very well. Many error-bounded lossy compressors have been developed for a wide range of parallel and distributed use cases for years. They are designed with distinct compression models and principles, such that each of them features particular pros and cons. In this article, we provide a comprehensive survey of emerging error-bounded lossy compression techniques. The key contribution is fourfold. (1) We summarize a novel taxonomy of lossy compression into six classic models. (2) We provide a comprehensive survey of 10 commonly used compression components/modules. (3) We summarized pros and cons of 47 state-of-the-art lossy compressors and present how state-of-the-art compressors are designed based on different compression techniques. (4) We discuss how customized compressors are designed for specific scientific applications and use-cases. We believe this survey is useful to multiple communities including scientific applications, high-performance computing, lossy compression, and big data.

Error-Bounded Lossy Compression

Hardware acceleration for HPS algorithms in two and three dimensions

We provide a flexible, open-source framework for hardware acceleration, namely massively-parallel execution on general-purpose graphics processing units (GPUs), applied to the hierarchical Poincaré–Steklov (HPS) family of algorithms for building fast direct solvers for linear elliptic partial differential equations. To take full advantage of the power of hardware acceleration, we propose two variants of HPS algorithms to improve performance on two- and three-dimensional problems. In the two-dimensional setting, we introduce a novel recomputation strategy that minimizes costly data transfers to and from the GPU; in three dimensions, we modify and extend the adaptive discretization technique of Geldermans and Gillman [1] to greatly reduce peak memory usage. We provide an open-source implementation of these methods written in JAX, a high-level accelerated linear algebra package, which allows for the first integration of a high-order fast direct solver with automatic differentiation tools. We conclude with extensive numerical examples showing our methods are fast and accurate on two- and three-dimensional problems.

Fast direct solvers

Improving I/O-aware Workflow Scheduling via Data Flow Characterization and trade-off Analysis

The scientific computing paradigm has transitioned from compute-intensive to I/O-intensive and memory-intensive in the past decade, especially when data-driven science has become common practice. Numerous empirical I/O-aware scheduling optimizations have been developed by incorporating I/O capacity and bandwidth as constraints into scheduling. Unfortunately, there is a lack of data flow (I/O) characterization tool and an understanding of trade-offs between concurrency, locality, and I/O bandwidth. To bridge the gap, this work 1) presents a set of descriptors to characterize, organize, and visualize I/O profiles, including flow size, I/O bandwidth, and operation count, which group data flows by I/O types, tasks, and files; 2) proposes an I/O Roofline model-based trade-off analysis to find the optimal trade-off between flow operational intensity, concurrency, and flow performance. The I/O descriptors generate useful insights into complicated I/O behaviors, suggesting distinct concurrency, storage, and scheduling to be used by types, tasks, and files. The proposed trade-off analysis guides scheduling decisions that generate resource assignment with the best flow parallelism. We evaluate our I/O-aware scheduling methodology on a highly I/O-intensive workflow–1000 Genomes. The experimental results demonstrate speedups of up to 2.4× compared to the state-of-the- art methods.

Guo, Luanzheng [BATTELLE (PACIFIC NW LAB)]

Process Design and Techno-Economic Analysis of the Modular Staged Pressurized Oxy-Combustion (SPOC) Power Plant for Biomass

This work describes the process design and techno-economic analysis (TEA) of the modular SPOC power plant for biomass firing and coal-biomass co-firing. Two Rankine cycles were considered: a supercritical steam cycle (242 bar, 593°C, 593°C) with 550 MWe net output and a subcritical cycle (166 bar, 566°C, 566°C) with 200 MWe net output. For both cases, 95% carbon capture was modeled, and hybrid poplar biomass was chosen to generate carbon-negative power. In addition, the supercritical 500 MWe case included a 25% biomass co-firing (carbon neutral) case. For both cycles, a 100% Powder River Basin coal firing case was used for comparison purposes. In the SPOC process, oxygen is produced via a cryogenic air separation unit (ASU) and the heat generated from the compression of air is integrated into the steam cycle and utilized for boiler feed water pre-heating. Unique to the SPOC process, the boilers are pressurized and arranged in a series-parallel configuration, with minimized flue gas recirculation. The flue gas is cooled and scrubbed in the direct-contact cooler (DCC) column, and the moisture in the flue gas is condensed, leaving the bottom of the DCC at a sufficiently high temperature such that it can be used for boiler feed water pre-heating, improving plant thermal efficiency. Following drying and purification, CO2 in the flue gas is at the purity required for storage or utilization. The performance data were obtained from process modelling via Aspen Plus®. The stream data from Aspen Plus® were used as an input for the AACE Class 5 cost study. Ultimately, the capital costs, Levelized Cost of Electricity (LCOE), and cost of CO2 captured and avoided were obtained. The HHV efficiency of the carbon negative 550 MWe supercritical SPOC case (34.8%) was clearly above those reported by NETL for the BECCS baseline cases of supercritical pulverized coal with capture (B12B, 31.5%) and the 49% biomass co-firing case with capture (PA3, 29.2%). The HHV efficiency of the carbon-negative subcritical plant is also higher than the subcritical baseline PC plant with capture (case B11B.95) presented by NETL (32% vs 29.7%). The LCOE for the SPOC 100% biomass case was similar to the LCOE for the BECCS 49% biomass with carbon capture case ($147/MWh), and the SPOC carbon neutral case LCOE was lower ($110/MWh) than the cost for the NETL baseline SC coal firing case with 90% carbon capture ($114/MWh).

Magalhaes, Duarte

Global tuning of hadronic interaction models with accelerator-based and astroparticle data

In high-energy and astroparticle physics, event generators play an essential role, even in the simplest data analyses. As analysis techniques become more sophisticated, e.g. based on deep neural networks, their correct description of the observed event characteristics becomes even more important. Physical processes occurring in hadronic collisions are simulated within a Monte Carlo framework. A major challenge is the modeling of hadron dynamics at low momentum transfer, which includes the initial and final phases of every hadronic collision. QCD-inspired phenomenological models used for these phases cannot guarantee completeness or correctness over the full phase space. These models usually include parameters which must be tuned to suitable experimental data. Until now, event generators have been developed and tuned mainly on the basis of data from high-energy physics experiments at accelerators. The wealth of data available from the latest generation of astroparticle experiments has not yet been fully exploited, and in many cases is not satisfactorily described. Both kinds of data sets are complementary as astroparticle experiments provide sensitivity especially to hadrons produced nearly parallel to the collision axis and cover center-of-mass energies up to several hundred TeV, well beyond those reached at colliders so far. In this report, we provide an overview of state-of-the-art event generators and their tuning, including the most relevant inputs from high-energy accelerator and astroparticle experiments. We present a road map that shows, for the first time, how the unified tuning of event generators with accelerator-based and astroparticle data can be performed.

Albrecht, J. [Ruhr U., Bochum, RAPP Ctr.; Ruhr U.,

Enabling kilometer-scale E3SM land model simulation over North America: A new integrated framework solution

This study introduces a novel framework designed to enhance the performance, scalability, and portability of the kilometer-scale E3SM Land Model (km-ELM) within the E3SM modeling infrastructure. By seamlessly integrating cutting-edge data tools, we address existing challenges such as slow performance, limited scalability, and difficulties in software integration in current data-driven ELM simulation over large geographic areas. Our innovative approach leverages the KiloCraft data toolkit to generate unified inputs for simulations ranging from a single-cite case, to a 72,083-cell regional case to a continental configuration encompassing 21.6 million land grid cells at a 1 km × 1 km resolution. We conduct extensive strong- and weak-scaling experiments on three state-of-the-art supercomputers, utilizing up to 100,800 CPU cores across 2400 compute nodes to evaluate end-to-end metrics including wall-clock time, simulation-years-per-day (SYPD), initialization costs, and I/O throughput. Our results reveal the land (LND) component’s efficient scaling, demonstrating near-ideal weak scaling and strong-scaling parallel efficiencies reaching up to 87% at 50,400 cores. We confirm portability and reproducibility through bitwise-equivalent outputs across different machines using identical inputs over supported machines. Notably, at extreme scales, we identify I/O as a critical bottleneck and that leads to effective solution with the SCORPIO/ADIOS stack. Collectively, these findings validate the deployment of km-ELM at a continental scale with high parallel efficiency and provide essential guidance on configuration, decomposition, and I/O settings for optimized kilometer-scale land simulations in E3SM. This work emphasizes the innovative design and practical solutions that enhance the operational capabilities of km-ELM, focusing on software performance and scalability while leaving detailed scientific evaluations of simulated land processes for future investigations.

E3SM land model (ELM), km-ELM, scalability, perfor

Effect of Anoxic Iron Corrosion on WIPP Brine Geochemistry FY23 Final Report (U)

A 280-day study was completed to evaluate the effect of zero-valent iron (Fe 0 ) on the Waste Isolation Pilot Plant (WIPP) brine geochemistry under anticipated reducing conditions. Hydrogen (H 2 ) gas is expected to be present in the repository after closure due to the anoxic corrosion of a vast quantity of iron contained in the waste forms disposed at WIPP; therefore, a background argon atmosphere containing H 2 was chosen for this study. WIPP groundwater brine pH and E h will impact the mobility and fate of plutonium within the repository. Modeling and laboratory results for Castile WIPP brine indicate that equilibrium fa values relative to the standard hydrogen electrode (SHE) are 40 mV more reducing (i.e., more negative) than those for Salado WIPP brine (-480 mV vs. -440 mV, respectively) because of the higher pH of the Castile brine (pH 9 .3 for Castile vs. pH 8.8 for Salado). The E h and pH data were corrected for the effects of high ionic strength. The experimental results for both brines are consistent with thermodynamic predictions using OLI Systems' Mixed Solvent Electrolyte chemical equilibrium model. The measured and corrected pH and E h data from this study are provided in Table ES-I and Table ES-2, respectively. The experimental study, with four test conditions in triplicate, was performed in a dual glovebox with a nominally 3 vol.% H 2 in argon atmosphere (target H 2 range: 3 ± I vol.%). Simulants containing MgO only ( experimental control) and MgO+Fe 0 (WIPP base case) were prepared for both the Salado and Castile brines. MgO was included in all simulants to account for the use of bulk magnesium oxide in the WIPP repository. Fe 0 was included in some simulants to incorporate the effects of the anoxic corrosion of iron and in-situ hydrogen generation in the study. The brine compositions were developed by Sandia National Laboratory (SNL; Xiong, 2008) and have been used in previous WIPP evaluations. The test method (agitation, etc.) is partially based on ASTM D3987-12. Twelve rounds of periodic measurements of pH and E h were performed over the course of the study. Chemical analysis results for liquids and solids (ICP-MS, ICP-ES, IC Anion, TIC, SEM-EDX) are consistent with the pH, E h , and thermodynamic modeling results. This study included the following conditions that deviate from anticipated post-closure conditions following brine intrusion, but were selected to facilitate bench-scale testing to validate modeling of pH and E h for the post-closure WIP P repository: an anoxic glove box atmosphere containing ≤ 4 vol. % H 2 vs. substantially higher H 2 gas concentrations assumed in the WIPP Performance Assessment (PA); a significantly higher liquid-to-solid test ratio compared to the much lower phase ratio anticipated in the WIP P repository; agitation of the simulant bottles to maximize mass transfer; and finally the use of Fe 0 reagents having a much greater surface area than expected in the WIP P repository. Non-representative conditions were chosen for various reasons such as: to provide bounding conservative results, to provide a margin of safety for testing, or to facilitate simulant sub-sampling and analysis. In a parallel effort, aqueous electrolyte thermodynamic models were developed for the synthetic Salado and Castile brines to inform the experimental design, facilitate laboratory data interpretation, and allow extension of evaluations beyond the parameters tested. Thermodynamic modeling simulations including the MgO and Fe 0 additives that are directly relevant to the experimental measurements (e.g., pH calibration curve, ORP corrections) are included in this report. The measured fa of the simulants was close to the OLI model predictions for both brines and was largely controlled by the background H 2 partial pressure in the vapor phase as well as H 2 generated in situ in the aqueous phase by the Fe 0 corrosion. The H 2 gas-phase concentration tested and thermodynamically evaluated was much lower than is assumed in the WIPP PA; however, H 2 (g) concentrations significantly below this level are still predicted to result in very reducing conditions. In conclusion: • The experimental results are consistent with thermodynamic model predictions for fa, pH, and the effects of high ionic strength. • Evidence to date suggests that the H2 concentration in the glovebox atmosphere ultimately determined the final E h values of the simulants and resulted in highly reducing conditions. As a result, little difference was observed between the control simulants containing only MgO and the WIPP base-case simulants that contained MgO and Fe 0 . • This test methodology is recommended for future studies evaluating WIPP repository conditions. The methodology includes: (1) background H 2 in argon with agitation ( or could alternatively include in-situ-generated H 2 in sealed bottles); (2) carefully measured and corrected ORP data ( with much effort focused on allowing the probes to fully stabilize); and (3) ionic-strength-corrected pH data. Other best practices, such as simulant sparging/handling, ORP probe replacement, etc., should also be considered. • The coupling of experimental studies and thermodynamic modeling is also highly recommended because these methods inform and direct one another leading to greater confidence in and understanding of the results.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W

torch-einshard v1.0

torch-einshard is a Python library for describing local and distributed PyTorch tensor computations with compact, einsum-like notation. Its expressions name logical axes, specify how they are sharded across a PyTorch DeviceMesh, and represent partial reductions. The library automatically performs contractions, permutations, reshaping, splitting, gathering, reduction, reduce-scatter, and repartitioning while preserving autograd. Additional features include sharding-aware FFTs, tensor rolls, halo exchange, sliding windows, 1D–3D convolutions, uneven-shard handling, parameter initialization and gradient management, and cost-based execution planning. It is designed for scientific machine learning and large-model workloads, including tensor-, sequence-, and spatial-parallel MLPs, attention, convolutions, and spectral operations. Compared with manually combining torch.einsum and distributed collectives, torch-einshard expresses both the mathematical operation and data placement in one readable formula. This reduces boilerplate and synchronization errors, keeps forward and backward communication consistent, and allows the library to select optimized collective strategies without changing model code.

Morozov, Dmitriy [Lawrence Berkeley National Labor

Machine Learning-Driven Conservative-to-Primitive Conversion in Hybrid Piecewise Polytropic and Tabulated Equations of State

We present a novel machine learning (ML)-based method to accelerate conservative-to-primitive inversion, focusing on hybrid piecewise polytropic and tabulated equations of state. Traditional root-finding techniques are computationally expensive, particularly for large-scale relativistic hydrodynamics simulations. To address this, we employ feedforward neural networks (NNC2PS and NNC2PL), trained in PyTorch (2.0+) and optimized for GPU inference using NVIDIA TensorRT (8.4.1), achieving significant speedups with minimal accuracy loss. The NNC2PS model achieves 𝐿 1 and 𝐿 ∞ errors of 4.54 × 10 −7 and 3.44 × 10−6, respectively, while the NNC2PL model exhibits even lower error values. TensorRT optimization with mixed-precision deployment substantially accelerates performance compared to traditional root-finding methods. Specifically, the mixed-precision TensorRT engine for NNC2PS achieves inference speeds approximately 400 times faster than a traditional single-threaded CPU implementation for a dataset size of 1,000,000 points. Ideal parallelization across an entire compute node in the Delta supercomputer (dual AMD 64-core 2.45 GHz Milan processors and 8 NVIDIA A100 GPUs with 40 GB HBM2 RAM and NVLink) predicts a 25-fold speedup for TensorRT over an optimally parallelized numerical method when processing 8 million data points. Moreover, the ML method exhibits sub-linear scaling with increasing dataset sizes. We release the scientific software developed, enabling further validation and extension of our findings. By exploiting the underlying symmetries within the equation of state, these findings highlight the potential of ML, combined with GPU optimization and model quantization, to accelerate conservative-to-primitive inversion in relativistic hydrodynamics simulations.

conservative-to-primitive conversion

Numerical eigen-spectrum slicing, accurate orthogonal eigen-basis, and mixed-precision eigenvalue refinement using OpenMP data-dependent tasks and accelerator offload

Performing a variety of numerical computations efficiently and, at the same time, in a portable fashion requires both an overarching design followed by a number of implementation strategies. All of these are exemplified below as we present transitioning the PLASMA numerical library from relying on dependence-driven large tasks to achieving utilization of fine grain tasking and offload to hardware accelerators while keeping its core dependence sets: OpenMP source code pragmas and runtime for most system-level functionality and basic low-level numerical kernels provided directly by hardware vendors or open source projects with vendor contributions. We also present new algorithmic methods and their efficient parallel implementations including fine grained tasking for eigen-spectrum slicing and offload for mixed-precision eigenvalue refinement. We provide performance, scaling, and numerical results showing sizable gains over the available solutions from either the open source and vendor-provided packages.

Luszczek, Piotr

Flexible User-Defined Domain Decomposition in Kilometer-Scale E3SM Land Model Simulation

The Energy Exascale Earth System Model (E3SM) Land Model (ELM) has been extended to kilometer-scale (km-ELM) resolutions, enabling high-fidelity simulations of terrestrial processes at 1 km x 1 km grid spacing. In ELM, domain decomposition partitions the computational domain across processors, ensuring efficient parallel execution. Currently, round-robin decomposition is applied, providing a straightforward way to distribute computational workload. As ELM continues evolving at the kilometer-scale (km-scale), particularly with integrating lateral flow modeling, decomposition strategies must also account for the increased workload and data movement. This paper introduces a flexible user-defined domain decomposition framework, allowing users to customize domain partitioning based on application requirements. The impact of different decomposition strategies is evaluated across various applications concerning computation, communication, and I/O. Results demonstrate that while 1D partitioning yields superior I/O performance, k-nearest neighbors (KNN) clustering effectively reduces inter-process communication overhead. This study lays the groundwork for scalable partitioning in large-scale land surface simulations, enhancing next-generation Earth system modeling.

Wang, Dali [ORNL] (ORCID:0000000168065108)

DyG-DPCD: A Distributed Parallel Community Detection Algorithm for Large-Scale Dynamic Graphs

Dynamic (Temporal) graphs capture the valuable evolution of real-world systems, from the continuously evolving patterns of social interactions and genetic pathways to the dynamic fluctuations of economic forces. Detecting communities for such evolving networks poses unique challenges. Detecting and analyzing the evolution of communities within dynamic graphs unlocks valuable insights into the underlying structural and temporal patterns of real-world systems. However, the sheer volume of modern graph data and the inherent complexity of the temporal dimension pose significant challenges to scalable community detection algorithms. Addressing this gap, our work explores the limited landscape of scalable distributed-memory parallel methods specifically designed for dynamic network community detection. We propose a novel parallel algorithm, DyG-DPCD (Dynamic Graph Distributed Parallel Community Detection), to detect communities in dynamic networks using the Message Passing Interface (MPI) framework. We present a vertex-centric approach, allowing us to detect communities through local optimization. Furthermore, we enhance our baseline algorithm by incorporating three heuristics, which improve the algorithm’s performance significantly while maintaining the quality of the solutions. We demonstrate the efficiency of our algorithm by experimenting on several real-world large-scale networks with hundreds of millions of edges spanning diverse domains. Notably, DyG-DPCD achieves speedups between 25× and 30× for large networks that we experimented on using NERSC compute nodes. In conclusion, our algorithm outperforms the STINGER parallel re-agglomeration algorithm by 30×.

97 MATHEMATICS AND COMPUTING

Characterizing the effect of hypersonic boundary layer turbulence on antenna performance: A computational approach

The degradation of antenna performance during hypersonic re-entry is a well known phenomenon that can lead to complete radio blackout. Recent additions to the Empire code establish it as a tool for the study and analysis of the problem. Coupling to the Sandia Parallel Aerodynamics and Reentry Code (SPARC) enables the electromagnetic analysis of realistic re-entry plasma profiles. The geometric flexibility afforded by both Empire and SPARC allow the consideration of arbitrary vehicle and antenna configurations. We have used this tool to study antenna performance during re-entry when the boundary layer becomes turbulent. A concise description of line-of-sight transmissions, which employs advanced statistical methods, was developed. New insights into the low altitude reflectometer readings of RAM-C2 are offered. Techniques for the reconstruction of the re-entry plasma profile from reflectometer data were explored.

42 ENGINEERING

Spherical tokamak physics research in preparation for the operation of NSTX-U

The National Spherical Torus Experiment Upgrade (NSTX-U) is preparing to resume operation, representing a crucial step toward realizing compact, cost-effective fusion pilot plants. In advance of this, extensive modeling and data analysis have been conducted to advance the physics basis for low-aspect-ratio, high-performance plasma regimes, focusing on three core objectives: confinement and stability, power and particle handling, and steady-state operation. Significant progress has been made in understanding the electron temperature flattening in high-β plasmas, which is shown to be driven by a complex interplay of magnetohydrodynamic instabilities (e.g. non-resonant infernal modes), fast-ion-driven Alfvén eigenmodes, and electron and ion-scale micro-instabilities, particularly Kinetic Ballooning Modes (KBMs), whose destabilization is strongly dependent on parallel magnetic field fluctuations (δB ∥ ). Furthermore, a new gyrokinetic critical pedestal model was developed, accurately predicting pedestal structure by identifying KBMs as the primary stability limit, offering a critical constraint for future high-confinement scenarios. To address the challenge of high heat flux, novel liquid lithium plasma-facing components were modeled. The analysis confirmed that lithium vapor shielding is a self-regulating mechanism for heat mitigation, while also emphasizing that strong main ion parallel flow is essential to minimize core lithium contamination. Finally, progress toward steady-state operation was anchored by developing the required physics basis and control tools. This includes predictive modeling for reversed magnetic shear sustainment, demonstrating that magnetic island-induced bootstrap current reduction is negligible in STs, and advancing real-time control and disruption avoidance capabilities. The development of high-speed surrogate models (e.g. MMMNet) provides computationally efficient tools vital for non-inductive scenario optimization and integrated, low-disruptivity operations planned for NSTX-U.

NSTX-U