Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallel computation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29

Fast and Accurate Greenberger-Horne-Zeilinger Encoding Using All-to-All Interactions

The 𝑁-qubit Greenberger-Horne-Zeilinger (GHZ) state is an important resource for quantum technologies. Here, we consider the task of GHZ encoding using all-to-all interactions, which prepares the GHZ state in a special case, and is furthermore useful for quantum error correction, interaction-rate enhancement, and transmitting information using power-law interactions. The naive protocol based on parallelizing CNOT gates takes O(1)-time of Hamiltonian evolution. In this work, we propose a fast protocol that achieves GHZ encoding with high accuracy. The evolution time O⁡(log 2 ⁡𝑁/𝑁) almost saturates the theoretical limit Ω⁡(log⁡𝑁/𝑁). Moreover, the final state is close to the ideal encoded one with high fidelity >1–10 −3 , up to large system sizes 𝑁 ≲ 2000. The protocol only requires a few stages of time-independent Hamiltonian evolution; the key idea is to use the data qubit as control, and to use fast spin-squeezing dynamics generated by e.g., two-axis twisting.

quantum computation↗

3D modeling of deep borehole electromagnetic measurements with energized casing source for fracture mapping at the Utah Frontier Observatory for Research in Geothermal Energy

Here, we present a 3D numerical modelling analysis evaluating the deployment of a borehole electromagnetic measurement tool to detect and image a stimulated zone at the Utah Frontier Observatory for Research in Geothermal Energy geothermal site. As the depth to the geothermal reservoir is several kilometres and the size of the stimulated zone is limited to several 100 m, surface-based controlled-source electromagnetic measurements lack the sensitivity for detecting changes in electrical resistivity caused by the stimulation. To overcome the limitation, the study evaluates the feasibility of using a three-component borehole magnetic receiver system at the Frontier Observatory for Research in Geothermal Energy site. To provide sufficient currents inside and around the enhanced geothermal reservoir, we use an injection well as an energized casing source. To efficiently simulate energizing the injection well in a realistic 3D resistivity model, we introduce a novel modelling workflow that leverages the strengths of both 3D cylindrical-mesh-based electromagnetic modelling code and 3D tetrahedral-mesh-based electromagnetic modelling code. The former is particularly well-suited for modelling hollow cylindrical objects like casings, whereas the latter excels at representing more complex 3D geological structures. In this workflow, our initial step involves computing current densities along a vertical steel-cased well using a 3D cylindrical electromagnetic modelling code. Subsequently, we distribute a series of equivalent current sources along the well's trajectory within a complex 3D resistivity model. We then discretize this model using a tetrahedral mesh and simulate the borehole electromagnetic responses excited by the casing source using a 3D finite-element electromagnetic code. This multi-step approach enables us to simulate 3D casing source electromagnetic responses within a complex 3D resistivity model, without the need for explicit discretization of the well using an excessive number of fine cells. We discuss the applicability and limitations of this proposed workflow within an electromagnetic modelling scenario where an energized well is deviated, such as at the Frontier Observatory for Research in Geothermal Energy site. Using the workflow, we demonstrate that the combined use of the energized casing source and the borehole electromagnetic receiver system offer measurable magnetic field amplitudes and sensitivity to the deep localized stimulated zone. The measurements can also distinguish between parallel-fracture anisotropic reservoirs and isotropic cases, providing valuable insights into the fracture system of the stimulated zone. Besides the magnetic field measurements, vertical electric field measurements in the open well sections are also highly sensitive to the stimulated zone and can be used as additional data for detecting and imaging the target. We can also acquire additional multiple-source data by grounding the surface electrode at various locations and repeating borehole electromagnetic measurements. This approach can increase the number of monitoring data by several factors, providing a more comprehensive dataset for analysing the deep-localized stimulated zone. The numerical analysis indicates that it is feasible to use the combination of the energized casing and downhole electromagnetic measurements in monitoring localized stimulated zone at large depths.

58 GEOSCIENCES↗

Optimization and Experimental Validation of Annular Finned PCM-HX for a Domestic Hot Water Heater Application

The load profile for domestic water heating is time-dependent and can result in high energy demand during peak operating times. Shifting this peak load can have significant environmental and economic impacts. Phase change material (PCM)-based thermal energy storage (TES) is a potentially useful technology for peak load shifting in domestic hot water (DHW) applications thanks to its high latent heat and energy density. In this study, an annular finned-tube PCM-HX design concept was optimized for a load-shifting TES unit to meet the Department of Energy standard for a medium-usage DHW heater using a resistance-capacitance model (RCM) integrated with a Multi-Objective Genetic Algorithm. The optimized design comprised 70 identical annular finned-tube PCM-HX units connected in parallel and utilizing RT62HC as the PCM. A single PCM-HX unit was prototyped and tested in a vertically oriented setup with upward heat transfer fluid (HTF) flow. The hot water supply time was defined based on a cutoff temperature of 51.7°C. The as-designed mass flow rate (1.5 g/s) was tested to assess the performance of the prototyped PCM-HX unit for RCM validation. For the experimental investigation, RTD sensor bundles measured HTF temperature at the PCM-HX inlet and outlet, and a Coriolis flow meter accurately measured the HTF mass flow rate. The simulated discharging power underpredicted the experimental result by about 12%, and the simulated hot water supply time underpredicted the experimental result by approximately 13% for the as-designed mass flow rate (1.5 g/s). The average deviation of the hot water supply temperature between the experimental and RCM results during the complete PCM solidification process was 1.3 K for the as-designed mass flow rate. The overall good agreement between the experimental and RCM results provides confidence that computationally efficient models such as RCM can be utilized for design optimization of PCM-HXs.

42 ENGINEERING↗

Using Apptainer in a Pilot-based Distributed Workload

GlideinWMS is a pilot and pressure-based workload manager for distributed scientific computing. Many experiments like CMS and Fermilab’s Neutrino experiments use it to provision elastic clusters for their analysis and simulations, split into close to a million concurrent jobs. Most user jobs require containers, and the pilots use Apptainer to set up the desired platform. For the pilots that run as regular batch jobs, Apptainer is safer, lighter, and easier to use than other containerization solutions. Many images used by the pilots are expanded SIF images distributed via the CernVM-FS: this combination is very efficient. At Fermilab, for example, we store on GitHub Dockerfiles that mimic the platform in the worker nodes of local clusters. GitHub workflows build and push the images to Docker Hub, and a service periodically pulls and converts them to the expanded SIF images in the CernVM-FS, so the scientists can find a familiar environment everywhere. Apptainer has also been used to run services inside the pilot jobs, like benchmarks that characterize the worker node being used, or a Triton Inference Server that allows sharing a GPU with all the jobs that run in parallel on a node.

Mambelli, Marco [Fermilab] (ORCID:0000000294892681↗

Expert and operator perspectives on barriers to energy efficiency in data centers

Abstract It was last estimated in 2016 that data centers (DCs) comprise approximately 2% of total US electricity consumption. However, this estimate is currently being updated to account for the massive increase in computing needs due to streaming, cryptocurrency, and artificial intelligence (AI). To prevent energy consumption that tracks with increasing computing needs, it is imperative we identify energy efficiency strategies and investments beyond the low-hanging fruit solutions. In a two-phased research approach, we ask: What non-technical barriers still impede energy efficiency (EE) practices and investments in the data center sector, and what can be done to overcome these barriers? In particular, we are focused on social and organizational barriers to EE. In Phase I, we performed a literature review and found that technical solutions are abundant in the literature, but fail to address the top-down cultural shifts that need to take place in order to adapt new energy efficiency strategies. In Phase II, reported here, we interviewed 16 data center operators/experts to ground-truth our literature findings. Our interview protocols focus on three aspects of DC decision-making: procurement practices, metrics and monitoring, and perceived barriers to energy efficiency. We find that vendors are the key drivers of procurement decisions, advanced efficiency metrics are facility-specific, and there is convergence in the design of advanced facilities due to the heat density of parallelized infrastructure. Our ultimate goals for our research are to design DC decarbonization policies that target organizational structure, empower individual staff, and foster a supportive external market.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Evaluation of DED and LPBF Fe-based Alloys Process Application Envelopes based on Performance, Process Economics, Supply Chain Risks, and Reactor-specific Targeted Components

The U.S. Department of Energy (DOE), Office of Nuclear Energy (NE), Advanced Materials and Manufacturing Technologies (AMMT) program aims to develop extreme-environment materials solutions for use in the deployment of advanced nuclear reactors and the sustainment of the current fleet. To achieve this objective, a combination of experiment, a computational tool, and machine learning (ML) for the design of materials is adopted for the maturation of materials for nuclear technology. Through advanced manufacturing techniques such as laser powder bed fusion (LPBF) and laser powder direct energy deposition (LP-DED), components with complex geometries can be fabricated with reduced time and effort. Such advanced manufacturing methods can also provide the opportunity to improve materials performance through optimized microstructures and mechanical properties. However, existing engineering alloys are not always well suited for fabrication with additive manufacturing (AM), as their compositions have been tuned to optimize fabrication via conventional methods. Thus, similar alloys with modified compositions that are better suited for AM can be studied for improved performance. Over the past three years, the AMMT teams from Argonne National Laboratory (ANL) and Pacific Northwest National Laboratory (PNNL) studied various known Fe-based alloys by evaluating their initial printability using LPBF, and an AMMT-developed down-selection and decision matrix reduced the number of alloys to be studied from six to three in fiscal year (FY) 2024. Additionally, in FY 2024, for parallel evaluation, these three alloys were studied using LPDED. While LPBF is better for small- to medium-sized components with high detail and internal features, LP-DED combines a material feed system to place the powder onto the exact spot where the laser will melt the material. This AM method can be easily scaled to extremely large components and provides high build rate speeds compared to those of conventional LPBF systems. Additionally, DED is a better choice for complex geometries and compositional gradients.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Enhancing high-fidelity neural network potentials through low-fidelity sampling

The efficacy of neural network potentials (NNPs) critically depends on the quality of the configurational datasets used for training. Prior research using empirical potentials has shown that well-selected liquid–solid transitional configurations of a metallic system can be translated to other metallic systems. This study demonstrates that such validated configurations can be relabeled using density functional theory (DFT) calculations, thereby enhancing the development of high-fidelity NNPs. Training strategies and sampling approaches are efficiently assessed using empirical potentials and subsequently relabeled via DFT in a highly parallelized fashion for high-fidelity NNP training. Our results reveal that relying solely on energy and force for NNP training is inadequate to prevent overfitting, highlighting the necessity of incorporating stress terms into the loss functions. To optimize training involving force and stress terms, we propose employing transfer learning to fine-tune the weights, ensuring that the potential surface is smooth for these quantities composed of energy derivatives. This approach markedly improves the accuracy of elastic constants derived from simulations in both empirical potential-based NNPs and relabeled DFT-based NNPs. Overall, this study offers significant insights into leveraging empirical potentials to expedite the development of reliable and robust NNPs at the DFT level.

97 MATHEMATICS AND COMPUTING↗

A survey on checkpointing strategies: Should we always checkpoint à la Young/Daly?

The Young/Daly formula provides an approximation of the optimal checkpointing period for a parallel application executing on a supercomputing platform. It was originally designed to handle fail-stop errors for preemptible tightly-coupled applications, but has been extended to other application and resilience frameworks. Here, we provide some background and survey various scenarios to assess the usefulness and limitations of the formula, both for preemptible applications and workflow applications represented as a graph of tasks. We also discuss scenarios with uncertainties, and extend the study to silent errors. We exhibit cases where the optimal period is of a different order than that dictated by the Young/Daly formula, and finally we explain how checkpointing can be further combined with replication.

97 MATHEMATICS AND COMPUTING↗

Enabling high-throughput enzyme discovery and engineering with a low-cost, robot-assisted pipeline

Abstract As genomic databases expand and artificial intelligence tools advance, there is a growing demand for efficient characterization of large numbers of proteins. To this end, here we describe a generalizable pipeline for high-throughput protein purification using small-scale expression in E. coli and an affordable liquid-handling robot. This low-cost platform enables the purification of 96 proteins in parallel with minimal waste and is scalable for processing hundreds of proteins weekly per user. We demonstrate the performance of this method with the expression and purification of the leading poly(ethylene terephthalate) hydrolases reported in the literature. Replicate experiments demonstrated reproducibility and enzyme purity and yields (up to 400 µg) sufficient for comprehensive analyses of both thermostability and activity, generating a standardized benchmark dataset for comparing these plastic-degrading enzymes. The cost-effectiveness and ease of implementation of this platform render it broadly applicable to diverse protein characterization challenges in the biological sciences.

36 MATERIALS SCIENCE↗

Symbol alphabets in QCD and flag cluster algebras

The full 245-letter symbol alphabet for all planar massless two-loop six-point Feynman integrals was recently determined in arXiv:2412.19884 and arXiv:2501.01847. In a parallel mathematical development, it was shown in arXiv:2408.14956 that there is an embedding of the cluster algebra associated to the partial flag variety $\mathcal{Fl}$ $2,n-2;n$ , which describes the kinematics of n massless particles, into that of the Grassmannian Gr(n–2, 2n–4). In this paper we connect these developments by showing that most of the rational symbol letters can be expressed in terms of flag cluster variables, and that all of the algebraic symbol letters arise from infinite mutation sequences.

97 MATHEMATICS AND COMPUTING↗

Porting Classical Approaches for Quantum Simulations to Quantum Computers

Simulating quantum many-body systems is one of the most promising problems in which we might anticipate that quantum computers should show quantum advantage. Unfortunately, there is still a gap between this promise and actual practice. New quantum algorithms need to be developed and the current quantum algorithms have various difficulties - e.g efficient state preparation - which must be overcome and improved upon. In many cases, classical approaches need to be ported over to quantum devices. In this project we have developed a suite of new quantum algorithms which makes progress in this regard. We developed a new optimization scheme for variational quantum eigensolvers, UBOS, which mitigates problems with local minimas and barren plateaus while improving convergence to the ground state by an order of magnitude. We developed a new way to utilize qubitization to find ground states of nearly frustration-free Hamiltonians faster than all previous methods. We developed a series of state preparation techniques which helps initialize parameterized quantum circuits into reasonable starting points on which quantum algorithms are then applied. In addition to the development of novel algorithms, it is critical to have classical simulation techniques for approximately simulating quantum circuits which can be used to benchmark and understand quantum algorithms. Toward that end, we developed a novel POVM formalism to simulate quantum circuits as well as exemplify the massive parallelization of tensor network methodologies. Finally, we developed physical understanding of entanglement phase transitions such as many-body localization and random tensor networks.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Method to simultaneously facilitate all jet physics tasks

Machine learning has become an essential tool in jet physics. Due to their complex, high-dimensional nature, jets can be explored holistically by neural networks in ways that are not possible manually. However, innovations in all areas of jet physics are proceeding in parallel. We show that specially constructed machine learning models trained for a specific jet classification task can improve the accuracy, precision, or speed of all other jet physics tasks. This is demonstrated by training on a particular multiclass generation and classification task and then using the learned representation for different generation and classification tasks, for datasets with a different (full) detector simulation, for jets from a different collision system ($pp$ versus $ep$), for generative models, for likelihood ratio estimation, and for anomaly detection. We consider our omnilearn approach thus as a jet-physics foundation model. It is made publicly available for use in any area where state-of-the-art precision is required for analyses involving jets and their substructure.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Nonlocal, diamagnetic electromagnetic effects in magnetically insulated transmission lines

We identify the time-dependent physics responsible for the critical reduction of current losses in magnetically insulated transmission lines (MITLs) due to uninsulated space charge-limited currents of electrons emitted by field stress. A drive current of sufficiently short pulse length introduces a strong enough time dependence that steady-state results alone become inadequate for the complete understanding of current losses. The time-dependent physics can be described as a nonlocal, diamagnetic electromagnetic response of space charge limited currents. As the pulse length is increased or equivalently, the MITL length reduced, these time-dependent effects diminish and current losses converge to those predicted by the well-known Child–Langmuir law in the external (vacuum) fields. We present a simple one-dimensional (1D) model that encapsulates the essence of this physics. We find excellent agreement with 2D particle-in-cell simulations for two MITL geometries, Cartesian parallel plate and azimuthally symmetric straight coaxial. Based on the 1D model, we explore various scaling dependencies of MITL losses with relevant parameters, e.g., peak current, pulse length, geometrical dimensions, etc. We propose an improved physics model of magnetic insulation in the form of a Hull curve, which could also help improve predictions of current losses by common circuit element codes, such as BERTHA. Finally, we describe how to calculate the temperature rise due to electron impact within the 1D model.

Computer simulation↗

Low two-level-system noise in hydrogenated amorphous silicon

At sub-Kelvin temperatures, two-level systems (TLSs) present in amorphous dielectrics source a permittivity noise, degrading the performance of a wide range of devices using superconductive resonators such as qubits or kinetic inductance detectors. We report here on measurements of TLS noise in hydrogenated amorphous silicon (a-Si:H) films deposited by plasma-enhanced chemical vapor deposition in superconductive lumped element resonators using parallel-plate capacitors. In conclusion, the TLS noise results presented in this article for two recipes of a-Si:H improve on the best results achieved in the literature by a factor >5 for a-Si:H and other amorphous dielectrics and are comparable to those observed for resonators deposited on crystalline dielectrics.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Asynchronous-many-task systems: Challenges and opportunities - Scaling an AMR astrophysics code on exascale machines using Kokkos and HPX

Dynamic and adaptive mesh refinement is pivotal in high-resolution, multi-physics, multi-model simulations, necessitating precise physics resolution in localized areas across expansive domains. Today’s supercomputers’ extreme heterogeneity presents a significant challenge for dynamically adaptive codes, highlighting the importance of achieving performance portability at scale. Our research focuses on astrophysical simulations, particularly stellar mergers, to elucidate early universe dynamics. Here, we present Octo-Tiger, leveraging Kokkos, HPX, and SIMD for portable performance at scale in complex, massively parallel adaptive multi-physics simulations. Octo-Tiger supports diverse processors, accelerators, and network backends. Experiments demonstrate exceptional scalability across several heterogeneous supercomputers including Perlmutter, Frontier, and Fugaku, encompassing major GPU architectures and x86, ARM, and RISC-V CPUs. Parallel efficiency of 47.59% (110,080 cores and 6880 hybrid A100 GPUs) on a full-system run on Perlmutter (26% HPCG peak performance) and 51.37% (using 32,768 cores and 2048 MI250X) on Frontier are achieved.

97 MATHEMATICS AND COMPUTING↗

Sierra/SD – Theory Manual – 5.22

Sierra/SD provides a massively parallel implementation of structural dynamics finite element analysis, required for high fidelity, validated models used in modal, vibration, static and shock analysis of structural systems. This manual describes the theory behind many of the constructs in Sierra/SD. For a more detailed description of how to use Sierra/SD, we refer the reader to User’s Manual. Many of the constructs in Sierra/SD are pulled directly from published material. Where possible, these materials are referenced herein. However, certain functions in Sierra/SD are specific to our implementation. We try to be far more complete in those areas. The theory manual was developed from several sources including general notes, a programmer_notes manual, the user’s notes and of course the material in the open literature.

97 MATHEMATICS AND COMPUTING↗

First-Principles-Based Study of the Decomposition of Phenol and Hydroquinone on Pt(111) Combined with Quantitative Information from XPS Spectra to Address the Impact of Coverage and Number of Hydroxyl Functional Groups

A combined first-principles-based and experimental X-ray photoelectron spectroscopy approach was used to investigate the thermal decomposition of two model biofuel compounds, phenol and hydroquinone, on Pt(111) at both low and high coverages. The DFT-based approach yields adsorption geometries and energies, activation barriers and core-level binding energy shifts for C 1s and O 1s. Increasing the coverage in the theoretical model leads to slight shifts in core-level binding energies─toward higher values for C 1s and lower values for O 1s. It also alters the energy profiles of the decomposition reaction pathway, resulting in weaker adsorption energies and changes in both reaction and activation barriers. At low temperatures, we observe a multilayer for phenol and hydroquinone upon adsorption, with desorption occurring at 200 and 270 K, respectively. Following desorption of the multilayer, decomposition proceeds via initial O–H bond scission, followed by two parallel pathways involving either C–H or C–C bond scission, whereby in the case of phenol C–H bond scission occurs first. Here, we further provide characteristic core level binding energies by theoretical calculations that are subsequently used in experimental analyses, establishing a reference database for key spectra of phenolic functionalities applicable to a range of catalytic reactions.

09 BIOMASS FUELS↗

A two-level GPU-accelerated incomplete LU preconditioner for general sparse linear systems

This paper presents a parallel preconditioning approach based on incomplete LU (ILU) factorizations in the framework of Domain Decomposition (DD) for general sparse linear systems. We focus on distributed memory parallel architectures, specifically, those that are equipped with graphic processing units (GPUs). In addition to block-Jacobi, we present general purpose two-level ILU Schur complement-based approaches, where different strategies are presented to solve the coarse-level reduced system. These strategies are combined with modified ILU methods in the construction of the coarse-level operator, in order to effectively remove smooth errors by targeting an algebraically smooth vector. We leverage available GPU-based sparse matrix kernels to accelerate the setup and the solve phases of the proposed ILU preconditioner. We evaluate the efficiency of the proposed methods as a smoother for algebraic multigrid (AMG) and as a preconditioner for Krylov subspace methods on challenging anisotropic diffusion problems and a collection of general sparse matrices.

97 MATHEMATICS AND COMPUTING↗