Search NASA⌕ Search

SEARCH · Search NASA

Results for “cluster computer”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25

Visualization of Noisy and Less Noisy Computational Basis States in Quantum Computing

Quantum computing technology holds substantial promise as a reliable computational paradigm. However, current noisy intermediate scale quantum (NISQ) systems, are significantly impacted by noise originating from hardware inconsistencies. This noise causes errors and lowers output fidelity. So we must find which basis states cause errors. However, there are two main challenges in analyzing noise corresponding to basis states. First, the noise distribution data is high dimensional in nature, thereby making its analysis challenging. Second, although functional box plots have been used in the state of the art research to understand such a high dimensional data, they suffer from clutter and occlusion issues because of overplotting. In this study, we introduce an innovative visualization pipeline to address the aforementioned challenges to provide a clear depiction of noisy and less-noisy basis states. Specifically, our proposed visualization pipeline comprises three stages namely, low dimensional embedding, clustering, and violin plot visualization, to reduce visual clutter and effectively analyze high-dimensional noise distribution data. Our analysis uses quantum machine learning (QML) circuits as case study for drawing a distinction between noisy and less noisy basis states.

Senapati, Priyabrata [Kent State University]↗

Net Present Value Optimization of a Natural Gas Combined Cycle Plant with CO 2 Capture using a Water-Lean Solvent Considering Transient Electricity Price for Multiple Regions

Global CO 2 emissions are increasing at about a 1.5% rate per year. Fossil fuel-based plants are one of the main contributors to this rise. In the power generation industry, fossil fuel plants are dominant, and many plants are under development. In this study, a natural gas combined cycle (NGCC) power plant with postcombustion capture using a leading water-lean solvent is considered. For optimal design and operating schedule, large-scale dynamic optimization is undertaken for net present value (NPV) optimization. The first principle dynamic model of NGCC is developed, including a model of the highly efficient H-class gas turbines. For computational tractability of the dynamic optimization problem, a reduced-order model is developed by using the Hankel singular value decomposition. A waterlean solvent, N-(2-ethoxyethyl)-3-morpholinopropan-1-amine, is used for carbon capture. A model of the capture system is developed in Aspen Plus, which is used to develop a reduced-order model by using ALAMO, a machine learning software. In addition, a reduced model of the CO 2 compression system with a dehydration unit is also considered. The integrated system is used for NPV optimization by using the Python-based PYOMO platform. The PCC process is analyzed for three configurations-conventional packed bed, rotating packed bed (RPB), and a combination of RPB and direct contact cooler. The NPV optimization is performed for 14 regional markets by considering year-long clustered and continuous locational marginal price data with a 1 h interval. Optimization results show that the PCC can achieve 90% CO 2 capture with a positive NPV for six regions. Sensitivity studies conducted by using the PCC configurations indicate that the process is economically feasible for 9 regions out of 14 regional electricity markets with NPV values in the range of 33−540 $MM.

cabon capture↗

UnigeneFinder: An Automated Pipeline for Gene Calling From Transcriptome Assemblies Without a Reference Genome

ABSTRACT For most species, transcriptome data are much more readily available than genome data. Without a reference genome, gene calling is cumbersome and inaccurate because of the high degree of redundancy in de novo transcriptome assemblies. To simplify and increase the accuracy of de novo transcriptome assembly in the absence of a reference genome, we developed UnigeneFinder. Combining several clustering methods, UnigeneFinder substantially reduces the redundancy typical of raw transcriptome assemblies. This pipeline offers an effective solution to the problem of inflated transcript numbers, achieving a closer representation of the actual underlying genome. UnigeneFinder performs comparably or better, compared with existing tools, on plant species with varying genome complexities. UnigeneFinder is the only available transcriptome redundancy solution that fully automates the generation of primary transcript, coding region, and protein sequences, analogous to those available for high‐quality reference genomes. These features, coupled with the pipeline’s cross‐platform implementation, focus on automation, and an accessible, user‐friendly interface, make UnigeneFinder a useful tool for many downstream sequence‐based analyses in nonmodel organisms lacking a reference genome, including differential gene expression analysis, accurate ortholog identification, functional enrichments, and evolutionary analyses. UnigeneFinder also runs efficiently both on high‐performance computing (HPC) systems and personal computers, further reducing barriers to use.

Xue, Bo [Plant Resilience Institute Michigan State↗

The Three Hundred Project: Modeling baryon and hot-gas fraction evolution in simulated clusters

The baryon fraction of galaxy clusters, expressed as the ratio between the mass in baryons (including both stars and cold or hot gas) and the total mass, is a powerful tool to provide information on the cosmological parameters, while the hot-gas fraction provides indications on the physics of the intracluster plasma and its interplay with the processes that drive galaxy formation. Using cosmological hydrodynamical simulations of about 300 simulated massive galaxy clusters with a median mass M 500 ≈ 7 × 10 14 M ⊙ at z = 0, we model the relations between total mass and either baryon fraction or the hot gas fractions at overdensities Δ = 2500, 500, and 200 with respect to the cosmic critical density, and their evolution from z ∼ 0 to z ∼ 1.3. We utilized the simulated galaxy clusters from the Three Hundred project, which include star formation and feedback from both supernovae and active galactic nuclei. We fit the simulation results for such scaling relations against three analytic forms (linear, quadratic, and logarithmic in a logarithmic plane) and three forms for the redshift dependence, and we considered as a variable both the inverse of the cosmic scale factor, (1 + z), and the Hubble expansion rate, E(z). We show that power-law dependencies on cluster mass poorly describe the investigated relations. A power law fails to simultaneously capture the flattening of the total baryon and gas fractions at high masses, their drop at low masses, and the transition between these two regimes. The other two functional forms provide a more accurate description of the curvature in mass scaling. The fractions measured within smaller radii exhibit a stronger evolution than those measured within larger radii. From the analysis of these simulations, we evince that as long as we include systems in the mass range herein investigated, the baryon or gas fraction can be accurately related to the total mass through either a parabola or a logarithm in the logarithmic plane. The trends are common to all modern hydro simulations, although the amplitude of the drop at low masses might differ. Being able to observationally determine the gas fraction in groups will thus provide constraints on the baryonic physics.

galaxy clusters↗

Are Marine Low Cloud Droplet Concentrations Buffered by Entrained Aitken‐Mode Aerosol (Final Technical Report)

During the summertime, the high-latitude oceans come to life with green phytoplankton, which gain their energy from sunlight and are food for sea creatures small and large. When the phytoplankton are eaten or die, sulfur-rich gases are released under the ocean surface and mix into the air. Observations suggest that, over the Southern Ocean, frequent storms lift this air high into the atmosphere while raining out particulates like salt. As a result, the sulfur-rich air then spawns high concentrations of small ‘Aitken-mode’ aerosol particles. We hypothesize that these particles work their way down into the marine boundary layer, where they can replenish the supply of cloud-condensation nuclei scavenged by frequent precipitation. Further, this process maintains high concentrations of liquid cloud droplets in austral summer, promoting more sunlight to be reflected to space. We call this ‘Aitken buffering’. The primary objective of this project has been to document and test what role Aitken-mode aerosols play in clouds over the Southern Ocean and elsewhere. This effort has included three main components: 1) developing a computer model that realistically simulates the aerosol processes and the small-scale turbulent air motions that move aerosols around and create the clouds, 2) using that model to interpret and extend these observations for process understanding, by allowing different factors that contribute to the aerosol budget, such as surface wind speed, precipitation, surface gas exchange, etc. to be separated, and 3) studying Aitken-mode aerosol and its variability with a focus over the Southern Ocean and Antarctica, using data from Atmospheric Radiation Measurement (ARM) sites and other available observations. Initial computer studies in more idealized conditions found that elevated concentrations of Aitken-mode aerosols above the clouds could help prevent the breakup of those clouds by acting as cloud-condensation nuclei after they were entrained into the cloudy boundary layer. Simulations of a day during the ACE-ENA field campaign showed that Aitken-mode aerosols could also prevent cloud breakup under more realistic conditions. A new method that extracts information about Aitken-mode aerosols from measurements of aerosols onto which cloud droplets can form finds that Aitken-mode aerosols do vary seasonally over the Southern Ocean, with a peak in summertime, as described above. Other work during this project has focused on understanding how patterns of water vapor, clouds and precipitation are coupled within low-lying clouds over the oceans, and also on how cloud droplets cluster within clouds and how the distribution of cloud droplet sizes change as dry air is mixed into clouds, with the latter studies also using observations from ACE-ENA.

54 ENVIRONMENTAL SCIENCES↗

ROOT RNTuple and EOS: The Next Generation of Event Data I/O

For several years, the ROOT team is developing the new RNTuple I/O subsystem in preparation of the next generation of collider experiments. Both HL-LHC and DUNE are expected to start data taking by the end of this decade. They pose unprecedented challenges to event data I/O in terms of data rates, event sizes, and event complexity. At the same time, the I/O landscape is becoming more diverse. HPC cluster file systems and object stores, NVMe disk cache layers in analysis facilities, and S3 storage on cloud resources are mixing with traditional XRootD-managed spinning disk pools.The ROOT team will finalize a first production version of the RNTuple binary format by the end of 2024. After this point, ROOT will provide backward compatibility for RNTuple data. This contribution provides an overview of the RNTuple feature set, the related R&D activities and the long-term vision for RNTuple. We report on performance, interface design, tooling, robustness, integration with experiment frameworks, and validation results, as well as recent R&D on parallel reading and writing and exploitation of modern hardware and storage systems. We will give an outlook on possible future features after a first production release.Collaboratively, the IT and EP departments at CERN have launched a formal project within the Research and Computing sector to evaluate the novel data format for physics analysis data utilized in LHC experiments and other fields. This part of the project focuses on validating the scalability of the EOS storage backend during the transition from the over 25 years old TTree production format to the newly developed RNTuple format, using both replicated and erasure-coded storage profiles.

Blomer, Jakob [CERN]↗

Improving Bond Dissociations of Reactive Machine Learning Potentials through Physics-Constrained Data Augmentation

In the field of computational chemistry, predicting bond dissociation energies (BDEs) presents well-known challenges, particularly due to the multireference character of reactive systems. Many chemical reactions involve configurations where single-reference methods fall short, as the electronic structure can significantly change during bond breaking. As generating training data for partially broken bonds is a challenging task, even state-of-the-art reactive machine learning interatomic potentials (MLIPs) often fail to predict reliable BDEs and smooth dissociation curves. By contrast, simple and inexpensive physics-based models, such as the well-established Morse potential, do not suffer from any such limitations. This work leverages the Morse potential to improve reactive MLIPs by augmenting the training data set with inexpensive Morse data along the dissociation pathways. Further, this physics-constrained data augmentation (PCDA) approach results in MLIPs with smooth bond dissociation curves as well as near coupled-cluster level BDEs, all without requiring any expensive multireference quantum mechanical calculations. A case study for methane combustion demonstrates how the PCDA approach can improve an existing reactive MLIP, namely, ANI-1xnr. In conclusion, not only are the BDEs and bond dissociation curves for all radicals and molecules significantly improved compared to ANI-1xnr but the PCDA-trained MLIP retains the reliability of ANI-1xnr when performing reactive molecular dynamics simulations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Fast calculation of diffraction patterns from an ensemble of aligned molecules

We report an algorithm to calculate electron diffraction patterns for molecules with anisotropic angular distribution, which is significantly faster than existing methods. The algorithm uses a transform to convert the molecular orientation distribution, which is a function of three Euler angles, to the atom-pair distribution functions which depend on the polar and azimuthal angles. The diffraction signal can then be calculated from the atom-pair distributions. We demonstrate the computation method numerically by calculating electron diffraction patterns for a symmetric top molecule (trifluoroiodomethane) and an asymmetric top molecule (formaldehyde) and show that it reduces the calculation time by approximately two orders of magnitude compared to the standard brute-force method. Here, the method can also be applied to the calculation of x-ray diffraction patterns.

74 ATOMIC AND MOLECULAR PHYSICS↗

Electric dipole polarizability of 58 Ni

The electric dipole strength distribution in 58 Ni between 6 and 20 MeV has been determined from proton inelastic scattering experiments at very forward angles at RCNP, Osaka. The experimental data are rather well reproduced by quasiparticle random-phase approximation calculations including vibration coupling, despite a mild dependence on the adopted Skyrme interaction. They allow an estimate of the experimentally inaccessible high-energy contribution above 20 MeV, leading to an electric dipole polarizability α D ⁡ ( 58 Ni) = 3.48 (31)⁢ fm 3 . This serves as a test case for recent extensions of coupled-cluster calculations with chiral effective field theory interactions to nuclei with two nucleons on top of a closed-shell system.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

FitCache: A Transparent Drop-In Framework for Multi-Tier Caching to Accelerate Distributed Deep Learning Workloads

Training in Deep learning (DL) remains highly compute- and data-intensive, with I/O becoming a critical bottleneck as models and datasets scale. Recent studies report that data loading can dominate training time, especially on large-scale HPC systems with shared parallel file systems (PFS). Existing caching approaches either rely on single-tier designs or require intrusive modifications to training pipelines, limiting their portability and effectiveness. In this work, we present FitCache, a transparent drop-in framework for multi-tier caching to accelerate distributed DL training by coordinating fast local memory (e.g., DRAM, Persistent Memory (PMem)) and NVMe as hierarchical caches atop PFS. Our design adapts to hardware diversity, i.e., if NVMe is missing, memory transparently acts as a caching tier, ensuring stable performance. FitCache transparently intercepts I/O requests and issues concurrent fetches across all tiers, returning data from the fastest responder without centralized metadata or static redirection paths. FitCache adapts to dynamic workloads and heterogeneous clusters while maintaining POSIX compatibility. Experiments on Frontier (2048 GPUs) and smaller research clusters show that FitCache reduces training time by up to 40% and per-batch I/O latency by up to 71.6% compared to Lustre Orion PFS, offering a drop-in solution for scalable DL training.

Hu, Guangxing [ORNL] (ORCID:0009000283203614)↗

Unveiling the microbial realm with VEBA 2.0: a modular bioinformatics suite for end-to-end genome-resolved prokaryotic, (micro)eukaryotic and viral multi-omics from either short- or long-read sequencing

Abstract The microbiome is a complex community of microorganisms, encompassing prokaryotic (bacterial and archaeal), eukaryotic, and viral entities. This microbial ensemble plays a pivotal role in influencing the health and productivity of diverse ecosystems while shaping the web of life. However, many software suites developed to study microbiomes analyze only the prokaryotic community and provide limited to no support for viruses and microeukaryotes. Previously, we introduced the Viral Eukaryotic Bacterial Archaeal (VEBA) open-source software suite to address this critical gap in microbiome research by extending genome-resolved analysis beyond prokaryotes to encompass the understudied realms of eukaryotes and viruses. Here we present VEBA 2.0 with key updates including a comprehensive clustered microeukaryotic protein database, rapid genome/protein-level clustering, bioprospecting, non-coding/organelle gene modeling, genome-resolved taxonomic/pathway profiling, long-read support, and containerization. We demonstrate VEBA’s versatile application through the analysis of diverse case studies including marine water, Siberian permafrost, and white-tailed deer lung tissues with the latter showcasing how to identify integrated viruses. VEBA represents a crucial advancement in microbiome research, offering a powerful and accessible software suite that bridges the gap between genomics and biotechnological solutions.

59 BASIC BIOLOGICAL SCIENCES↗

Unsupervised physics-informed disentanglement of multimodal data

Here, we introduce physics-informed multimodal autoencoders (PIMA) - a variational inference framework for discovering shared information in multimodal datasets. Individual modalities are embedded into a shared latent space and fused through a product-of-experts formulation, enabling a Gaussian mixture prior to identify shared features. Sampling from clusters allows cross-modal generative modeling, with a mixture-of-experts decoder that imposes inductive biases from prior scientific knowledge and thereby imparts structured disentanglement of the latent space. This approach enables cross-modal inference and the discovery of features in high-dimensional heterogeneous datasets. Consequently, this approach provides a means to discover fingerprints in multimodal scientific datasets and to avoid traditional bottlenecks related to high-fidelity measurement and characterization of scientific datasets.

97 MATHEMATICS AND COMPUTING↗

Surrogate models for development of unconventional shale reservoirs by an integrated numerical approach of hydraulic fracturing, flow and geomechanics, and machine learning

We develop well-completion surrogate models by taking an integrated workflow of hydraulic fracturing, flow, geomechanics, and machine learning simulation. There are three steps in the proposed workflow. First, history-matching processes are conducted with the field data including pumping and production data for characterization. Second, full-physics simulation is performed with various parameters of the field development (e.g., cluster spacing, clusters per stage, pumping rates and times, amount of proppant, and well spacing) to generate multiple simulation results by changing the parameters of the completion design with well-known hydraulic fracturing, reservoir, geomechanics simulators to calculate fracture geometry, reservoir depressurization, induced stress changes. The workflow is demonstrated over a field in the Southern Midland Basin. Here, we take two completion scenarios: a single well case followed by a multi-well case. Finally, a Long Short-Term Memory (LSTM) machine learning algorithm is employed to create surrogate models that can replicate the full-physics simulation results. Furthermore, results show that the trained models applied in the single well and multi-well cases for a particular geological system can provide good accuracy close to those provided by full-physics simulations. Specifically, the site-specific surrogate models can predict fracture parameters (length, height, and surface area) and cumulative production accurately with computational efficiency, suggesting our proposed workflow can be used as a pragmatic tool for expediting the well completion optimization process.

Geomechanics↗

DESI peculiar velocity survey – Fundamental Plane

The Dark Energy Spectroscopic Instrument (DESI) peculiar velocity survey aims to measure the peculiar velocities of early- and late-type galaxies within the DESI footprint using both the Fundamental Plane and optical Tully–Fisher relations. Direct measurements of peculiar velocities can significantly improve constraints on the growth rate of structure, reducing uncertainty by a factor of approximately 2.5 at redshift 0.1 compared to the DESI Bright Galaxy Survey’s redshift space distortion measurements alone. We assess the quality of stellar velocity dispersion measurements from DESI spectroscopic data. These measurements, along with photometric data from the Legacy Survey, establish the Fundamental Plane relation and determine distances and peculiar velocities of early-type galaxies. During survey validation, we obtain spectra for 6698 unique early-type galaxies, up to a photometric redshift of 0.15. 64 per cent of observed galaxies (4267) have relative velocity dispersion errors below 10 per cent. This percentage increases to 75 per cent if we restrict our sample to galaxies with spectroscopic redshifts below 0.1. We use the measured central velocity dispersion, along with photometry from the DESI Legacy Imaging Surveys, to fit the Fundamental Plane parameters using a 3D Gaussian maximum likelihood algorithm that accounts for measurement uncertainties and selection cuts. In addition, we conduct zero-point calibration using the absolute distance measurements to the Coma cluster, leading to a value of the Hubble constant, H 0 = 76.05 ± 0.35 (statistical) ±0.49 (systematic Fundamental Plane) ±4.86 (statistical due to calibration) km s –1 Mpc –1 ⁠. This H 0 value is within 2σ of Planck cosmic microwave background results and within 1σ of other low-redshift distance indicator-based measurements.

cosmological parameters↗

Fast and Scalable FFT-Based GPU-Accelerated Algorithms for Block-Triangular Toeplitz Matrices with Application to Linear Inverse Problems Governed by Autonomous Dynamical Systems

In this work, we present an efficient and scalable algorithm for performing matrix-vector multiplications (matvecs) for block Toeplitz matrices. Such matrices, which are shift-invariant with respect to their blocks, arise in the context of solving inverse problems governed by autonomous systems, and time-invariant systems in particular. In this article, we consider inverse problems that infer unknown parameters from observational data of a linear time-invariant dynamical system given in the form of partial differential equations (PDEs). Matrix-free Newton-conjugate-gradient methods are often the gold standard for solving these inverse problems, but they require numerous actions of the Hessian on a vector. Matrix-free adjoint-based Hessian matvecs require solution of a pair of linearized forward/adjoint PDE solves per Hessian action, which may be prohibitive for large-scale inverse problems. Time invariance of the forward PDE problem leads to a block Toeplitz structure of the discretized parameter-to-observable (p2o) map defining the mapping from inputs (parameters) to outputs (observables) of the PDEs. This block Toeplitz structure enables us to exploit two key properties: (1) compact storage of the p2o map and its adjoint, and (2) efficient fast Fourier transform–based Hessian matvecs. The proposed algorithm is mapped onto large multi-GPU clusters and achieves more than 80% of peak bandwidth on NVIDIA A100 GPUs. Excellent weak scaling is shown for up to 48 A100 GPUs. For the targeted problems, the implementation executes Hessian matvecs within fractions of a second, which is orders of magnitude faster than can be achieved by conventional matrix-free Hessian matvecs via forward/adjoint PDE solves.

97 MATHEMATICS AND COMPUTING↗

Flow annealed importance sampling bootstrap meets differentiable particle physics

High-energy physics requires the generation of large numbers of simulated data samples from complex but analytically tractable distributions called matrix elements. Surrogate models, such as normalizing flows, are gaining popularity for this task due to their computational efficiency. We adopt an approach based on flow annealed importance sampling bootstrap (FAB) that evaluates the differentiable target density during training and helps avoid the costly generation of training data in advance. We show that FAB reaches higher sampling efficiency with fewer target evaluations in high dimensions in comparison to other methods.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Linear-Scaling Asymmetric Triples Correction through the Solution of the DLPNO–CCSD Lambda Equations: DLPNO–CCSD(T) Λ

In this research, we derive equations for solving for the stationary points of the DLPNO–CCSD Lagrangian, in the t 1 -transformed formalism introduced earlier and as currently implemented in the P SI 4 quantum chemistry software package. These lambda equations in the local pair natural orbital basis allow for the evaluation of CCSD(T) Λ energetics with linear-scaling computational effort, also known as the asymmetric triples correction. This DLPNO–CCSD(T) Λ method allows for accurate triples contributions to be computed for larger molecules, especially in cases that CCSD(T) is known to be insufficient, such as with multireference systems and bond-breaking systems. We showcase the accuracy of our code on reaction energies, barrier heights, and noncovalent interaction energies. Also showcased are the capabilities of our code by evaluating DLPNO–CCSD(T) Λ energetics on large noncovalent dimers up to 112 atoms, as well as a rhodium catalyst complex containing 66 atoms.

Cluster chemistry↗

Exploring Architectural-Aware Affinity Policies in Modern HPC Runtimes

Modern commodity and High-Performance Computing (HPC) systems are evolving with complex CPU architectures. These architectures now feature higher core and NUMA domain counts and implement features such as hyperthreading. When considering significant differences in hardware configurations, library availability, and hardware-tailored system/software stacks, which could substantially vary from one system to another, performance portability is hard to achieve. Throughout the years, this trend resulted in an increasingly high burden on application developers to fine-tune their workloads for each architecture. This work explores how hardware-dependent aspects such as locality/process/thread affinity affect performance in modern CPU architectures. We focus our study on the Global Memory and Threading (GMT) distributed runtime system as a representative of Partitioned Global Address Space (PGAS) software stacks commonly adopted for productivity. In particular, to appreciate performance implications, we evaluate GMT’s thread affinity policies, and, introduce two new ones which exploit architectural awareness. Finally, we explore alternative NUMA configurations via different process bindings and perform a scalability study on three HPC clusters with varying CPU architectures and NUMA layouts. Our analysis indicates that more complex architectures are more affected by affinity and binding policies and highlights the importance of setting proper runtime configurations to achieve superior performance.

Di Dio Lavore, Ian↗