Search NASASearch

SEARCH · Search NASA

Results for “cluster computer”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Equation-of-motion internally contracted multireference unitary coupled-cluster theory

The accurate computation of excited states remains a challenge in electronic structure theory, especially for systems with a ground state that requires a multireference treatment. In this work, we introduce a novel equation-of-motion (EOM) extension of the internally contracted multireference unitary coupled-cluster framework (ic-MRUCC), termed EOM-ic-MRUCC. EOM-ic-MRUCC follows the transform-then-diagonalize approach, in analogy to its non-unitary counterpart. By employing a projective approach to optimize the ground state, the method retains additive separability and proper scaling with system size. We show that excitation energies are size-intensive if the EOM operator satisfies the “killer” and the projective conditions. Furthermore, we propose to represent changes in the reference state upon electron excitation via projected many-body operators that span the active orbitals and show that the EOM equations formulated in this way are invariant with respect to active orbital rotations. We test the EOM-ic-MRUCC method truncated to single and double excitations by computing the potential energy curves for several excited states of a BeH2 model system, the HF molecule, and water undergoing symmetric dissociation. Across these systems, our method delivers accurate excitation energies and potential energy curves within 5 mE h (∼0.14 eV) from full configuration interaction. Here, we find that truncating the Baker–Campbell–Hausdorff series to fourfold commutators contributes negligible errors (on the order of 10 −5 E h or less), offering a practical route to highly accurate excited-state calculations with reduced computational overhead.

74 ATOMIC AND MOLECULAR PHYSICS

Toward the “platinum standard” of quantum chemistry on quantum computers: Perturbative quadruple corrections in unitary coupled cluster theory

We propose a non-iterative, post-hoc correction to the unitary coupled cluster theory with the single, double, and triple excitations (UCCSDT) Ansatz, which considers the leading-order effects of neglected quadruple excitations. We present two ways to derive this correction, henceforth referred to as [Q-6], which leads to an improvement in the correlation energy shown to be truncated to sixth-order in many-body perturbation theory. Furthermore, a comparison between the UCC-based [Q-6] correction proposed in this work and analogous, “platinum standard” quadruple corrections proposed in conventional coupled cluster theory recognizes that [Q-6] is distinct from prior corrections since it is constructed entirely from internally connected components. Although trotterized (t) and full operator variants of UCCSDT exhibit errors in scans of small molecule potential energy surfaces that routinely exceed 1.6 mH, we find that t/UCCSDT[Q-6] is, nevertheless, able to achieve chemical accuracy as measured by the mean unsigned error.

Correlation energy

How Cloud is Accelerating Research at NREL

This presentation coincides with AWS's announcement of their new Parallel Computing Service (PCS) which allows for easy creation of HPC-style clusters in their AWS cloud computing platform. I helped them beta test this service before it was made generally available in August. AWS asked if we would be interested in discussing our experience with the PCS service, and our experience with HPC workloads in the cloud in general, so this slideshow discusses a brief history of scientific computing at NREL and shares a bit of our experiences and approach to utilizing cloud services for HPC-style workloads.

97 MATHEMATICS AND COMPUTING

Automated Classification of Vehicle Movements at Signalized Intersections Using Vehicle Trajectories

Accurate vehicle movement classification through signalized intersections is of paramount importance to the analysis of intersection performance and the optimization of traffic control strategies. Conventional techniques for tracking vehicle turning movements depend on infrastructure-based strategies like human counts, loop detectors, and video analytics, all of which are costly, prone to errors, and spatially constrained. High-frequency trajectory data can be utilized to determine vehicle movement patterns in a scalable and infrastructure-independent method due to the adoption of connected vehicles (CVs). In recent years, several studies have utilized connected vehicle data to generate performance measures. Most of the trajectory-based performance measures approaches, however, require map matching-i.e., extracting geospatial references from maps to identify the movements that individual vehicles make at a signalized intersection. These approaches are often time-consuming and hinder scalability since geographic features need to be provided for an analysis to be conducted. Map matching methods are prone to errors as different map versions change these geographic features. This research presents a novel automatic classification pipeline that uses CV trajectory data to classify vehicle movements at signalized crossings, specifically pass-through left-turn and right-turn maneuvers. The process starts by filtering trips that cross a spatial bounding box that has been defined at the target intersection. Approach and departure headings for each trajectory crossing the boundary are computed and are clustered together to identify dominant movements. The proposed algorithm is used to classify the movement of vehicles at 10 intersections in the state of California, and the results indicate that the algorithm can classify movements at these intersections with varying traffic volumes and road network configurations, all in a map-less framework with no need for conflation of vehicle trajectories to a digital base map.

24 POWER TRANSMISSION AND DISTRIBUTION

Poplar: a phylogenomics pipeline

Motivation Generating phylogenomic trees from the genomic data is essential in understanding biological systems. Each step of this complex process has received extensive attention and has been significantly streamlined over the years. Given the public availability of data, obtaining genomes for a wide selection of species is straightforward. However, analyzing that data to generate a phylogenomic tree is a multistep process with legitimate scientific and technical challenges, often requiring a significant input from a domain-area scientist. Results We present Poplar, a new, streamlined computational pipeline, to address the computational logistical issues that arise when constructing the phylogenomic trees. It provides a framework that runs state-of-the-art software for essential steps in the phylogenomic pipeline, beginning from a genome with or without an annotation, and resulting in a species tree. Running Poplar requires no external databases. In the execution, it enables parallelism for execution for clusters and cloud computing. The trees generated by Poplar match closely with state-of-the-art published trees. The usage and performance of Poplar is far simpler and quicker than manually running a phylogenomic pipeline. Availability and implementation Freely available on GitHub at https://github.com/sandialabs/poplar. Implemented using Python and supported on Linux.

Koning, Elizabeth [Sandia National Laboratories (S

Studying CPU and memory utilization of applications on Fujitsu A64FX and Nvidia Grace Superchip

ARM-based manycore CPU architectures are well-positioned to provide the rising memory throughput requirements of modern data intensive scientific applications in High Performance Computing (HPC). The Fujitsu A64FX CPU platform is based on the ARM v8.2A architecture, and is the processor of the flagship Japanese supercomputer - "Fugaku", which was previously ranked as the #1 supercomputer in the world according to the Top500 list. The Nvidia Grace superchip features 144 Neoverse V2 cores based on the ARMv9 architecture with 4x128b SVE2, providing exceptional computational power. The chip supports up to 480GB of memory, making it ideal for AI, machine learning, and scientific computing workloads. In this paper, we conduct a thorough performance exploration of a variety of parallel bandwidth-sensitive benchmarks and applications compiled with the native Fujitsu compiler on a Fugaku A64FX compute node and ARM (LLVM) Compiler on an NVIDIA Grace superchip compute node, engaging all the computational cores per cluster using OpenMP multithreading (assuming the cores can drive the available bandwidth). Our ultimate goals are to study the resource utilization of scientific applications and benchmarks on A64FX and Grace superchip, considering graph application scenarios ( GAP Benchmark suite) and eleven appli- cation proxies from the Rodinia heterogeneous benchmark suite (considering domains such as Data Mining, Bioinformatics, Fluid Dynamics, Pattern Recognition, etc.). Through exhaustive performance monitoring, we quantify the resource utilization of diverse OpenMP-based HPC applications on both the Fujitsu A64FX and the Nvidia Grace Superchip platforms.

benchmarking, Performance Analysis, High performan

Concordant Mode Approach (CMA): Vibrational Analysis of New and Upgraded Intermolecular Benchmarks for Noncovalent Bonding

The Concordant Mode Approach (CMA) is a novel method that offers tremendous potential for increasing the system size and the level of theory attainable in quantum chemical computations of molecular vibrational frequencies. To investigate the extension of CMA to intermolecular vibrations, computations with coupled cluster singles and doubles with perturbative triples theory [CCSD(T)] using two augmented correlation-consistent polarized-valence triple-ζ basis sets (aug-cc-pVTZ or h-aug-ccpVTZ) were performed on 17 prototypical loosely bound complexes of hydrogen-bonded, dispersion, and mixed character. These Level A results provide new and upgraded benchmarks for noncovalent bonding and a severe test for CMA vibrational analyses. The Level A target frequencies were recovered remarkably well using second-order Møller−Plesset perturbation theory (MP2) with h-aug-cc-pVTZ for generating the underlying (Level B) normal modes of the CMA scheme. Employing this Level B within the lowest-rung CMA-0A method reproduces the 435 benchmark frequencies with a mean absolute error (MAE) of 0.23 cm −1 and a corresponding standard deviation (σ) of 0.84 cm −1 ; strikingly, the corresponding subset of 106 interfragment frequencies exhibits MAE = 0.34 cm −1 and σ = 0.90 cm −1 . Subsequent application of the higher-rung CMA-2A scheme eliminates all outliers and reduces the overall MAE to a minuscule 0.08 cm −1 with the inclusion of only 3.0% of the off-diagonal couplings not accounted for by CMA-0A. Accordingly, the highly efficient CMA methodology proves to be robust even for vibrations on flat potential energy surfaces.

Aromatic compounds

panhandle

A project to provide user activity monitoring for High Performance Computing systems and clusters. The goal is to provide effective user activity monitoring with minimal performance impact on the host running this service.

McGee, David [@LANL @USMC @DoD]

Using containers to speed up development, to run integration tests and to teach about distributed systems

GlideinWMS is a workload manager provisioning resources for many experiments including CMS and DUNE. The software is distributed both as native packages and specialized production containers. Following an approach used in other communities like web development we built our workspaces, system-like containers to ease development and testing. Developers can change the source tree or check out a different branch and quickly reconfigure the services to see the effect of their changes. In this paper, we’ll talk about what differentiates workspaces from other containers. We’ll describe our base system composed of three containers. A one-node cluster including a compute element and a batch system. A GlideinWMS Factory controlling pilot jobs. And a scheduler and Frontend, to submit jobs and provision resources. Additional containers can be used for optional components. This system can easily run on a laptop and we’ll share our evaluation of different container runtimes, with an eye for ease of use and performance. Finally, we’ll talk about our experience as developers and with students. The GlideinWMS workspaces are easily integrated with IDEs like VS Code, simplifying debugging and allowing development and testing of the system also when offline. They simplified the training and onboarding of new team members and Summer interns. And they were useful in workshops where students could have first-hand experience with the mechanisms and components that, in production, run millions of jobs.

Mambelli, Marco

Using Containers to Speed Up Development, to Run Integration Tests and to Teach About Distributed Systems

GlideinWMS is a workload manager provisioning resources for many experiments, including CMS and DUNE. The software is distributed both as native packages and specialized production containers. Following an approach used in other communities like web development, we built our workspaces, system-like containers to ease development and testing. Developers can change the source tree or check out a different branch and quickly reconfigure the services to see the effect of their changes. In this paper, we will talk about what differentiates workspaces from other containers. We will describe our base system, composed of three containers: a one-node cluster including a compute element and a batch system, a GlideinWMS Factory controlling pilot jobs, and a scheduler and Frontend to submit jobs and provision resources. Additional containers can be used for optional components. This system can easily run on a laptop, and we will share our evaluation of different container runtimes, with an eye for ease of use and performance. Finally, we will talk about our experience as developers and with students. The GlideinWMS workspaces are easily integrated with IDEs like VS Code, simplifying debugging and allowing development and testing of the system even when offline. They simplified the training and onboarding of new team members and summer interns. And they were useful in workshops where students could have first-hand experience with the mechanisms and components that, in production, run millions of jobs.

Mambelli, Marco [Fermilab] (ORCID:0000000294892681

A new biogeochemical modelling framework (FLaMe-v1.0) for lake methane emissions on the regional scale: development and application to the European domain

This study presents a new physical-biogeochemical modelling framework for simulating lake methane (CH 4 ) emissions at regional scales. The new model, FLaMe-v1.0 (Fluxes of Lake Methane), rests on an innovative, computationally efficient lake clustering approach that enables the simulation of CH 4 emissions across a large number of lakes. Building on the Canadian Small Lake Model (CSLM) that simulates the lake physics, we develop a suite of biogeochemical modules to simulate transient dynamics of organic Carbon (C), Oxygen (O 2 ), and CH 4 . We first test the performance of FLaMe-v1.0 by analyzing physical and biogeochemical processes in two theoretical lakes with characteristics that can be considered representative for many lakes (an oligotrophic, deep lake driven by cold climate versus a eutrophic, shallow lake driven by warm climate). Next, we evaluate the model by comparing simulated and observed timeseries of CH 4 emissions in four well-surveyed lakes. We then apply FLaMe-v1.0 at the European scale to evaluate simulated diffusive and ebullitive lake CH 4 fluxes against in-situ measurements in both boreal and central European regions. Finally, we provide a first assessment of the spatio-temporal variability in CH 4 emissions from European lakes with a surface area comprised between 0.1–1000 km 2 (n= 108 407, total area = 1.33 × 105 km 2 ), indicating a total emission of 0.97 ± 0.23 Tg CH 4 yr −1 , with the uncertainty constrained by combining FLaMe-v1.0 and machine learning techniques. Moreover, 30 % and 70 % of these CH 4 emissions are through diffusive and ebullitive pathways, respectively. Annually averaged CH 4 emission rates per unit lake area during 2010–2016 have a South-to-North decreasing gradient, resulting in a mean over the European domain as 7.39 g CH 4 m −2 yr −1 . Our simulations reveal a strong seasonality (with ice-blocking effects accounted for) in European lake CH 4 emissions, with nearly ten times higher emissions during late summer than during winter. This pronounced seasonal variation highlights the importance of accounting for the sub-annual variability in CH 4 emissions to accurately constrain regional CH 4 budgets. In the future, FLaMe-v1.0 could be embedded into Earth System Models to investigate the feedback between climate warming and global lake CH 4 emissions.

Maisonnier, Manon [Free Univ. of Brussels (Belgium

precipbestats (c0)

Best estimates of precipitation from ARM instruments derived through clustering and other computational techniques.

54 ENVIRONMENTAL SCIENCES

precipbetseries (c1)

Best estimates of precipitation from ARM instruments derived through clustering and other computational techniques.

54 ENVIRONMENTAL SCIENCES

precipbetseries (c0)

Best estimates of precipitation from ARM instruments derived through clustering and other computational techniques.

54 ENVIRONMENTAL SCIENCES

Exploiting a Shortcoming of Coupled-Cluster Theory: The Extent of Non-Hermiticity as a Diagnostic Indicator of Computational Accuracy

The fundamental non-Hermitian nature of the forms of the coupled-cluster (CC) theory widely used in quantum chemistry has usually been viewed as a negative, but the present paper shows how this can be used to an advantage. Specifically, the non-symmetric nature of the reduced one-particle density matrix (in the molecular orbital basis) is advocated as a diagnostic indicator of computational quality. In the limit of the full coupled-cluster theory [which is equivalent to full configuration interaction (FCI)], the electronic wave function and correlation energy are exact within a given one-particle basis set, and the symmetric character of the exact density matrix is recovered. The extent of the density matrix asymmetry is shown to provide a measure of “how difficult the problem is” (like the well-known T 1 diagnostic), but its variation with the level of theory also gives information about “how well this particular method works”, irrespective of the difficulty of the problem at hand. The proposed diagnostic is described and applied to a select group of small molecules, and an example of its overall utility for the practicing quantum chemist is illustrated through its application to the beryllium dimer (Be 2 ). Future application of this idea to excited states, open-shell systems, and symmetry-breaking problems and an extension of the method to the two-particle density are then proposed.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Computational ranking identifies Plexin-B2 in circulating tumor cell clustering with monocytes in breast cancer metastasis

Abstract Multicellular circulating tumor cell (CTC) clusters can be up to 50 times more efficient than single CTCs in mediating viable metastasis. Here, combining computational ranking and functional determination, we identify the transmembrane protein Plexin-B2 (PLXNB2) as one of the top molecular targets associated with unfavorable distant metastasis-free survival, showing enriched expression in CTC clusters versus single CTCs from patients with advanced breast cancer (mostly female). Loss of PLXNB2 (Plxnb2) reduces the formation of homotypic tumor cell clusters and heterotypic tumor-myeloid cell clusters, reducing spontaneous metastases in female mice bearing human (mouse) breast cancer. Interactions of PLXNB2 with its ligands SEMA4C on tumor cells and SEMA4A on myeloid cells (monocytes) promote homotypic and heterotypic CTC cluster formation, respectively, thereby driving lung metastasis. Global proteomic analysis reveals downstream effectors of the PLXNB2 pathway associated with tumor cell clustering. Thus, PLXNB2 is a therapeutic target for preventing new metastasis in breast cancer.

Science & Technology - Other Topics

Investigating the crust of neutron stars with neural-network quantum states

An accurate description of low-density nuclear matter is crucial for explaining the physics of neutron star crusts. In the density range between approximately 0.01 fm −3 and 0.1 fm −3 , matter transitions from neutron-rich nuclei to various higher-density pasta shapes, before ultimately reaching a uniform liquid. In this work, we introduce a variational Monte Carlo method based on a neural Pfaffian-Jastrow quantum state, which allows us to model the transition from the liquid phase to neutron-rich nuclei microscopically. At low densities, nuclear clusters dynamically emerge from the microscopic interactions among protons and neutrons, which we model based on pionless effective field theory. Our variational Monte Carlo approach represents a significant improvement over the state-of-the-art auxiliary-field diffusion Monte Carlo method, which is severely hindered by the fermion-sign problem in this low-density regime and cannot capture the onset of clusters. In addition to computing the energy per particle of symmetric nuclear matter and pure neutron matter, we analyze an intermediate isospin-asymmetry configuration to elucidate the formation of nuclear clusters. We also provide evidence that the presence of such nuclear clusters influences the amount of protons in the crust compared to protons in beta-equilibrated, neutrino-transparent matter.

Nuclear astrophysics