Search NASASearch

SEARCH · Search NASA

Results for “cluster computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Hierarchical Truncations for Many-Body Expansion Potentials

In this work, a new strategy to truncate high-order terms in the many-body expansion (MBE) is proposed. This new approach, which we call a hierarchical many-body expansion (HMBE), is based on a hierarchical partition of the system into multitier fragments and can in principle be applied to any large molecular system. Numerical tests on a series of (H 2 O) 64 structures are presented, demonstrating satisfactory relative energies between the clusters and binding energies of individual clusters compared with full-cluster calculations, with significantly fewer high-order terms computed than conventional MBE. The hierarchical truncation can be augmented by certain many-body terms for fragments at the interface between the partitions (called “Schengen terms”) to further improve accuracy. This work establishes the HMBE scheme as a promising framework to model very large systems (e.g., proteins), which are naturally built on a hierarchical structure.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

The final WaZP galaxy cluster catalog of the Dark Energy Survey and comparison with SZE data

In this work, we present and characterize the galaxy cluster catalog detected by the WaZP cluster finder, which is not based on red-sequence identification, on the full six years of observations of the Dark Energy Survey (DES-Y6). The full catalog contains over 400k detected clusters with richnesses, Ngals, above 5 and that reach redshifts up to 1.3. We also provide a version of the catalog where the observation depth and richness computation are homogenized to be used for cosmology, containing 33k rich (Ngals >25) clusters. We compare our results with the previous WaZP catalog obtained from the DES first-year data release (DES-Y1). We find that essentially all clusters within the common footprint and depth limit are recovered. The deeper observations on DES-Y6 and the more complete available spectroscopic redshift sample lead to improvements in the redshifts of the clusters, resulting in an average scatter of 1.4% and offset of 0.2%. The optical clusters are also cross-matched with Sunyaev Zel'dovich Effect (SZE) cluster samples detected by the South Pole Telescope (SPT) and the Atacama Cosmology Telescope (ACT). We find that essentially all SZE clusters with reasonable overlapping footprint have a corresponding WaZP cluster. Conversely, 90% of the optical detections with richness greater than 150 have a counterpart in the deeper regions of the SZE surveys. Based on cross-match with the SZE catalogs, we also find that 15-20% of the SZE matched systems have more than one possible WaZP counterpart at the same redshift and within the SZE R500c, indicating possible interacting or unrelaxed systems. Finally, given the optical and SZE beams, WaZP and SZE centerings are found to be consistent. A more detailed study of the SZE-WaZP mass-richness relation will be presented in a separate paper.

Benoist, C. [OCA, Nice, Lab. Lagrange; LIneA, Rio

Environmental controls on the kinetics of iron-sulfur cluster nucleation and nanoparticle formation

Anoxic, sulfidic conditions have been prevalent since the early Proterozoic and favor aqueous iron-sulfur (FeS aq ) clusters as a major fraction of the soluble, reduced iron and sulfur pool. FeS aq cluster formation and nucleation is driven by the high affinity between ferrous iron (Fe(II)) and sulfide (HS − ), ultimately yielding particles that precipitate as iron sulfide minerals. FeS aq clusters were recently shown to be bioavailable sources of iron and sulfur for a variety of anaerobes, yet little is known of the factors that influence the kinetics of their formation and nucleation. Here we apply computational and spectroscopic approaches to investigate the dynamics of FeS aq nucleation, cluster growth, precipitation, and redissolution as a function of Fe(II)/HS − concentration, temperature, and pH. Experiments were conducted under excess HS − to mimic euxinic conditions common to contemporary anaerobic aquatic ecosystems and those of the Proterozoic. Density functional theory calculations reveal the key role of water oxygen-iron interactions in stabilizing small FeS aq clusters and promoting solubility. Dynamic light scattering revealed a concentration-dependent increase in the kinetics of FeS aq nucleation and cluster aggregation. Increasing temperature promoted FeS aq cluster nucleation and aggregation while also enhancing dissolution. Alkaline pH also promoted FeS aq nucleation and cluster aggregation. At 25 °C, pH 7.0, and at reactant concentrations of 30 µM, FeS aq clusters < 10 nm in diameter remained in solution for > 2 h. These results underscore the importance of temperature, pH, and reactant concentration in the kinetics of FeS aq nucleation and cluster growth that, in turn, influence their bioavailability in anaerobic ecosystems.

Aquatic ecosystems

A Comparison of Electronic Structure Methods for Predicting the Hydrogenation Energies of Candidate Molecules for Hydrogen Storage

The development of novel energy materials and fuels is required to expand current available energy sources. Aiming to reach this goal, there is growing interest in using molecular hydrogen as an energy carrier due to its abundance and high energy density. Liquid organic hydrogen carriers (LOHCs) are a promising route to the large-scale storage and transport of hydrogen for use in the energy economy. The search for thermodynamically viable LOHC molecules for real world use has led to a set of constraints on the dehydrogenation enthalpy and the minimum gravimetric hydrogen capacity. These constraints allow one to formulate the search for an ideal LOHC candidate molecule as an optimization problem well suited to the strengths of machine learning and artificial intelligence computational approaches. A critical barrier to a large-scale, high-throughput screening of LOHC candidate molecules is the lack of reliable training data. Computational electronic structure methods including density functional theory, coupled cluster approximations, and diffusion Monte Carlo can be used to provide training data where experimental data are either unreliable or do not exist. In this work, we use these methods to calculate the dehydrogenation energies and enthalpies of candidate LOHC molecules.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Machine learning at the Spallation Neutron Source accelerator and target

We describe the ongoing efforts to apply Machine Learning techniques to improve the performance of our accelerator and target. Specially, we are looking to minimize halo beam losses in the absence of a proper physics model, automatically detect and log anomalies in the target support systems such as cooling, and detect and prevent errant beam pulses in the linac. We also describe the infrastructure we use to acquire and stream data to the GPU cluster for training, our code development cycle, and edge computing for model inference. To minimize halo beam losses, we use a Reinforcement Learning technique tested on a virtual accelerator. The target anomaly detection is trained on archived data using incomplete physics models and is made part of the existing target reporting system. The errant beam prevention analyzes beam current and beam phase waveforms as well as accelerator configuration data to predict errant pulses. We also develop continual learning to adapt to changes in the accelerator.

Accelerator Physics

Thorium Monosilicide, ThSi: An Experimental and Theoretical Study

The present theoretical and experimental combination study investigates the ThSi molecule in detail. Computationally, we utilized high-level multireference and coupled-cluster levels of theory conjoined with large correlation consistent basis sets to study a series of electronic and spin–orbit states of ThSi. Here, we report potential energy curves (PECs), electron configurations at equilibrium distances, spectroscopic constants, energetics, and spin–orbit coupling effects for 16 electronic states of ThSi. The studied 16 electronic states are arranged tightly within 0.9 eV, highlighting the complexity of the electronic spectrum of ThSi. The ground electronic state of ThSi is a single-reference 1 1 Σ + state that derives from the 1σ 2 2σ 2 1π 4 electronic configuration. The Ω = 0 + spin–orbit ground state of ThSi is composed of 1 1 Σ + (47%) and 13Π (44%) electronic states. Our measured bond energy (D0) of ThSi, obtained using resonant two-photon ionization (R2PI) spectroscopy is 3.146(4) eV, where the assigned error limit is given in parentheses in units of the last quoted digits. The computed D0 of ThSi (Ω = 0 + ) at the CBS-C-CCSD(T)-δT(Q)-δDK-δSO level (3.181 eV) is in good agreement with the experimental value. Our derived enthalpy of formation for ThSi, Δ f H 0K o (ThSi(g)), is 971.8(6.0) kJ/mol. Finally, we have performed density functional theory (DFT) calculations for ThSi(1 1 Σ + ) using 16 exchange correlation functionals that span multiple rungs of “Jacob’s ladder” of density functional approximation (DFA) to assess the DFT errors on D 0 , r e , and ω e of ThSi with respect to experimental and ab initio coupled-cluster values.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Hiperclust

This software leverages transfer learning to analyze atom probe tomography (APT) data. It is trained on synthetic data and then applies this knowledge to predict the optimal number of clusters for a given APT dataset. Initially, the software used preliminary clustering to estimate the general structure of the data. Based on this, it provides suggestions for key parameters like minimum cluster size and minimum number of points. These parameters are critical for algorithms like HDBSCAN, ensuring accurate cluster formation without the need for trial-and-error testing. The software runs on High-Performance computing (HPC) systems, enabling fast, scalable analysis of large APT datasets, ultimately saving time and improving the reliability of clustering outcomes.

Tang, Yalei [Idaho National Laboratory (INL), Idah

Control of Permanent Porosity in Type 3 Porous Liquids via Solvent Clustering

Porous liquids (PLs) are an exciting new class of materials for carbon capture due to their high gas adsorption capacity and ease of industrial implementation. They are composed of sorbent particles suspended in a nonadsorbed solvent, forming a liquid with permanent porosity. While PLs have a vast number of potential compositions based on the number of solvents and sorbent materials available, most of the research has been focused on the selection of the sorbent rather than the solvent. Therefore, PL design criteria on the supramolecular structures of the solvent are explored to create a fundamental understanding of how the solvent enables PL formation for rapid discovery of new PL compositions. Atomistic molecular dynamics simulation of eight solvents with a range of molecular sizes, shapes, and intramolecular bonding was performed, identifying that the shape and size of molecular clusters formed in the solvent are the driving predictor of PL formation rather than the size of the individual solvent molecule. The results demonstrate a significant departure from common approaches to PL formation based on the steric exclusion of solvent molecules from the sorbent via the size of the pore aperture. A modeling and experimental validation study further supports these findings. In conclusion, through this computational material design study, a previously unexplored mechanism in PL formation, solvent–solvent clustering, is identified as a critical factor for the accelerated discovery of liquid phase carbon capture materials.

Carbon capture

Deploying and Operating CephFS for Scientific Applications at Fermilab

Fermilab has been running a Ceph cluster in production for several years to support high-throughput scientific computing. Our primary use case is CephFS, which serves interactive data analysis workloads, with growing interest in using RGW for scalable object storage of scientific datasets. In this talk, we'll share lessons learned from successfully deploying and maintaining our Ceph cluster with cephadm, including challenges faced, performance tuning, and operational practices. We'll also present custom tools we've developed to streamline monitoring and management and discuss how Ceph fits into our broader storage architecture for large-scale scientific research.

Peisker, Alison [Fermilab]

Scalable edge clustering of dynamic graphs via weighted line graphs

Timestamped relational datasets consisting of records (or connections) between pairs of entities are ubiquitous in network science. For applications like peer-to-peer communication, email, various social network interactions, and computer network security, it is useful to organize these records into groups based on how and when they are occurring. Weighted line graphs offer a natural way to model how records are related in such datasets but for large real-world graph topologies, building and utilizing the line graph is prohibitively expensive. Here, we present the framework to cluster the edges of a dynamic graph via the associated line graph that contains two major contributions. The first is a method to work with the line graph implicitly and the second is a distributed scale implementation of an agglomerative hierarchical graph clustering algorithm. We outline a novel hierarchical dynamic graph edge clustering approach that efficiently breaks massive relational datasets into small sets of edges containing events at various timescales. This is in stark contrast to traditional graph clustering algorithms that prioritize highly connected (clique-like) community structures. Our approach relies on constructing a sufficient subgraph of a weighted line graph and applying a hierarchical agglomerative clustering. This approach is related to scalable techniques from spatial clustering, nonlinear-dimension reduction, topological data analysis, and draws particular inspiration from HDBSCAN. As an edge clustering, this method yields an overlapping node clustering. Our algorithm is parallelizable and we demonstrate efficient clustering of a billion-scale, real-world dynamic graph into small edge sets that correlate in topology and time. The entire clustering process for a graph with tens of billions of edges takes just a few minutes of run time on 256 nodes of a distributed compute environment. We argue how the output of the edge clustering is useful for a multitude of data visualization and powerful machine learning tasks, both involving the original massive dynamic graph data and metadata associated with the nodes and edges. Finally, we describe how this approach can be extended to dynamic hypergraphs and dynamic graphs/hypergraphs with unstructured data living on vertices and edges.

Data Analysis

Exploring the exact limits of the real-time equation-of-motion coupled cluster cumulant Green’s functions

In this paper, we analyze the properties of the recently proposed real-time equation-of-motion coupled-cluster (RT-EOM-CC) cumulant Green’s function approach [Rehr et al., J. Chem. Phys. 152, 174113 (2020)]. We specifically focus on identifying the limitations of the original time-dependent coupled cluster (TDCC) ansatz and propose an enhanced double TDCC ansatz, ensuring the exactness in the expansion limit. In addition, we introduce a practical cluster-analysis-based approach for characterizing the peaks in the computed spectral function from the RT-EOM-CC cumulant Green’s function approach, which is particularly useful for the assignments of satellite peaks when many-body effects dominate the spectra. Our preliminary numerical tests focus on reproducing, approximating, and characterizing the exact impurity Green’s function of the three-site and four-site single impurity Anderson models using the RT-EOM-CC cumulant Green’s function approach. The numerical tests allow us to have a direct comparison between the RT-EOM-CC cumulant Green’s function approach and other Green’s function approaches in the numerical exact limit.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Cluster-Graph Fingerprinting: A Framework for Quantitative Analysis of Machine-Learned Interatomic Model Training and Simulation Data

Machine-learned interatomic models represent a significant advancement in simulation methods, extending the predictive ability of first-principles methods to previously inaccessible length and time scales. However, the data-driven nature of these models can lead to difficult-to-detect errors that can compromise prediction accuracy. To address this challenge, we introduce a novel fingerprinting approach based on the Chebyshev Interaction Model for Efficient Simulation (ChIMES) ML-IAM graph-based descriptor. Our strategy enables efficient and statistically rigorous analysis of system configurations used in ML-IAM training and those generated by their application, e.g., in molecular dynamics simulations. We demonstrate that these fingerprints can effectively assess novelty of a configuration relative to an existing data set and determine dissimilarity among individual configurations, which are two key tasks in workflows for active learning-based ML-IAM training, data set curation, and on-the-fly uncertainty quantification.

36 MATERIALS SCIENCE

Effective many-body interactions in reduced-dimensionality spaces through neural network models

Accurately describing properties of challenging problems in physical sciences often requires complex mathematical models that are unmanageable to tackle head on. Therefore, developing reduced-dimensionality representations that encapsulate complex correlation effects in many-body systems is crucial to advance the understanding of these complicated problems. However, a numerical evaluation of these predictive models can still be associated with a significant computational overhead. To address this challenge, in this paper we discuss a combined framework that integrates recent advances in the development of active-space representations of coupled cluster (CC) downfolded Hamiltonians with neural network approaches. The primary objective of this effort is to train neural networks to eliminate the computationally expensive steps required for evaluating hundreds or thousands of Hugenholtz diagrams, which correspond to multidimensional tensor contractions necessary for evaluating a many-body form of downfolded effective Hamiltonians. Using small molecular systems (the H 2 O and HF molecules) as examples, we demonstrate that training neural networks employing effective Hamiltonians for a few nuclear geometries of molecules can accurately interpolate or extrapolate their forms to other geometrical configurations characterized by different intensities of correlation effects. We also discuss differences between effective interactions that define CC downfolded Hamiltonians with those of bare Hamiltonians defined by Coulomb interactions in the active spaces. Published by the American Physical Society 2024

97 MATHEMATICS AND COMPUTING

FORESTR: Finding, Organizing, Representing, Explaining, Summarizing, and Thinning Random forests

Random forests have become popular models used for data driven predictions. As a result, random forests are currently used or being considered for high-consequence mission applications in national security, such as the prediction of yield from optical signals and malware detection. While random forests may provide accurate predictions, the complexity of the algorithm causes a lack of interpretability. Random forests are an ensemble of regression or decision trees. Individual regression and decision trees are interpretable, but ensembles are inherently difficult to interpret due to the compilation of many models. We aim to increase the interpretability of random forests by finding patterns in the ensemble of trees that can be used to “thin” (or remove) trees. As a starting point, in this report, we develop a new distance metric for quantifying the similarity between trees based on their topologies (i.e., shapes). We base the metric on a novel distance metric for graphs that is a proper mathematical distance, is invariant to transformations, has registration between graphs, and computes topological evolutions between graphs. We use the tree distance metric to compute tree statistics such as a “mean tree” and to identify clusters of trees. We apply the developed methodology to a toy dataset and a mission relevant product inspection dataset to demonstrate how the metric can provide insight into random forests. Furthermore, we discuss the limitations of the approach and ideas for future research into how the metric could be used as a thinning tool to develop less complex models.

97 MATHEMATICS AND COMPUTING

Speedup of UEDGE Parameter Scans Using Machine-Learning Optimized OpenMP Parallelization and a Continuation Solver

This article presents the OpenMP parallelization of the preconditioning Jacobian assembly and right‐hand side residual evaluation in UEDGE. A continuation algorithm, utilizing the internal NKSOL implicit Jacobian‐Free Newton‐Krylov solver to efficiently scan physical parameters, is also presented. The implemented parallelization reduces the computational time for a benchmark scan run on 32 threads by compared to the serial version when using trained random forest regression models to identify the optimal decomposition of the system of equations. Random forest regression models applied to the UEDGE time‐dependent and continuation solver algorithms did not yield meaningful improvement in computational performance. A benchmark DIII‐D gas injection rate scan in the 0.35–0.75 kA interval, performed on a test cluster using the parallelized code and continuation solver, produced 1066 steady‐state solutions with a 22 s average wall‐clock computational time per steady‐state solution.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Correlated Anion Disorder in Heteroanionic Cubic TiOF 2

Resolving anion configurations in heteroanionic materials is crucial for understanding and controlling their properties. For anion-disordered oxyfluorides, conventional Bragg diffraction cannot fully resolve the anionic structure, necessitating alternative structure determination methods. We have investigated the anionic structure of anion-disordered cubic (ReO 3 -type) TiOF 2 using X-ray pair distribution function (PDF), 19 F MAS NMR analysis, density functional theory (DFT), cluster expansion modeling, and genetic-algorithm structure prediction. Our computational data predict short-range anion ordering in TiOF 2 , characterized by predominant cis-[O 2 F 4 ] titanium coordination, resulting in correlated anion disorder at longer ranges. To validate our predictions, we generated partially disordered supercells using genetic-algorithm structure prediction and computed simulated X-ray PDF data and 19 F MAS NMR spectra, which we compared directly to experimental data. To construct our simulated 19 F NMR spectra, we derived new transformation functions for mapping calculated magnetic shieldings to predicted magnetic chemical shifts in titanium (oxy)fluorides, obtained by fitting DFT-calculated magnetic shieldings to previously published experimental chemical shift data for TiF 4 . We find good agreement between our simulated and experimental data, which supports our computationally predicted structural model and demonstrates the effectiveness of complementary experimental and computational techniques in resolving anionic structure in anion-disordered oxyfluorides. From additional DFT calculations, we predict that increasing anion disorder makes lithium intercalation more favorable by, on average, up to 2 eV, highlighting the significant effect of variations in short-range order on the intercalation properties of anion-disordered materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Generalized representative structures for atomistic systems

A new method is presented to generate atomic structures that reproduce the essential characteristics of arbitrary material systems, phases, or ensembles. Previous methods allow one to reproduce the essential characteristics (e.g. the chemical disorder) of a large random alloy within a small crystal structure. The ability to generate small representations of random alloys, along with the restriction to crystal systems, results from using the fixed-lattice cluster correlations to describe structural characteristics. A more general description of the structural characteristics of atomic systems is obtained using complete sets of atomic environment descriptors. These are used within for generating representative atomic structures without restriction to fixed lattices. A general data-driven approach is provided here utilizing the atomic cluster expansion (ACE) basis. The N-body ACE descriptors are a complete set of atomic environment descriptors that span both chemical and spatial degrees of freedom and are used within for describing atomic structures. The generalized representative structure (GRS) method presented within generates small atomic structures that reproduce ACE descriptor distributions corresponding to arbitrary structural and chemical complexity. It is shown that systematically improvable representations of crystalline systems on fixed parent lattices, amorphous materials, liquids, and ensembles of atomic structures may be produced efficiently through optimization algorithms. With the GRS method, we highlight reduced representations of atomistic machine-learning training datasets that contain similar amounts of information and small 40–72 atom representations of liquid phases. The ability to use GRS methodology as a driver for informed novel structure generation is also demonstrated. The advantages over other data-driven methods and state-of-the-art methods restricted to high-symmetry systems are highlighted.

atomic cluster expansion