Search NASA⌕ Search

SEARCH · Search NASA

Results for “Linux cluster”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Andes Data Analysis System at the Oak Ridge Leadership Computing Facility

Andes is a (704)-node commodity-type Linux® cluster. Each of Andes’s 704 nodes contain two 16-core 3.0 GHz AMD EPYC 7302 processors with AMD’s Simultaneous Multithreading (SMT) Technology and 256GB of main memory. Andes also has nine large memory GPU nodes. These nodes each have 1TB of main memory and two NVIDIA K80 GPUs with two 14-core 2.30 GHz Intel Xeon processors with HT Technology.

AMD EPYC↗

Charon User Manual: v. 2.2 (revision1)

This manual gives usage information for the Charon semiconductor device simulator. Charon was developed to meet the modeling needs of Sandia National Laboratories and to improve on the capabilities of the commercial TCAD simulators; in particular, the additional capabilities are running very large simulations on parallel computers and modeling displacement damage and other radiation effects in significant detail. The parallel capabilities are based around the MPI interface which allows the code to be ported to a large number of parallel systems, including linux clusters and proprietary “big iron” systems found at the national laboratories and in large industrial settings.

42 ENGINEERING↗

Searching for the Most Harmful Field Errors in the HSR IR Superconducting Magnets

In this project, we improve beam stability for the Electron-Ion Collider. Magnetic field errors can reduce beam stability, making it essential to identify the field errors that have the greatest impact on accelerator performance. However, this is particularly challenging because beam stability depends on the complex interactions of many magnetic field errors, resulting in a high-dimensional and nonlinear optimization problem. We determine which field errors are the most influential for the large physical aperture superconducting magnet B2PF, a critical magnet in the Interaction Region (IR) in the Hadron Storage Ring (HSR). We complete and analyze nearly 30,000 simulations on the Brookhaven National Laboratory Linux Cluster by varying 18 nonlinear magnetic field errors. We evaluate beam stability using the dynamic aperture and the tune diffusion. We identify the field errors that most strongly influence beam stability and establish quantitative field error tolerances that improve accelerator performance.

43 PARTICLE ACCELERATORS↗

OctoFAS: A Two-Level Fair Scheduler That Increases Fairness in Network-Based Key-Value Storage

We identified a fairness problem in a network-based key-value storage system using Intel Storage Performance Development Kit (SPDK) in a multitenant environment. In such an environment, each tenant’s I/O service rate is not fairly guaranteed compared to that of other tenants. To address the fairness problem, we propose OctoFAS, a two-level fair scheduler designed to improve overall throughput and fairness among tenants. The two-level scheduler of OctoFAS consists of (i) inter-core scheduling and (ii) intra-core scheduling. Through inter-core scheduling, OctoFAS addresses the load imbalance problem that is inherent in SPDK on the storage server by dynamically migrating I/O requests from overloaded cores to underloaded cores, thereby increasing overall throughput. Intra-core scheduling prioritizes handling requests from starving tenants over well-fed tenants within core-specific event queues to ensure fair I/O services among multiple tenants. OctoFAS is deployed on a Linux cluster with SPDK. Through extensive evaluations, we found that OctoFAS ensures that the total system throughput remains high and balanced, while enhancing fairness by approximately 10% compared to the baseline, when both scheduling levels operate in a hybrid fashion.

97 MATHEMATICS AND COMPUTING↗

OCTOKV: An Agile Network-Based Key-Value Storage System with Robust Load Orchestration

In this paper, we propose OctoKV, an innovative network-based key-value storage system. OctoKV addresses the repetitive address translation overhead associated with traditional key-value stores running on file systems on the client side. To mitigate this overhead, we implemented the key-value store on the server side using NVMe-oF and a user-level NVMe driver. In particular, we employed fine-grained resource monitoring and load balancing based on heuristics to optimize I/O performance. OctoKV is deployed on a Linux cluster with Intel SPDK. The extensive evaluation shows that OctoKV achieves lower I/O response times in comparison to traditional approaches where key-value stores run on the client side. Also, the proposed load balancing strategies efficiently enhance I/O response times by equally distributing the workload from overloaded cores to other cores.

Khan, Awais↗

Implementation of a practical Markov chain Monte Carlo sampling algorithm in PyBioNetFit

Abstract Summary Bayesian inference in biological modeling commonly relies on Markov chain Monte Carlo (MCMC) sampling of a multidimensional and non-Gaussian posterior distribution that is not analytically tractable. Here, we present the implementation of a practical MCMC method in the open-source software package PyBioNetFit (PyBNF), which is designed to support parameterization of mathematical models for biological systems. The new MCMC method, am, incorporates an adaptive move proposal distribution. For warm starts, sampling can be initiated at a specified location in parameter space and with a multivariate Gaussian proposal distribution defined initially by a specified covariance matrix. Multiple chains can be generated in parallel using a computer cluster. We demonstrate that am can be used to successfully solve real-world Bayesian inference problems, including forecasting of new Coronavirus Disease 2019 case detection with Bayesian quantification of forecast uncertainty. Availability and implementation PyBNF version 1.1.9, the first stable release with am, is available at PyPI and can be installed using the pip package-management system on platforms that have a working installation of Python 3. PyBNF relies on libRoadRunner and BioNetGen for simulations (e.g. numerical integration of ordinary differential equations defined in SBML or BNGL files) and Dask.Distributed for task scheduling on Linux computer clusters. The Python source code can be freely downloaded/cloned from GitHub and used and modified under terms of the BSD-3 license (https://github.com/lanl/pybnf). Online documentation covering installation/usage is available (https://pybnf.readthedocs.io/en/latest/). A tutorial video is available on YouTube (https://www.youtube.com/watch?v=2aRqpqFOiS4&t=63s). Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Calculated Gamma Output from a 6-kilogram Sphere of Neptunium

We previously modeled a 6 kg neptunium sphere with pyDMTK 2.0.0b, a python-based intrinsic radiation (INRAD) modeling tool, on the MOONLIGHT machine. Here we report results from version 2.0.1b on SNOW, another TriLab Linux Capacity Cluster (TLCC) resource on the Laboratory’s Turquoise network. We also present gamma output from MCNP 6.2.0 in terms of discrete line strengths, full gamma spectra, and dose rate maps for visualization. Results from both models agree with recent gamma measurements.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Decode the Workload: Training Deep Learning Models for Efficient Compute Cluster Representation

Monitoring the status of a high throughput computing cluster running computationally intensive production jobs is a crucial yet challenging system administration task due to the complexity of such systems. To this end, we train autoencoders using the Linux kernel CPU metrics of the cluster. Additionally, we explore assisting these models with graph neural networks to share information across threads within a compute node. The models are compared in terms of their ability to: 1) Produce a compressed latent representation that captures the salient features of the input, 2) Detect anomalous activity, and 3) Make distinction between different kinds of jobs run at Jefferson Lab. The goal is to have a robust encoder whose compressed embeddings are used for several downstream tasks. We extend this study further by deploying these models in a human-in-the-loop production-based setting for the anomaly detection task and discuss the associated implementation aspects such as continual learning and the criterion to generate alarms. This study represents a first step in the endeavor towards building self-supervised large-scale foundation models for computing centers.

Mohammed, Ahmed↗

COG User's Manual: A Multiparticle Monte Carlo Transport Code (Sixth Edition)

COG is a high-resolution code for the Monte Carlo simulation of coupled particle transport in arbitrary 3-D geometry. COG will transport neutrons, protons, deuterons, alpha particles with energies up to hundreds of GeV, and photons with energy ranges limited by the available cross section sets and physics models. Electrons can be transported via the EGS5 electron transport kernel, electrons can also be transported. The COG code is a significant upgrade from earlier Monte Carlo transport codes and has been written specifically to make it more versatile, accurate, and easy to use. COG has provisions for calculating deep penetration (shielding) problems, criticality problems, and neutron activation problems while retains all of the standard capabilities found in other Monte Carlo transport codes. COG uses high-resolution pointwise cross-section databases and makes no compromises in the transport physics, so that the results of a COG run are limited only by the accuracy of the databases used. COG runs primarily on Linux Operating System workstations with MPICH software installed – currently, Red Hat 7 & 8, Windows 10 (Windows Subsystem for Linux –WSL), Ubuntu 16, 18 & 20, OpenSUSE Leap 15.2, Fedora 32, Apple Power Mac with Intel CPU (with MacPorts installed) workstations, and LLNL LC supercomputer CTS-1 cluster with TOSS 3 are supported.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Poplar: a phylogenomics pipeline

Motivation Generating phylogenomic trees from the genomic data is essential in understanding biological systems. Each step of this complex process has received extensive attention and has been significantly streamlined over the years. Given the public availability of data, obtaining genomes for a wide selection of species is straightforward. However, analyzing that data to generate a phylogenomic tree is a multistep process with legitimate scientific and technical challenges, often requiring a significant input from a domain-area scientist. Results We present Poplar, a new, streamlined computational pipeline, to address the computational logistical issues that arise when constructing the phylogenomic trees. It provides a framework that runs state-of-the-art software for essential steps in the phylogenomic pipeline, beginning from a genome with or without an annotation, and resulting in a species tree. Running Poplar requires no external databases. In the execution, it enables parallelism for execution for clusters and cloud computing. The trees generated by Poplar match closely with state-of-the-art published trees. The usage and performance of Poplar is far simpler and quicker than manually running a phylogenomic pipeline. Availability and implementation Freely available on GitHub at https://github.com/sandialabs/poplar. Implemented using Python and supported on Linux.

Koning, Elizabeth [Sandia National Laboratories (S↗