Search NASA⌕ Search

SEARCH · Search NASA

Results for “HPC Access”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Jumping the Queue: From NASA to the Commercial Cloud

NASA's High-End Computing Capability (HECC) Project has made it possible for its users to run on commercial cloud resources in a seamless way. In the first of three phases, we implemented a pilot project for a few users, enabling them to “jump the queue” and burst jobs from the HECC environment to Amazon Web Services (AWS). By using GPU-accelerated nodes at AWS, the users were able to make significant advances in their research. The second phase of the project made AWS access available to all HECC users and added accounting to make users responsible for cloud charges. We are also enabling export-controlled work through the use of AWS GovCloud. In the third phase, we will add web-based mechanisms to permit non-HECC users to access cloud resources for their HPC projects.

Hood, Robert↗

Social Networking Adapted for Distributed Scientific Collaboration

Share is a social networking site with novel, specially designed feature sets to enable simultaneous remote collaboration and sharing of large data sets among scientists. The site will include not only the standard features found on popular consumer-oriented social networking sites such as Facebook and Myspace, but also a number of powerful tools to extend its functionality to a science collaboration site. A Virtual Observatory is a promising technology for making data accessible from various missions and instruments through a Web browser. Sci-Share augments services provided by Virtual Observatories by enabling distributed collaboration and sharing of downloaded and/or processed data among scientists. This will, in turn, increase science returns from NASA missions. Sci-Share also enables better utilization of NASA s high-performance computing resources by providing an easy and central mechanism to access and share large files on users space or those saved on mass storage. The most common means of remote scientific collaboration today remains the trio of e-mail for electronic communication, FTP for file sharing, and personalized Web sites for dissemination of papers and research results. Each of these tools has well-known limitations. Sci-Share transforms the social networking paradigm into a scientific collaboration environment by offering powerful tools for cooperative discourse and digital content sharing. Sci-Share differentiates itself by serving as an online repository for users digital content with the following unique features: a) Sharing of any file type, any size, from anywhere; b) Creation of projects and groups for controlled sharing; c) Module for sharing files on HPC (High Performance Computing) sites; d) Universal accessibility of staged files as embedded links on other sites (e.g. Facebook) and tools (e.g. e-mail); e) Drag-and-drop transfer of large files, replacing awkward e-mail attachments (and file size limitations); f) Enterprise-level data and messaging encryption; and g) Easy-to-use intuitive workflow.

Karimabadi, Homa↗

Using Big Data Technologies with Earth Science Data in HDF5: HDF5 Scalable Solutions

HDF5 (Hierarchical Data Format 5) is open-source, high-performance software that consists of an abstract data model, library, and fileformat used for storing and managing extremely large and/or complex data collections. NASA Earth Observing System (EOS) Data and Information Systems use HDF5 as an archival format to store remote sensing data from EOS satellites. HDF5 is also used to store other types of Geoscience and Strophysical data, e.g., seismic data and data from Low-Frequency Array (LOFAR) radio telescopes. Data stored in HDF5 has reached tens of petabytes and is growing at an accelerated rate.With the growing amout of HDF5 Earth Science data to analyze and process, scientists need to adopt big data technologies including new storage paradigms such as cloud and object storage. To run models and perform data analysis they also need to utilizied efficient and diverse ways to access data, from high-performance computing's (HPC) Message Passing Interface (MPI) I/O and deep memory hierarchies (DMH) to non-HPC frameworks such as Apache Hadoop, Spark, and Drill. The HDF Group continually works to enable usage of big data technologies in HDF software.

Knox, Larry↗

The Efficiency and the Scalability of an Explicit Operator on an IBM POWER4 System

We present an evaluation of the efficiency and the scalability of an explicit CFD operator on an IBM POWER4 system. The POWER4 architecture exhibits a common trend in HPC architectures: boosting CPU processing power by increasing the number of functional units, while hiding the latency of memory access by increasing the depth of the memory hierarchy. The overall machine performance depends on the ability of the caches-buses-fabric-memory to feed the functional units with the data to be processed. In this study we evaluate the efficiency and scalability of one explicit CFD operator on an IBM POWER4. This operator performs computations at the points of a Cartesian grid and involves a few dozen floating point numbers and on the order of 100 floating point operations per grid point. The computations in all grid points are independent. Specifically, we estimate the efficiency of the RHS operator (SP of NPB) on a single processor as the observed/peak performance ratio. Then we estimate the scalability of the operator on a single chip (2 CPUs), a single MCM (8 CPUs), 16 CPUs, and the whole machine (32 CPUs). Then we perform the same measurements for a chache-optimized version of the RHS operator. For our measurements we use the HPM (Hardware Performance Monitor) counters available on the POWER4. These counters allow us to analyze the obtained performance results.

Frumkin, Michael↗

Distributed Accounting on the Grid

By the late 1990s, the Internet was adequately equipped to move vast amounts of data between HPC (High Performance Computing) systems, and efforts were initiated to link together the national infrastructure of high performance computational and data storage resources together into a general computational utility 'grid', analogous to the national electrical power grid infrastructure. The purpose of the Computational grid is to provide dependable, consistent, pervasive, and inexpensive access to computational resources for the computing community in the form of a computing utility. This paper presents a fully distributed view of Grid usage accounting and a methodology for allocating Grid computational resources for use on a Grid computing system.

Thigpen, William↗

Electra: A Modular-Based Expansion of NASA's Supercomputing Capability

NASA has increasingly relied on high-performance computing (HPC) re- sources for computational modeling, simulation, and data analysis to meet the science and engineering goals of its missions in space exploration, aeronautics, and Earth and space science. The NASA Advanced Supercomputing (NAS) Division at Ames Research Center in Silicon Valley, Calif., hosts NASA’s premier supercomputing resources, integral to achieving and enhancing the success of the agency’s missions. NAS provides a balanced environment, funded under the High-End Computing Capability (HECC) project, comprised of world-class supercomputers, including its flagship distributed-memory cluster, Pleiades; high-speed networking; and massive data storage facilities, along with multi-disciplinary support teams for user support, code porting and optimization, and large-scale data analysis and scientific visualization. However, as scientists have increased the fidelity of their simulations and engineers are conducting larger parameter-space studies, the requirements for supercomputing resources have been growing by leaps and bounds. With the facility housing the HECC systems reaching its power and cooling capacity, NAS undertook a prototype project to investigate an alternative approach for housing supercomputers. Modular supercomputing, or container-based computing, is an innovative concept for expanding NASA’s HPC capabilities. With modular supercomputing, additional containers—similar to portable storage pods—can be connected together as needed to accommodate the agency’s ever-increasing demand for computing resources. In addition, taking advantage of the local weather permits the use of cooling technologies that would additionally save energy and reduce annual water usage. The first stage of NASA’s Modular Supercomputing Facility (MSF) prototype, which resulted in a 1,000 square-foot module on a concrete pad with room for 16 compute racks, was completed in Fall 2016 and an SGI (now HPE) computer system, named Electra, was deployed there in early 2017. Cooling is performed via an evaporative system built into the module, and preliminary experience shows a Power Usage Effectiveness (PUE) measurement of 1.03. Electra achieved over a petaflop on the LINPACK benchmark, sufficient to rank number 96 on the November 2016 TOP500 list [14]. The system consists of 1,152 InfiniBand-connected Intel Xeon Broadwell-based nodes. Its users access their files on a facility-wide file system shared by all HECC compute assets via Mellanox MetroX InfiniBand extenders, which connect the Electra fabric to Lustre routers in the primary facility over fiber-optic links about 900 feet long. The MSF prototype has exceeded expectations and is serving as a blueprint for future expansions. In the remainder of this chapter, we detail how modular data center technology can be used to expand an existing compute resource. We begin by describing NASA’s requirements for supercomputing and how resources were provided prior to the integration of the Electra module-based system.

Biswas, Rupak↗

Performance and Portability of a Linear Solver Across Emerging Architectures

A linear solver algorithm used by a large-scale unstructured-grid computational fluid dynamics application is examined for a broad range of familiar and emerging architectures. Efficient implementation of a linear solver is challenging on recent CPUs offering vector architectures. Vector loads and stores are essential to effectively utilize available memory bandwidth on CPUs, and maintaining performance across different CPUs can be difficult in the face of varying vector lengths offered by each. A similar challenge occurs on GPU architectures, where it is essential to have coalesced memory accesses to utilize memory bandwidth effectively. In this work, we demonstrate that restructuring a computation, and possibly data layout, with regard to architecture is essential to achieve optimal performance by establishing a performance benchmark for each target architecture in a low level language such as vector intrinsics or CUDA. In doing so, we demonstrate how a linear solver kernel can be mapped to Intel® Xeon™ and Xeon Phi™, Marvell® ThunderX2®, NEC® SX-Aurora™ TSUBASA Vector Engine, and NVIDIA® and AMD® GPUs. We further demonstrate that the required code restructuring can be achieved in higher level programming environments such as OpenACC, OCCA, and Intel® OneAPI™/SYCL, and that each generally results in optimal performance on the target architecture. Relative performance metrics for all implementations are shown, and subjective ratings for ease of implementation and optimization are suggested.

Programming models↗

NASA Earth Exchange – Current overview of climate and wildfire-oriented works

NASA Earth Exchange (NEX) combines state-of-the-art supercomputing, Earth system modeling, and NASA remote sensing data feeds to deliver a work environment for exploring and analyzing petabyte-scale datasets covering large regions, continents, or the globe. As an accessible platform, NEX can accelerate fundamental research, develop new applications, and reduce overall project costs by providing the research community with data, software, and high-end computing power. Two research thrusts for the NEX community are developing and distributing the NEX-GDDP-CMIP6 downscaled climate dataset and the community-driven wildfire research. NEX-GDDP-CMIP6 dataset is comprised of global downscaled climate scenarios derived from the General Circulation Model (GCM) runs conducted under the Coupled Model Intercomparison Project Phase 6 (CMIP6) and across two of the four “Tier 1” greenhouse gas emissions scenarios known as Shared Socioeconomic Pathways (SSPs). The wildfire research thrust leverages geostationary and low earth orbit remote sensing platforms and advanced modeling (WRF). We present a high-level overview of the NEX community, discuss relevant datasets and show several visualization techniques of data used in the wildfire research activity.

NEX↗

Cloud Computing Methods for Near Rectilinear Halo Orbit Trajectory Design

Complicated mission design problems require innovative computational solutions. As spacecraft depart from a proposed Gateway in a Near Rectilinear Halo Orbit (NRHO), recontact analysis is required to avoid risk of collision and ensure safe operations. Escape dynamics from NRHOs are governed by multiple gravitational bodies, yielding a trajectory design space that is exhaustively large. This paper summarizes the recontact analysis for departure from the NRHO and describes how the Deep Space Trajectory Explorer (DSTE) trajectory design software incorporates high performance cloud computing to compute and visualize the orbit design space. Recent focus on exploration missions to cislunar space has kindled accelerated interest in multibody orbit solutions. Trajectory analysis in the presence of multiple gravity fields is complex, and innovative computational tools are needed to simplify complicated design spaces, to generate large quantities of data quickly, and to visualize the output for user accessibility. The Gateway mission is a prime example. The Gateway1 is proposed as a human outpost in deep space. The current baseline orbit for the Gateway is a Near Rectilinear Halo Orbit (NRHO) near the Moon.2 The NRHO exists in a regime that experiences the gravitational effects of the Earth and the Moon simultaneously, complicating orbit analysis. The mission design process benefits greatly from updated computational tools for multibody missions like the Gateway. As an example, consider the problem of assessing the risk of collision in an NRHO. As a staging location to missions to the lunar surface and beyond the Earth-Moon system, the Gateway will experience spacecraft and other objects regularly arriving and departing. Departing objects potentially include spent logistics modules, visiting crew vehicles, debris objects, wastewater particles, and cubesats. Each departure is governed by the dynamics of the Gateway orbit and the surrounding dynamical environment. Over time, any unmaintained object in such an orbit eventually departs due to the small instabilities associated with the NRHOs. A separation maneuver speeds the departure from the NRHO, but the effects of the maneuver on the spacecraft behavior depend on the location, magnitude, and direction of the burn. Escape dynamics from the NRHO with regard to these maneuver options open up an enormous potential trajectory design space where subtle changes in input can produce dramatically large changes in the results. Any departing object must avoid recontacting the Gateway as it leaves the lunar vicinity, and a recontact analysis thus involves a significant number of computations and extensive output data. To explore the dynamics of this extensive design space, the Deep Space Trajectory Explorer3 (DSTE) trajectory design software incorporates new High Performance Computing (HPC) services and novel interactive visualizations. This paper details the HPC and cloud infrastructure techniques that are implemented in the DSTE, applying the new capabilities to analysis of recontact risk with the Gateway in NRHO. NEAR RECTILINEAR HALO ORBITS The Gateway is planned to fly in a lunar NRHO as its baseline orbit. The NRHO families of orbits are subsets of the larger halo families, which originate from planar orbits near the L1 and L2 libration points; the Earth-Moon L2 halo family appears in Figure 1. Each halo orbit is perfectly periodic in the Circular Restricted 3-Body Problem (CR3BP) and becomes a quasi-periodic orbit in a higher fidelity ephemeris force model. The NRHOs are defined as those members of the halo family with bounded stability properties;2 they pass near the Moon at perilune and are nearly polar. Families exist with apolunes located both above the lunar north pole and above the lunar south pole; the Gateway is planned to reside in a southern L2 NRHO in a 9:2 resonance with the lunar synodic period. The 9:2 NRHO is characterized by a period of about 6.5 days, a perilune radius of about 3,500 km, and an apolune radius of about 71,000 km; it is strongly affected by the gravity of both the Earth and the Moon simultaneously. This NRHO offers extended communications with assets on the south pole of the Moon,4 as well as low-cost orbit maintenance and attitude control,5 favorable eclipse avoidance properties,6 and inexpensive transfers from Earth and to other destinations.5,7 The NRHO portion of the southern L2 halo family is highlighted in black in Figure 1, and the 9:2 NRHO appears in blue.

Phillips, Sean M.↗