Search NASASearch

SEARCH · Search NASA

Results for “graphics processing units”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Leveraging the Usage of GPUs in SAR Processing for the NISAR Mission

The NASA ISRO Synthetic Aperture Radar (NISAR) mission will redefine the future of earth science in terms of both the quality as well as the quantity of data that will be downlinked daily. The current software architecture used to process this data is the InSAR Scientific Computing Environment (ISCE), a powerful and modular platform that applies a combination of novel and legacy processing modules to many sources of SAR data. Until recently, this architecture could process most images in a reasonable amount of time; however in the case of the NISAR mission (where the daily influx as well as the size of the images themselves are significantly larger) the current architecture can take hours to process even a single image. This paper explores new efforts to use a Graphics Processing Unit (GPU) to accelerate one of the processing modules to achieve unprecedented runtimes with no loss in precision, potentially setting a new standard in radar processing in the world of “Big Data”.

Cohen, Joshua

Standardizing Microprocessor and GPU Radiation Test Approaches

Microprocessor, Graphics Processing Units (GPUs) and DDRx memory devices have emerged as promising next-generation technologies that enables both high performance processing and acceleration of complex algorithms for the latest challenges in human spaceflight, autonomous vehicles and artificial intelligence (AI). The feature sets of these devices offer exponential increases to throughput, calculation capability and system autonomy when compared to legacy flight systems. NASA's Electronic Part and Packaging (NEPP) Program has conducted an investigation into the radiation susceptibility of leading edge devices and process technologies by establishing standardized test approaches. Unlike most discrete devices, these require state of the art test systems to induce specific hardware activity similar to application software, thus allowing the characterization of failure modes within the system. To best characterize the tested part, NEPP eliminates variables that may impact device performance under radiation. Simplification of remaining system-level variables leads to an improved understanding of complex computational devices and their intended applications. The failure modes and error signatures that are recorded during testing are used to determine radiation sensitivity of the semiconductor process and the microcode architecture of the design. This presentation will discuss the test methodology that NASA Electronic Parts and Packaging (NEPP) is working to establish for its microprocessor, GPU and DDRx memory test programs to provide guidance on these devices and their underlying technology, in regards to their potential usage in future space flight systems.

GPU

TES-8: Advanced Exo-Brake, VR and COM Experiments

The TES-8 was jettisoned from the International Space Station on January 31, 2019. As an orbital laboratory and 8th in on-going series, the design makes use of a standard set of interfaces and safety features that permit rapid re-flight. On this flight, an advanced Exo-Brake is flown with de-orbit targeting capability that will engender sample return capability from LEO platforms. A Virtual Reality data recording system uses stereo imaging and efficient data-compression with an NVIDIA GPU (Graphics Processing Unit) to permit compression and transmission of very large data files. An SDR (Software Defined Radio) will download data to the NEN (Near Earth Network) for the first time - demonstrating potential use in cis-lunar space using S-band. For the first time, a comparison will be made regarding the functionality of the Iridium and Globalstar short burst data modems - as essential communication tools for future nano-sat projects. Lastly, the 7 micro-processors and 4 cameras provide an excellent learning platform for university students and NASA young professionals.

Exo-Brake

Advanced Astrophysics Discovery Technology in the Era of Data Driven Astronomy

Astrophysics is at the threshold of a new epoch in which increasinglycomplex, heterogeneous datasets will challenge our existing information infrastructure and traditional approaches to analysis. The rapid advancement of graphics processing units, compact field programmable gate arrays and dedicated artificial intelligence accelerator chips is now permitting the use of scientific methods, processes and algorithms to extract knowledge and insights from structured and unstructured data in ways never before seen. Miniaturization of spacecraft architectures and supporting infrastructure is opening new observing strategies and new discovery spaces for science. The community is just beginning to awaken to these imminent challenges as evidenced by their relative lack of emphasis in the New Worlds, New Horizons ASTRO2010 decadal survey, in the ExoPAG Science Analysis Group 11 report andin the formulation of the WFIRST Data Challenge. We suggest that the Astrophysics Science Division (ASD), which has clearly recognized this new epoch of rapidly evolving information technology, could be more affirmative in its approach. We offer a modest structural solution.

Barry, Richard K.

Parallelization of a Six Degree of Freedom Entry Vehicle Trajectory Simulation Using OpenMP and OpenACC

The art and science of writing parallelized software, using methods such as Open Multi-Processing (OpenMP) and Open Accelerators (OpenACC), is dominated by computer scientists. Engineers and non-computer scientists looking to apply these techniques to their project applications face a steep learning curve, especially when looking to adapt their original single threaded software to run multi-threaded on graphics processing units (GPUs). There are significant changes in mindset that must occur; such as how to manage memory, the organization of instructions, and the use of if statements (also known as branching). The purpose of this work is twofold: 1) to demonstrate the applicability of parallelized coding methodologies, OpenMP and OpenACC, to tasks outside of the typical large scale matrix mathematics; and 2) to discuss, from an engineer’s perspective, the lessons learned from parallelizing software using these computer science techniques. This work applies OpenMP, on both multi-core central processing units (CPUs) and Intel® Xeon Phi™ 7210, and OpenACC on GPUs. These parallelization techniques are used to tackle the simulation of thousands of entry vehicle trajectories through the integration of six degree of freedom (DoF) equations of motion (EoM). The forces and moments acting on the entry vehicle, and used by the EoM, are estimated using multiple models of varying levels of complexity. Several benchmark comparisons are made on the execution of six DoF trajectory simulation: single thread Intel® Xeon® E5-2670 CPU, multi-thread CPU using OpenMP, multi-thread Xeon Phi™ 7210 using OpenMP, and multi-thread NVIDIA® Tesla® K40 GPU using OpenACC. These benchmarks are run on the Pleiades Supercomputer Cluster at the National Aeronautics and Space Administration (NASA) Ames Research Center (ARC), and a Xeon Phi™ 7210 node at NASA Langley Research Center (LaRC).

Green, Justin S.

Simulations of a Turbulent Flow Subjected to Favorable and Adverse Pressure Gradients

This paper reports the results from a direct numerical simulation of an initially turbulent boundary layer passing over a wall-mounted “speed bump” geometry. The speed bump, represented in the form of a Gaussian distribution profile, generates a favorable pressure gradient region over the upstream half of the geometry, followed by an adverse pressure gradient over the downstream half. The boundary layer approaching the bump undergoes strong acceleration in the favorable pressure gradient region before experiencing incipient or very weak separation within the adverse pressure gradient region. These types of flows have proven to be particularly challenging to predict using lower-fidelity simulation tools based on various turbulence modeling approaches and warrant the use of the highest-fidelity simulation techniques. Simulation results are utilized to examine the key phenomena present in the flowfield, such as relaminarization/stabilization in the strong acceleration region succeeded by retransition to turbulence near the onset of adverse pressure gradient, incipient/weak separation, and development of internal layers where the sense of streamwise pressure gradient changes at the foot, apex and tail of the bump. The present direct numerical simulation is performed using a flow solver developed exclusively for graphics processing units, which is found to provide a significant speedup compared to an earlier solver optimized for central processing unit architectures.

Ali Uzun

Recent Improvements to the LAURA and HARA Codes

This paper describes recent improvements to the LAURA and HARA codes. LAURA is a CFD code for aerothermodynamics, and HARA evaluates the shock-layer radiation that provides the radiative source term for the flowfield energy equations and radiative heating to a surface. The next release of LAURA and HARA includes a variety of new capabilities. These new capabilities include an automated uncertainty quantification workflow for radiative heat transfer, options for specifying surface roughness and turbulent transition location in the algebraic turbulence models, and improved grid and solution interpolation techniques. Additionally, the computational efficiency of both LAURA and HARA have been improved. Optimization of the MPI communication routines in LAURA are shown to improve the parallel efficiency of the primary flow when running with multiple processes per block, and recent optimization of HARA leverage graphics processing unit (GPU) acceleration in the radiation calculations. Using GPU acceleration of HARA is shown to decrease the cost of the radiation line-of-sight calculation by approximately one order of magnitude for a 10.5 km/s Earth entry simulation.

LAURA HARA CFD 5.6

NASA GPU Hackathon Yields Significant Code Improvements

The NASA GPU Hackathon 2020 brought together application developers and computer experts to help get important NASA applications running effectively on graphics processing unit (GPU) nodes. Nine teams of application developers participated in this virtual event, a major impetus for teams to modernize codes of interest for NASA missions to CPU nodes containing GPU accelerators, with a focus on hands-on problem solving. The photo in Figure1 shows 30 of the more than50 participants. The HECC project and NVIDIA jointly organized the event, and HECC provided five Pleiades nodes each with 4 V100 GPUs for teams to use. The virtual event, which took place over four days from September 28–October 7, 2020, used Microsoft Teams and Slack as collaboration tools. Each team consisted of three to six members from NASA Centers and supporting organizations. The teams were paired with one to two mentors from industry, government, and academia. The experience levels of the teams ranged from being GPU novices to advanced CUDA programming experts. OpenACC and the emerging Kokkos API were used in addition to CUDA for GPU programming. During the event, which focused on accelerating AeroSciences and CFD applications, most teams achieved considerable performance improvements on both GPUs and CPUs. For example, a team with no GPU experience completed a first port of a time-critical loop to a GPU. Another team of expert CUDA programmers were able to restructure their algorithm, yielding a factor of five speed-up. And another team sped up some of their CUDA kernels by a factor of 20, which directly translated into their production code. This article highlights some of the many successes resulting from the event.

HECC

Synthetic Tracking on a Small Telescope

Synthetic tracking uses high speed (up to 10 Hz) low noise (<2e-) large format sensors ~16 Mpix along with a multi-vector shift/add algorithm that coadds multiple image frames to increase the signal to noise ratio (SNR) needed to detect (if present) multiple moving objects in the field of view (FOV). We published the application of synthetic tracking to look for asteroids in 2014 (Shao 2014), but recently have applied it more as well to Earth orbiting objects. We have begun testing the data processing graphical processing unit (GPU) array with a small telescope, a 28 cm Celestron RASA telescope and a low cost low noise 16 Mpix CMOS camera at a dark site in California. This system is now operational with a 2 sqdeg FOV and a limiting magnitude between ~16-17.5 stellar magnitudes (mag) depending on a number of observational parameters for short integration times. The instrument can be used to search for NEOs, where we use much longer integration times to get sensitivity ~ 20.5 mag (at new moon). Synthetic tracking provides significant improvements in both sensitivity and astrometric accuracy.

Turyshev, Slava G.

GPU Supported Simulation of Transition-edge Sensor Arrays

We present numerical simulations of full transition-edge sensor (TES) arrays utilizing graphical processing units (GPUs). With the support of GPUs, it is possible to perform simulations of large pixel arrays to assist detector development. Comparisons with TES small-signal and noise theory confirm the representativity of the simulated data. In order to demonstrate the capabilities of this approach, we present its implementation in xifusim, a simulator for the X-ray Integral Field Unit, a cryogenic X-ray spectrometer on board the future Athena X-ray observatory.

M Lorenz

Radiation specification and testing of heterogenous microprocessor SOCs

Modern commercial microprocessor devices include multiple processor architectures, buses, basic peripherals, and application hardware such as Graphics Processing Units (GPUs) and Digital Signal Processors (DSPs) in one device. Developing RHBD versions of similar devices risks sacrificing processing performance for system-wide radiation requirements. The heterogenous structure of modern commercial system on a chip (SOC) devices, in design and performance goals for subsystems, suggests a similar approach to specifying Radiation Hardened by Design (RHBD) requirements.

Ballast, Jon

Memory Optimizations for Sparse Linear Algebra on GPU Hardware

An effort to maximize memory bandwidth utilization for a sparse linear algebra kernel executing on NVIDIA® Tesla V100 and A100 Graphics Processing Units (GPUs) is described. The kernel consists of a block-sparse matrix-vector product and a series of forward/backward triangular solves. The computation is memory-bound and exhibits low arithmetic intensity. Along with a relatively small block size, the data layout poses a challenge to effectively utilize the available memory bandwidth on common GPU architectures. An earlier implementation using a warp to process a single row of the matrix was found to yield good memory performance on the V100 architecture. However, anew approach, which assigns a warp to six rows of the matrix, is proposed for the A100. In addition, two new features offered by the A100 architecture are explored.L2residency control enables a portion of theL2cache to be used for persistent data access, and the asynchronous copy instruction allows data to be loaded directly from main memory into shared memory. Demonstrations show that the new implementation improves memory bandwidth utilization from 71.5% to 81.2% of the peak available on theA100 architecture.

GPU

Application of a Detached Eddy Simulation Approach with Finite-Rate Chemistry to Mars-Relevant Retropropulsion Operating Environments

Human-scale Mars vehicles will require retropropulsion for descent and landing, replacing heritage supersonic parachute systems with an extended phase of powered flight. Due to the limitations of terrestrial testing in Mars-relevant conditions, design and analysis will increasingly rely on computational modeling and simulation. This paper provides an overview of a computational campaign investigating the aerodynamics of a Mars lander concept along various points on a powered descent trajectory including supersonic, transonic, and subsonic conditions using finite-rate chemistry. Simulations using unstructured grids containing billions of elements are performed at scale using thousands of Graphics Processing Units, enabling run-times of a few days for each simulation presented. At each freestream condition, significant minor species concentrations are observed external to the nozzles in the large mixing region upstream of the vehicle. While the flowfields are highly non-stationary in all cases, the mean integrated forces and moments on the vehicle remain small in comparison to the deceleration provided through retropropulsion.

Ashley Korzun

Computational Investigation of the Effect of Chemistry on Mars Supersonic Retropropulsion Environments

Retropropulsion ground tests require significant compromises on physical scale, instrumentation, configuration, and environments. Matching the full Martian environment is simply not possible on Earth. Ground tests of retropropulsion configurations thus far have neglected effects of chemistry due to physical constraints of wind tunnel models and facilities; most experiments use inert simulant gases at relatively low temperatures. As such, a strong reliance on high-fidelity computational analyses is required to expand the knowledge of retropropulsion aerodynamics. In this work, we investigate the effects of chemistry using scale-resolving computational fluid dynamics (CFD) with finite-rate chemistry and a graphics processing unit (GPU)-enabled implementation of the NASA FUN3D flow solver on a human-scale Mars lander concept at supersonic freestream conditions. Results are compared to previous perfect gas simulations.

retropropulsion

Computational Investigation of the Effect of Chemistry on Mars Supersonic Retropropulsion Environments

Retropropulsion ground tests require significant compromises on physical scale, instrumentation, configuration, and environments. Matching the full Martian environment is simply not possible on Earth. Ground tests of retropropulsion configurations thus far have neglected effects of chemistry due to physical constraints of wind tunnel models and facilities; most experiments use inert simulant gases at relatively low temperatures. As such, a strong reliance on high-fidelity computational analyses is required to expand the knowledge of retropropulsion aerodynamics. In this work, we investigate the effects of chemistry using scale-resolving computational fluid dynamics (CFD) with finite-rate chemistry and a graphics processing unit (GPU)-enabled implementation of the NASA FUN3D flow solver on a human-scale Mars lander concept at supersonic freestream conditions. Results are compared to previous perfect gas simulations.

retropropulsion

Computational Investigation of the Effect of Chemistry on Mars Retropropulsion Environments using a Massively Parallel GPU Approach

In this work, we investigate the effects of chemistry on a human-scale Mars lander concept using scale-resolving computational fluid dynamics (CFD) with finite-rate chemistry and a graphics processing unit (GPU)-enabled implementation of the NASA FUN3D flow solver, enabling run-times of a few days for the simulations presented. Simulations are carried out on Summit at Oak Ridge Leadership Computing Facility using thousands of GPUs. Retropropulsion ground tests require significant compromises on physical scale, instrumentation, configuration, and environments. Ground tests of retropropulsion configurations thus far have neglected effects of chemistry due to physical constraints of wind tunnel models and facilities; most experiments use inert simulant gases at low temperatures. As such, a strong reliance on high-fidelity computational analyses such as those presented in this work is required to expand the knowledge of retropropulsion aerodynamics. An overview of the GPU approach will be presented. Results are compared to a previous scaled perfect gas (air) campaign.

retropropulsion

A Multi-Architecture Approach for Implicit Computational Fluid Dynamics on Unstructured Grids

High-performance computing (HPC) architectures are trending toward manycore paradigms such as graphics processing units (GPUs). Approximately half of the top 100 publicly disclosed supercomputers in the world utilize GPU accelerators for performance. This is in contrast to a decade ago, where there were only a few such machines in the top 100. It is not currently possible to compile and run legacy central processing unit (CPU) software efficiently on GPUs without significant refactoring. Though a number of frameworks offering performance portability exist, none offer a standardized specification that is supported by all major hardware vendors. Additionally, experiences show that obtaining a high percentage of peak performance often requires architecture-specific code. This work details a pragmatic multi-architecture computational fluid dynamics library focused on aerospace problems across the speed range from low subsonic to hypersonic flows involving thermochemical nonequilibrium. A thin abstraction layer above NVIDIA CUDA C++ is utilized, which enables primarily single-source software currently capable of running efficiently on multicore CPUs, NVIDIA GPUs, AMD GPUs, and Intel GPUs. Results on various problems of interest across the speed range are presented and performance is compared between various architectures.

GPU

A Multi-Architecture Approach for Implicit Computational Fluid Dynamics on Unstructured Grids

High-performance computing (HPC) architectures are trending toward manycore paradigms such as graphics processing units (GPUs). Approximately half of the top 100 publicly disclosed supercomputers in the world utilize GPU accelerators for performance. This is in contrast to a decade ago, where there were only a few such machines in the top 100. It is not currently possible to compile and run legacy central processing unit (CPU) software efficiently on GPUs without significant refactoring. Though a number of frameworks offering performance portability exist, none offer a standardized specification that is supported by all major hardware vendors. Additionally, experiences show that obtaining a high percentage of peak performance often requires architecture-specific code. This work details a pragmatic multi-architecture computational fluid dynamics library focused on aerospace problems across the speed range from low subsonic to hypersonic flows involving thermochemical nonequilibrium. A thin abstraction layer above NVIDIA CUDA C++ is utilized, which enables primarily single-source software currently capable of running efficiently on multicore CPUs, NVIDIA GPUs, AMD GPUs, and Intel GPUs. Results on various problems of interest across the speed range are presented and performance is compared between various architectures.

GPU