Search NASASearch

SEARCH · Search NASA

Results for “GPU computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Standardizing Microprocessor and GPU Radiation Test Approaches

Microprocessor, Graphics Processing Units (GPUs) and DDRx memory devices have emerged as promising next-generation technologies that enables both high performance processing and acceleration of complex algorithms for the latest challenges in human spaceflight, autonomous vehicles and artificial intelligence (AI). The feature sets of these devices offer exponential increases to throughput, calculation capability and system autonomy when compared to legacy flight systems. NASA's Electronic Part and Packaging (NEPP) Program has conducted an investigation into the radiation susceptibility of leading edge devices and process technologies by establishing standardized test approaches. Unlike most discrete devices, these require state of the art test systems to induce specific hardware activity similar to application software, thus allowing the characterization of failure modes within the system. To best characterize the tested part, NEPP eliminates variables that may impact device performance under radiation. Simplification of remaining system-level variables leads to an improved understanding of complex computational devices and their intended applications. The failure modes and error signatures that are recorded during testing are used to determine radiation sensitivity of the semiconductor process and the microcode architecture of the design. This presentation will discuss the test methodology that NASA Electronic Parts and Packaging (NEPP) is working to establish for its microprocessor, GPU and DDRx memory test programs to provide guidance on these devices and their underlying technology, in regards to their potential usage in future space flight systems.

GPU

New Features of the NEQAIR Radiation Code

The longest-lived code for predicting shock layer radiation, NEQAIR, is now in its 5th decade of service. Substantial changes to the code have been made over the previous decade, the most recent report of which was at the 5th Workshop on Radiation in High Temperature Gases in 2014, for the version referred to as NEQAIR14. This paper will review some of the improvements made to the NEQAIR code since then, which is now at v15.2. Some of these features are discussed briefly below. NEQAIR15 and subsequent versions have enabled parallel evaluation of multiple lines of sight. This is accomplished by utilizing the HDF5 file format and placing multiple lines into a single file, LOS.h5, which is used for both input and output. This approach enables straightforward parallel execution both over the number of lines of sight and the number of points per line. For large problems, runtime reduces linearly with the number of nodes deployed since each line is processed independently by a subset of MPI ranks. Three applications of the multi-line solver are discussed. The first has to do with performing loosely coupled radiation-flowfield solutions. In this case the computed absorption and emission coefficients are used to evaluate the total energy absorbed or emitted at each point, allowing evaluation of the volumetric source term in the flowfield. The second computation is for obtaining heat flux from nonuniform flows, which require integration over spherical co-ordinates. These are of particular interest for evaluating radiation on the vehicle backshell. This 3D option improves the angular integration scheme and allows adaptive line selection that together reduce the number of lines required by about an order of magnitude. The final application is for remote observation, which is essentially the 3D integration problem over a small solid angle. For all three of these computations, data can be stored in the HDF5 file which allows a NEQAIR run to be restarted when it times out, or to add atmospheric absorption or instrument scan functions. An additional level of parallelism is enabled in NEQAIR15.2 using GPU routines. The GPU parallelism has realized up to 8x speed-up when running on a single core but diminishes as CPU parallelism is increased. For running multi-line simulations, it may be easier to reserve a large number of CPU nodes than to obtain the number of GPU nodes required for similar performance. A GUI, known as NEQTPY, allows for reading and creating input files, running NEQAIR, and displaying results. A significant feature of NEQTPY is the ability to perform spectral fits to data. The fits can operate on a single line spectrum (radiance vs. wavelength) or a 3D input file with multiple columns of data. Other new features include improved constants, additional species, more detailed non-Boltzmann modelling, advanced user controls, the ability to read and calculate spectra from HITRAN datafiles, photodissociation and photoionization cross-sections. A “fast” automatic grid option may reduce the size and time of spectral calculations while still maintaining good accuracy for total heat flux.

Brett A Cruden

Recent Improvements to the LAURA and HARA Codes

This paper describes recent improvements to the LAURA and HARA codes. LAURA is a CFD code for aerothermodynamics, and HARA evaluates the shock-layer radiation that provides the radiative source term for the flowfield energy equations and radiative heating to a surface. The next release of LAURA and HARA includes a variety of new capabilities. These new capabilities include an automated uncertainty quantification workflow for radiative heat transfer, options for specifying surface roughness and turbulent transition location in the algebraic turbulence models, and improved grid and solution interpolation techniques. Additionally, the computational efficiency of both LAURA and HARA have been improved. Optimization of the MPI communication routines in LAURA are shown to improve the parallel efficiency of the primary flow when running with multiple processes per block, and recent optimization of HARA leverage graphics processing unit (GPU) acceleration in the radiation calculations. Using GPU acceleration of HARA is shown to decrease the cost of the radiation line-of-sight calculation by approximately one order of magnitude for a 10.5 km/s Earth entry simulation.

LAURA HARA CFD 5.6

FUN3D Manual: 14.0.2

This manual describes the installation and execution of FUN3D version 14.0.2, including optional dependent packages. FUN3D is a suite of computational fluid dynamics simulation and design tools that uses mixed-element unstructured grids in a large number of formats, including structured multiblock and overset grid systems. A discretely-exact adjoint solver may be used for formal design optimization, error estimation, and mesh adaptation. FUN3D also offers a reacting, real-gas capability and provides GPU acceleration of many common simulation options.

William K Anderson

FUN3D Manual: 14.1

This manual describes the installation and execution of FUN3D version 14.1, including optional dependent packages. FUN3D is a suite of computational fluid dynamics simulation and design tools that uses mixed-element unstructured grids in a large number of formats, including structured multiblock and overset grid systems. A discretely-exact adjoint solver may be used for formal design optimization, error estimation, and mesh adaptation. FUN3D also offers a reacting, real-gas capability and provides GPU acceleration of many common simulation options.

William K Anderson

Jumping the Queue: From NASA to the Commercial Cloud

NASA's High-End Computing Capability (HECC) Project has made it possible for its users to run on commercial cloud resources in a seamless way. In the first of three phases, we implemented a pilot project for a few users, enabling them to “jump the queue” and burst jobs from the HECC environment to Amazon Web Services (AWS). By using GPU-accelerated nodes at AWS, the users were able to make significant advances in their research. The second phase of the project made AWS access available to all HECC users and added accounting to make users responsible for cloud charges. We are also enabling export-controlled work through the use of AWS GovCloud. In the third phase, we will add web-based mechanisms to permit non-HECC users to access cloud resources for their HPC projects.

Hood, Robert

FUN3D Manual: 14.0

This manual describes the installation and execution of FUN3D version 14.0,including optional dependent packages. FUN3D is a suite of computational fluid dynamics simulation and design tools that uses mixed-element unstructured grids in a large number of formats, including structured multiblock and overset grid systems. A discretely-exact adjoint solver may be used for for-mal design optimization, error estimation, and mesh adaptation. FUN3D also offers a reacting, real-gas capability and provides GPU acceleration of many common simulation options.1

William K Anderson

"Sensor Web Evolution - Webs of Webs for NASA Science - Focus on small Uninhabited Aerial Systems (sUAS)"

This paper will describe the evolution of information collection, derivation and delivery mechanisms in webs of NASA sensor webs, with a focus on recent advancements in small Uninhabited Aerial Systems (sUAS). I will discuss the movement to "Fog Computing", also known as Edge Computing. Fog Computing facilitates the distribution of common operations and networking between edge devices and cloud computing facilities, optimizing the production of actionable intelligence. Initially, sUASs utilized onboard data collection as standard, with minimal data downloaded directly. Information products were derived in conventional computational environments, generally desk top computers, and information products made available to the Science Community in weeks or months. With the increased availability, and increasingly lower costs, of beyond line of sight (BLOS) satellite based communication, transmission rates and data volumes increased, and processing migrated to Cloud based services. Contemporary sUASs are moving some of that information product derivation to on vehicle services, and are creating a distributed Cloud/Fog environment. I will describe the technological advances that have made this possible, including low power multi-core Central Processing Units (CPU), and, more recently, the availability of high end Graphical Processing Units (GPU) that consume only a few watts. Intelligent system software, leveraging these hardware advances, finally allows for information product generation on-board, rather than simple data collection. Additionally, intelligent flight control systems now support mutual vehicle to vehicle collaboration, allowing sUASs to create ad-hoc sensor webs on demand, as required. Also discussed will be the lessons learned by the Authors' development of data systems for NASA's large High Altitude Long Endurance (HALE) UASs like Predator and Global Hawk, and how those lessons are being applied to sUAS development. This paper will focus on application, rather a deep dive into the technology, and will highlight improving data management through these new technologies.

Sensor Web

Short–Period Variables in TESS Full–Frame Image Light Curves Identified via Convolutional Neural Networks

The Transiting Exoplanet Survey Satellite (TESS) mission measured light from stars in ∼85% of the sky throughout its 2 yr primary mission, resulting in millions of TESS 30-minute-cadence light curves to analyze in the search for transiting exoplanets. To search this vast data set, we aim to provide an approach that is computationally efficient, produces accurate predictions, and minimizes the required human search effort. We present a convolutional neural network that we train to identify short-period variables. To make a prediction for a given light curve, our network requires no prior target parameters identified using other methods. Our network performs inference on a TESS 30-minute-cadence light curve in ∼5 ms on a single GPU, enabling large-scale archival searches. We present a collection of 14,156 short-period variables identified by our network. The majority of our identified variables fall into two prominent populations, one of close-orbit main-sequence binaries and another of δ Scuti stars. Our neural network model and related code are additionally provided as open-source code for public use and extension.

Convolutional neural networks

Real-Time Background Oriented Schlieren: Catching Up With Knife Edge Schlieren

Background Oriented Schlieren (BOS) is a widely used technique that provides density gradient information in flow fields of interest, without imposing stringent optical quality requirements on the facility/experiment windows and/or optics used in the BOS setup. Typically, the BOS reference image is acquired before the test begins (flow off) and then the "live" image data are acquired during the actual testing/experiment (flow on). The raw BOS image data, while displayed in real-time as they are acquired from the camera, unfortunately provide little if any visual indication of the density gradients in the flow. Generally, the "live" images must be processed off-line after the testing is completed, providing no indication of the success of the BOS setup and no feedback on the operational success of the test. Advances in computer processing hardware enables the implementation of real-time processing and display of the BOS image data. Two different approaches to implementing the real-time BOS (RT-BOS) processing capability are described herein. First, a traditional multi-core Central Processing Unit (CPU) based approach using scheduled parallel threads is used to build a RT-BOS processing engine. In the second approach, a Graphical Processing Unit (GPU) approach is used to costruct a RT-BOS processing engine. Generally, high core count CPU processors can provide a useful processing rate for RT-BOS. However, the GPU based approach exceeds the processing capability of the CPU approach, at a fraction of the cost. The GPU approach places no restrictions on the Host PC processing capability, except that it be capable of acquiring the BOS image data from the camera in real-time.

Wernet, Mark P.

Identifying Planetary Transit Candidates in TESS Full-frame Image Light Curves via Convolutional Neural Networks

The Transiting Exoplanet Survey Satellite(TESS)mission measured light from stars in∼75% of the sky throughout its 2 yr primary mission, resulting in millions of TESS 30-minute-cadence light curves to analyze in the search for transiting exoplanets. To search this vast data trove for transit signals, we aim to provide an approach that both is computationally efficient and produces highly performant predictions. This approach minimizes the required human search effort. We present a convolutional neural network, which we train to identify planetary transit signals and dismiss false positives. To make a prediction for a given light curve, our network requires no prior transit parameters identified using other methods. Our network performs inference on a TESS 30-minute-cadence light curve in∼5 ms on a single GPU, enabling large-scale archival searches. We present 181 new planet candidates identified by our network, which pass subsequent human vetting designed to rule out false positives.Our neural network model is additionally provided as open-source code for public use and extension

Gregory Olmschenk

Leveraging the Usage of GPUs in SAR Processing for the NISAR Mission

The NASA ISRO Synthetic Aperture Radar (NISAR) mission will redefine the future of earth science in terms of both the quality as well as the quantity of data that will be downlinked daily. The current software architecture used to process this data is the InSAR Scientific Computing Environment (ISCE), a powerful and modular platform that applies a combination of novel and legacy processing modules to many sources of SAR data. Until recently, this architecture could process most images in a reasonable amount of time; however in the case of the NISAR mission (where the daily influx as well as the size of the images themselves are significantly larger) the current architecture can take hours to process even a single image. This paper explores new efforts to use a Graphics Processing Unit (GPU) to accelerate one of the processing modules to achieve unprecedented runtimes with no loss in precision, potentially setting a new standard in radar processing in the world of “Big Data”.

Cohen, Joshua

New Rover Conops with High-Performance Onboard Computing: Give Up Raw Data to Reduce Ops Cost and Do More Science

A major portion of time during the tactical operation of Mars rovers is spent for selecting, prioritizing, and coordinating sciences and engineering activities such that they fit within resource constraints, including the downlink data volume, energy, and time. In particular, the downlink data volume constraint is getting particularly tighter in recent missions because modern instruments produce increasingly high data volume while the communication bandwidth is essentially bounded by the law of physics. Tactical operation would be substantially simplified, hence the operation cost could be reduced, if the data volume constraint is relaxed or even removed. In this abstract, we propose a new operation paradigm for achieving this goal. The key observation is that, both in science and engineering applications, the bit size of raw data is typically much greater than the volume of processed information that is needed for scientific or engineering analysis. For example, a full-resolution image from Mastcam-Z, the main science camera on Perseverance, is about 700 kB in volume and we downlinked 29,685 images up to Sol 243, totaling ~20 GB of data. But of course, scientists do not use every pixel of these images; what they really look for in the images are geological features, typically represented by specific geometric configurations or textures. An end product after processing hundreds of Mascam-Z images could be a single geological map summarizing the spatial distribution of the features. For another example, a 100-meter drive of Perseverance produces 7-12 MB of drive telemetry, which records every detail of the rover's motion at 8 Hz, including position, attitude, steering angles, encoder readings, motor currents and many other information. But what the ground engineers eventually pay attention to is the signs of anomaly, such as excessive motor currents or high slip; if a drive is nominal, the vast majority of this data is unused. What if, then, we process the raw data onboard and only downlink the processed data that is relevant to scientific or engineering analyses, such as a list of detected science features (with cropped images) or a list of potential signs of anomaly while driving? A major roadblock for such onboard, high-level information processing has been the onboard computational resource. RAD750, the main onboard computer of Perseverance, is obviously not sufficient for performing complex image or signal processing such as object detection, semantic segmentation, or anomaly detection. Interestingly, RAD750 is not the best processor that Perseverance has; Qualcomm's Snapdragon 801, a modern mobile processor, is on her Heli Base Station, a device for communicating with Mars Helicopter Ingenuity; also, Intel's Atom E3845 processors are on engineering cameras. In the reminder of this paper, we will introduce two particular uses cases of these high-performance co-processors (meaning auxiliary CPU, GPU, or other types of processors that are separate from the main processor that runs the main flight software) for lowering operation cost and accommodating more science activities for a given communication constraint.

Didier, A.

UASs in the VOG/Edge/FOG Sensor Web Environment

This paper will describe the evolution of information collection, derivation and delivery mechanisms in sensor webs utilizing Uninhabited Aerial Systems (UAS).We will discuss the movement to "Fog Computing", also known as Edge Computing. Fog Computing facilitates the distribution of common operations and networking between edge devices and cloud computing facilities, optimizing the production of actionable intelligence. Initially, UASs utilized onboard data collection as standard, with minimal data downloaded directly. Information products were derived in conventional computational environments, generally desk top computers, and information products made available to the Science Community in weeks or months. With the increased availability, and increasingly lower costs, of beyond line of sight (BLOS) satellite based communication, transmission rates and data volumes increased, and processing migrated to Cloud based services. Contemporary UASs are moving some of that information product derivation to on vehicle services, and are creating a distributed Cloud/Fog environment. The Author will describe the technological advances that have made this possible, including low power multi-core Central Processing Units (CPU), and, more recently, the availability of high end Graphical Processing Units (GPU) that consume only a few watts. Intelligent system software, leveraging these hardware advances, finally allows for information product generation on-board, rather than simple data collection. Additionally, intelligent flight control systems now support mutual vehicle to vehicle collaboration, allowing UASs to create ad-hoc sensor webs on demand, as required. Also discussed will be the lessons learned by the Authors' development of data systems for NASA's large High Altitude Long Endurance (HALE) UASs like Predator and Global Hawk, and how those lessons are being applied to other UAS development This paper will focus on applications, rather a deep dive into the technology, and will highlight improving data management through these new technologies.

UAS

The Influence of Computer Architecture on Performance and Scaling for Hypersonic Flow Simulations

It is critical to understand how hypersonic simulation tools perform on a range of computational platforms. This information will aid in the acquisition of appropriate hardware and the potential refactoring of hypersonic codes to run on different systems. In this paper, we consider two representative high-speed reacting flow cases: a model Mach 8 hypersonic waverider glide vehicle and a model hydrocarbon-fueled hypersonic ramjet propulsion system. In both scenarios, the flow fields are in chemical non-equilibrium and are modeled by the multi-species reacting Navier-Stokes equations. For these simulations we use several hypersonic simulation tools, including US3D, Kestrel, FUN3D, and the JENRER© flow solver. We explore several high performance computing systems containing IntelR© XeonR© Platinum processors, AMD EPYCTM7702 processors, and NVIDIAR© Tesla V100 devices. We compare performance and strong scaling between the different systems.

CPU