Search NASA⌕ Search

SEARCH · Search NASA

Results for “Parallel Performance Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 883 records · Page 49

Design of object-oriented distributed simulation classes

Distributed simulation of aircraft engines as part of a computer aided design package is being developed by NASA Lewis Research Center for the aircraft industry. The project is called NPSS, an acronym for 'Numerical Propulsion Simulation System'. NPSS is a flexible object-oriented simulation of aircraft engines requiring high computing speed. It is desirable to run the simulation on a distributed computer system with multiple processors executing portions of the simulation in parallel. The purpose of this research was to investigate object-oriented structures such that individual objects could be distributed. The set of classes used in the simulation must be designed to facilitate parallel computation. Since the portions of the simulation carried out in parallel are not independent of one another, there is the need for communication among the parallel executing processors which in turn implies need for their synchronization. Communication and synchronization can lead to decreased throughput as parallel processors wait for data or synchronization signals from other processors. As a result of this research, the following have been accomplished. The design and implementation of a set of simulation classes which result in a distributed simulation control program have been completed. The design is based upon MIT 'Actor' model of a concurrent object and uses 'connectors' to structure dynamic connections between simulation components. Connectors may be dynamically created according to the distribution of objects among machines at execution time without any programming changes. Measurements of the basic performance have been carried out with the result that communication overhead of the distributed design is swamped by the computation time of modules unless modules have very short execution times per iteration or time step. An analytical performance model based upon queuing network theory has been designed and implemented. Its application to realistic configurations has not been carried out.

Schoeffler, James D.↗

Aeroacoustic Characteristics of a Rectangular Multi-Element Supersonic Jet Mixer-Ejector Nozzle

This paper provides a unique, detailed evaluation of the acoustics and aerodynamics of a rectangular multi-element supersonic jet mixer-ejector noise suppressor. The performance of such mixer-ejectors is important in aircraft engine application for noise suppression and thrust augmentation. In contrast to most prior experimental studies on ejectors that reported either aerodynamic or acoustic data, our work documents both types of data. We present information on the mixing, pumping, ejector wall pressure distribution, thrust augmentation and noise suppression characteristics of four simple, multi-element, jet mixer-ejector configurations. The four configurations included the effect of ejector area ratio (AR = ejector area/primary jet area) and the effect of non-parallel ejector walls. We also studied in detail the configuration that produced the best noise suppression characteristics. Our results show that ejector configurations that produced the maximum maximum pumping (entrained flow per secondary inlet area) also exhibited the lowest wall pressures in the inlet region, and the maximum thrust augmentation. When cases having the same total mass flow were compared, we found that noise suppression trends corresponded with those for pumping. Surprisingly, the mixing (quantified by the peak Mach number, and flow uniformity) at the ejector exit exhibited no relationship to the noise suppression at moderate primary jet fully expanded Mach numbers (Mj is less than 1.4). However, the noise suppression dependence on the mixing was apparent at higher Mj. The above observations are justified by noting that the mixing at the ejector exit is ot a strong factor in determining the radiated noise when noise produced internal to the ejector dominates the noise field outside the ejector.

Raman, Ganesh↗

Design of Object-Oriented Distributed Simulation Classes

Distributed simulation of aircraft engines as part of a computer aided design package being developed by NASA Lewis Research Center for the aircraft industry. The project is called NPSS, an acronym for "Numerical Propulsion Simulation System". NPSS is a flexible object-oriented simulation of aircraft engines requiring high computing speed. It is desirable to run the simulation on a distributed computer system with multiple processors executing portions of the simulation in parallel. The purpose of this research was to investigate object-oriented structures such that individual objects could be distributed. The set of classes used in the simulation must be designed to facilitate parallel computation. Since the portions of the simulation carried out in parallel are not independent of one another, there is the need for communication among the parallel executing processors which in turn implies need for their synchronization. Communication and synchronization can lead to decreased throughput as parallel processors wait for data or synchronization signals from other processors. As a result of this research, the following have been accomplished. The design and implementation of a set of simulation classes which result in a distributed simulation control program have been completed. The design is based upon MIT "Actor" model of a concurrent object and uses "connectors" to structure dynamic connections between simulation components. Connectors may be dynamically created according to the distribution of objects among machines at execution time without any programming changes. Measurements of the basic performance have been carried out with the result that communication overhead of the distributed design is swamped by the computation time of modules unless modules have very short execution times per iteration or time step. An analytical performance model based upon queuing network theory has been designed and implemented. Its application to realistic configurations has not been carried out.

Schoeffler, James D.↗

Commercial Off-The-Shelf GPU Qualification for Space Applications

With increased sensor data rates, and limited downlink capability, NASA missions have increased demands for onboard processing for applications ranging from synthetic aperture radar (SAR) data reduction to hyperspectral image processing and recognition, and even artificial intelligence (AI). Graphics Processor Units (GPUs) offer an attractive processing architecture for many of the applications due to their massive parallelism. As no radiation hardened GPU devices currently exist, any near term GPU-based onboard processors must use commercially available devices. To address this need NASA GSFC is collaborating with Cubic Aerospace Incorporated to, (a) characterize the capability of GPUs to meet the demands of a candidate onboard processing application, thereby demonstrating their ability to improve mission performance, reduce spacecraft SWaP, and potentially enable new missions, and (b) evaluate the radiation tolerance of capable COTS GPU devices to determine their suitability for spaceflight applications and understand any mitigations that are needed. A candidate onboard processing image has been prototyped and evaluated on a commercial GPU board and has demonstrated significantly increased processing throughput. Radiation tests for commercial GPU devices are planned for early fiscal year 2019.

Onboard processing↗

Rocket-based measurement of Birkeland currents related to an auroral arc and electrojet.

A rocket-borne experiment performed to study currents associated with a quiet auroral arc is discussed. The magnetic field in the vicinity of the arc was measured with a vector magnetometer, while the orientation of the payload was determined with a lunar-aspect sensor. Possible current configurations were inferred by constructing model current systems that reproduced the magnetic field variations measured along the flight path. The data are interpreted in terms of a model current system consisting of a northwestward electrojet and two oppositely directed Birkeland sheet currents, all lying in planes approximately parallel to the auroral arc. The current density integrated through the 16-km thickness of each Birkeland current sheet was found to be roughly 0.16 A/m. The electrojet current was roughly 6000 A.

Park, R. J.↗

Development of a Grid-Independent Geos-Chem Chemical Transport Model (v9-02) as an Atmospheric Chemistry Module for Earth System Models

The GEOS-Chem global chemical transport model (CTM), used by a large atmospheric chemistry research community, has been re-engineered to also serve as an atmospheric chemistry module for Earth system models (ESMs). This was done using an Earth System Modeling Framework (ESMF) interface that operates independently of the GEOSChem scientific code, permitting the exact same GEOSChem code to be used as an ESM module or as a standalone CTM. In this manner, the continual stream of updates contributed by the CTM user community is automatically passed on to the ESM module, which remains state of science and referenced to the latest version of the standard GEOS-Chem CTM. A major step in this re-engineering was to make GEOS-Chem grid independent, i.e., capable of using any geophysical grid specified at run time. GEOS-Chem data sockets were also created for communication between modules and with external ESM code. The grid-independent, ESMF-compatible GEOS-Chem is now the standard version of the GEOS-Chem CTM. It has been implemented as an atmospheric chemistry module into the NASA GEOS- 5 ESM. The coupled GEOS-5-GEOS-Chem system was tested for scalability and performance with a tropospheric oxidant-aerosol simulation (120 coupled species, 66 transported tracers) using 48-240 cores and message-passing interface (MPI) distributed-memory parallelization. Numerical experiments demonstrate that the GEOS-Chem chemistry module scales efficiently for the number of cores tested, with no degradation as the number of cores increases. Although inclusion of atmospheric chemistry in ESMs is computationally expensive, the excellent scalability of the chemistry module means that the relative cost goes down with increasing number of cores in a massively parallel environment.

CTM↗

Communications Technology Assessment for the Unmanned Aircraft System (UAS) Control and Non-Payload Communications (CNPC) Link

The National Aeronautics and Space Administration (NASA) Glenn Research Center (GRC) is performing communications systems research for the Unmanned Aircraft System (UAS) in the National Airspace System (NAS) Project. One of the goals of the communications element is to select and test a communications technology for the UAS Control and Non-Payload Communications (CNPC) link. The GRC UAS Modeling and Simulation (M/S) Sub Team will evaluate the performance of several potential technologies for the CNPC link through detailed software simulations. In parallel, an industry partner will implement a technology in hardware to be used for flight testing. The task necessitated a technical assessment of existing Radio Frequency (RF) communications technologies to identify the best candidate systems for use as the UAS CNPC link. The assessment provides a basis for selecting the technologies for the M/S effort and the hardware radio design. The process developed for the technical assessments for the Future Communications Study1 (FCS) was used as an initial starting point for this assessment. The FCS is a joint Federal Aviation Administration (FAA) and Eurocontrol study on technologies for use as a future aeronautical communications link. The FCS technology assessment process methodology can be applied to the UAS CNPC link; however the findings of the FCS are not directly applicable because of different requirements between a CNPC link and a general aeronautical data link. Additional technologies were added to the potential technologies list from the State of the Art Unmanned Aircraft System Communication Assessment developed by NASA GRC2. This document investigates the state of the art of communications as related to UAS. A portion of the document examines potential communications systems for a UAS communication architecture. Like the FCS, the state of the art assessment surveyed existing communications technologies. It did not, however, perform a detailed assessment of the technology necessary to recommend a technology for the UAS CNPC link. The technical assessment process, as shown in Figure 1, consists of the following steps. First, candidate RF communications technologies are identified. An initial review of each of these technologies is then performed to determine if the technology appears to be a good candidate and requires further review. Any technology that can be shown to be inadequate at that point is removed from consideration to allow for more detailed analysis of the remaining technologies. Criteria for the detailed assessments are defined and a scoring methodology is devised. This is followed by the detailed review and scoring of each technology. The least favorable technologies are removed during the process until only the few best candidates remain.

Aircraft Command and Control↗

Dependence of field-aligned electron precipitation occurrence on season and altitude

An examination of factors affecting the occurrence of field-aligned 2.3-keV electron precipitation has been performed by using data from more than 7500 orbits of the polar-orbiting satellite Ogo 4. Both season and altitude were found to be parameters that are directly related to the probability of occurrence. The highest probabilities occurred when the measurements were made at altitudes from 800 km to apogee (914 km), except during summer. In this altitude interval, the electron precipitation was more likely to be field-aligned during winter than during any other season. The analysis suggests the establishment by electrostatic charge layers of localized electric fields parallel to the magnetic field. The resulting potential distribution focuses the electron beam along the field lines in the region between the charge layers but destroys the focused beam below the lower layer, and thus an altitude dependence is created.

Berko, F. W.↗

Development of a Computer Architecture to Support the Optical Plume Anomaly Detection (OPAD) System

The NASA OPAD spectrometer system relies heavily on extensive software which repetitively extracts spectral information from the engine plume and reports the amounts of metals which are present in the plume. The development of this software is at a sufficiently advanced stage where it can be used in actual engine tests to provide valuable data on engine operation and health. This activity will continue and, in addition, the OPAD system is planned to be used in flight aboard space vehicles. The two implementations, test-stand and in-flight, may have some differing requirements. For example, the data stored during a test-stand experiment are much more extensive than in the in-flight case. In both cases though, the majority of the requirements are similar. New data from the spectrograph is generated at a rate of once every 0.5 sec or faster. All processing must be completed within this period of time to maintain real-time performance. Every 0.5 sec, the OPAD system must report the amounts of specific metals within the engine plume, given the spectral data. At present, the software in the OPAD system performs this function by solving the inverse problem. It uses powerful physics-based computational models (the SPECTRA code), which receive amounts of metals as inputs to produce the spectral data that would have been observed, had the same metal amounts been present in the engine plume. During the experiment, for every spectrum that is observed, an initial approximation is performed using neural networks to establish an initial metal composition which approximates as accurately as possible the real one. Then, using optimization techniques, the SPECTRA code is repetitively used to produce a fit to the data, by adjusting the metal input amounts until the produced spectrum matches the observed one to within a given level of tolerance. This iterative solution to the original problem of determining the metal composition in the plume requires a relatively long period of time to execute the software in a modern single-processor workstation, and therefore real-time operation is currently not possible. A different number of iterations may be required to perform spectral data fitting per spectral sample. Yet, the OPAD system must be designed to maintain real-time performance in all cases. Although faster single-processor workstations are available for execution of the fitting and SPECTRA software, this option is unattractive due to the excessive cost associated with very fast workstations and also due to the fact that such hardware is not easily expandable to accommodate future versions of the software which may require more processing power. Initial research has already demonstrated that the OPAD software can take advantage of a parallel computer architecture to achieve the necessary speedup. Current work has improved the software by converting it into a form which is easily parallelizable. Timing experiments have been performed to establish the computational complexity and execution speed of major components of the software. This work provides the foundation of future work which will create a fully parallel version of the software executing in a shared-memory multiprocessor system.

Katsinis, Constantine↗

ZBLAN Viscosity Instrumentation

The past year's contribution from Dr. Kaukler's experimental effort consists of these 5 parts: a) Construction and proof-of-concept testing of a novel shearing plate viscometer designed to produce small shear rates and operate at elevated temperatures; b) Preparing nonlinear polymeric materials to serve as standards of nonlinear Theological behavior; c) Measurements and evaluation of above materials for nonlinear rheometric behavior at room temperature using commercial spinning cone and plate viscometers available in the lab; d) Preparing specimens from various forms of pitch for quantitative comparative testing in a Dynamic Mechanical Analyzer, Thermal Mechanical Analyzer; and Archeological Analyzer; e) Arranging to have sets of pitch specimens tested using the various instruments listed above, from different manufacturers, to form a baseline of the viscosity variation with temperature using the different test modes offered by these instruments by compiling the data collected from the various test results. Our focus in this project is the shear thinning behavior of ZBLAN glass over a wide range of temperature. Experimentally, there are no standard techniques to perform such measurements on glasses, particularly at elevated temperatures. Literature reviews to date have shown that shear thinning in certain glasses appears to occur, but no data is available for ZBLAN glass. The best techniques to find shear thinning behavior require the application of very low rates of shear. In addition, because the onset of the thinning behavior occurs at an unknown elevated temperature, the instruments used in this study must provide controlled low rates of shear and do so for temperatures approaching 600 C. In this regard, a novel shearing parallel plate viscometer was designed and a prototype built and tested.

Kaukler, William↗

Rapid Corner Detection Using FPGAs

In order to perform precision landings for space missions, a control system must be accurate to within ten meters. Feature detection applied against images taken during descent and correlated against the provided base image is computationally expensive and requires tens of seconds of processing time to do just one image while the goal is to process multiple images per second. To solve this problem, this algorithm takes that processing load from the central processing unit (CPU) and gives it to a reconfigurable field programmable gate array (FPGA), which is able to compute data in parallel at very high clock speeds. The workload of the processor then becomes simpler; to read an image from a camera, it is transferred into the FPGA, and the results are read back from the FPGA. The Harris Corner Detector uses the determinant and trace to find a corner score, with each step of the computation occurring on independent clock cycles. Essentially, the image is converted into an x and y derivative map. Once three lines of pixel information have been queued up, valid pixel derivatives are clocked into the product and averaging phase of the pipeline. Each x and y derivative is squared against itself, as well as the product of the ix and iy derivative, and each value is stored in a WxN size buffer, where W represents the size of the integration window and N is the width of the image. In this particular case, a window size of 5 was chosen, and the image is 640 480. Over a WxN size window, an equidistance Gaussian is applied (to bring out the stronger corners), and then each value in the entire window is summed and stored. The required components of the equation are in place, and it is just a matter of taking the determinant and trace. It should be noted that the trace is being weighted by a constant k, a value that is found empirically to be within 0.04 to 0.15 (and in this implementation is 0.05). The constant k determines the number of corners available to be compared against a threshold sigma to mark a valid corner. After a fixed delay from when the first pixel is clocked in (to fill the pipeline), a score is achieved after each successive clock. This score corresponds with an (x,y) location within the image. If the score is higher than the predetermined threshold sigma, then a flag is set high and the location is recorded.

Morfopoulos, Arin C.↗

Numerical eigen-spectrum slicing, accurate orthogonal eigen-basis, and mixed-precision eigenvalue refinement using OpenMP data-dependent tasks and accelerator offload

Performing a variety of numerical computations efficiently and, at the same time, in a portable fashion requires both an overarching design followed by a number of implementation strategies. All of these are exemplified below as we present transitioning the PLASMA numerical library from relying on dependence-driven large tasks to achieving utilization of fine grain tasking and offload to hardware accelerators while keeping its core dependence sets: OpenMP source code pragmas and runtime for most system-level functionality and basic low-level numerical kernels provided directly by hardware vendors or open source projects with vendor contributions. We also present new algorithmic methods and their efficient parallel implementations including fine grained tasking for eigen-spectrum slicing and offload for mixed-precision eigenvalue refinement. We provide performance, scaling, and numerical results showing sizable gains over the available solutions from either the open source and vendor-provided packages.

Luszczek, Piotr↗

TRAPEDS: Producing traces for multicomputers via execution-driven simulation

Trace-driven simulation is an important aid in performance analysis of computer systems. Capturing address traces for these simulations is a difficult problem for single processors and particularly for multicomputers. Even when existing trace methods can be used on multicomputers, the amount of collected data typically grows with the number of processors, so I/O and trace storage costs increase. A new technique is presented which modifies the executable code to dynamically collect the address trace from the user code and analyzes this trace during the execution of the program. This method helps resolve the I/O and storage problems and facilitates parallel analysis of the address trace. If a trace stored on disk is desired, the generated trace information can also be written to files during execution, with a resultant drop in program execution speed. An initial implementation on the Intel iPSC/2 hypercube multicomputer is detailed, and sample simulation results are presented. The effect of this trace collection method on execution time is illustrated.

Stunkel, Craig B.↗

TRAPEDS - Producing traces for multicomputers via execution driven simulation

Trace-driven simulation is an important aid in performance analysis of computer systems. Capturing address traces for these simulations is a difficult problem for single processors and particularly for multicomputers. Even when existing trace methods can be used on multicomputers, the amount of collected data typically grows with the number of processors, so I/O and trace storage costs increase. A new technique is presented which modifies the executable code to dynamically collect the address trace from the user code and analyzes this trace during the execution of the program. This method helps resolve the I/O and storage problems and facilitates parallel analysis of the address trace. If a trace stored on disk is desired, the generated trace information can also be written to files during execution, with a resultant drop in program execution speed. An initial implementation on the Intel iPSC/2 hypercube multicomputer is detailed, and sample simulation results are presented. The effect of this trace collection method on execution time is illustrated.

Stunkel, Craig B.↗

High-performance parallel analysis of coupled problems for aircraft propulsion

Applications are described of high-performance parallel, computation for the analysis of complete jet engines, considering its multi-discipline coupled problem. The coupled problem involves interaction of structures with gas dynamics, heat conduction and heat transfer in aircraft engines. The methodology issues addressed include: consistent discrete formulation of coupled problems with emphasis on coupling phenomena; effect of partitioning strategies, augmentation and temporal solution procedures; sensitivity of response to problem parameters; and methods for interfacing multiscale discretizations in different single fields. The computer implementation issues addressed include: parallel treatment of coupled systems; domain decomposition and mesh partitioning strategies; data representation in object-oriented form and mapping to hardware driven representation, and tradeoff studies between partitioning schemes and fully coupled treatment.

Felippa, C. A.↗

High-Efficiency High-Resolution Global Model Developments at the NASA Goddard Data Assimilation Office

The Data Assimilation Office (DAO) has been developing a new generation of ultra-high resolution General Circulation Model (GCM) that is suitable for 4-D data assimilation, numerical weather predictions, and climate simulations. These three applications have conflicting requirements. For 4-D data assimilation and weather predictions, it is highly desirable to run the model at the highest possible spatial resolution (e.g., 55 km or finer) so as to be able to resolve and predict socially and economically important weather phenomena such as tropical cyclones, hurricanes, and severe winter storms. For climate change applications, the model simulations need to be carried out for decades, if not centuries. To reduce uncertainty in climate change assessments, the next generation model would also need to be run at a fine enough spatial resolution that can at least marginally simulate the effects of intense tropical cyclones. Scientific problems (e.g., parameterization of subgrid scale moist processes) aside, all three areas of application require the model's computational performance to be dramatically improved as compared to the previous generation. In this talk, I will present the current and future developments of the "finite-volume dynamical core" at the Data Assimilation Office. This dynamical core applies modem monotonicity preserving algorithms and is genuinely conservative by construction, not by an ad hoc fixer. The "discretization" of the conservation laws is purely local, which is clearly advantageous for resolving sharp gradient flow features. In addition, the local nature of the finite-volume discretization also has a significant advantage on distributed memory parallel computers. Together with a unique vertically Lagrangian control volume discretization that essentially reduces the dimension of the computational problem from three to two, the finite-volume dynamical core is very efficient, particularly at high resolutions. I will also present the computational design of the dynamical core using a hybrid distributed-shared memory programming paradigm that is portable to virtually any of today's high-end parallel super-computing clusters.

Lin, Shian-Jiann↗

High-Efficiency High-Resolution Global Model Developments at the NASA Goddard Data Assimilation Office

The Data Assimilation Office (DAO) has been developing a new generation of ultra-high resolution General Circulation Model (GCM) that is suitable for 4-D data assimilation, numerical weather predictions, and climate simulations. These three applications have conflicting requirements. For 4-D data assimilation and weather predictions, it is highly desirable to run the model at the highest possible spatial resolution (e.g., 55 kin or finer) so as to be able to resolve and predict socially and economically important weather phenomena such as tropical cyclones, hurricanes, and severe winter storms. For climate change applications, the model simulations need to be carried out for decades, if not centuries. To reduce uncertainty in climate change assessments, the next generation model would also need to be run at a fine enough spatial resolution that can at least marginally simulate the effects of intense tropical cyclones. Scientific problems (e.g., parameterization of subgrid scale moist processes) aside, all three areas of application require the model's computational performance to be dramatically improved as compared to the previous generation. In this talk, I will present the current and future developments of the "finite-volume dynamical core" at the Data Assimilation Office. This dynamical core applies modem monotonicity preserving algorithms and is genuinely conservative by construction, not by an ad hoc fixer. The "discretization" of the conservation laws is purely local, which is clearly advantageous for resolving sharp gradient flow features. In addition, the local nature of the finite-volume discretization also has a significant advantage on distributed memory parallel computers. Together with a unique vertically Lagrangian control volume discretization that essentially reduces the dimension of the computational problem from three to two, the finite-volume dynamical core is very efficient, particularly at high resolutions. I will also present the computational design of the dynamical core using a hybrid distributed- shared memory programming paradigm that is portable to virtually any of today's high-end parallel super-computing clusters.

Lin, Shian-Jiann↗

Arguments for the Physical Nature of the Triggered Ion-Acoustic Waves Observed on the Parker Solar Probe

Triggered ion-acoustic waves are a pair of coupled waves observed in the previously unexplored plasma regime near the Sun. They may be capable of producing important effects on the solar wind. Because this wave mode has not been observed or studied previously and it is not fully understood, the issue of whether it has a natural origin or is an instrumental artifact can be raised. This paper discusses this issue by examining 13 features of the data such as whether the triggered ion-acoustic waves are electrostatic, whether they are both narrowband, whether they satisfy the requirement that the electric field is parallel to the k-vector, whether the phase difference between the electric field and the density fluctuations is 90°, whether the two waves have the same phase velocity as they must if they are coupled, whether the phase velocity is that of an ion-acoustic wave, whether they are associated with other parameters such as electron heating, whether the electric field instrument otherwise performed as expected, etc. The conclusion reached from these analyses is that triggered ion-acoustic waves are highly likely to have a natural origin although the possibility that they are artifacts unrelated to processes occurring in the natural plasma cannot be eliminated. This inability to absolutely rule out artifacts as the source of a measured result is a characteristic of all measurements.

Electric fields↗