Search NASA⌕ Search

SEARCH · Search NASA

Results for “Parallel Performance Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30

NAS Experiences of Porting CM Fortran Codes to HPF on IBM SP2 and SGI Power Challenge

Current Connection Machine (CM) Fortran codes developed for the CM-2 and the CM-5 represent an important class of parallel applications. Several users have employed CM Fortran codes in production mode on the CM-2 and the CM-5 for the last five to six years, constituting a heavy investment in terms of cost and time. With Thinking Machines Corporation's decision to withdraw from the hardware business and with the decommissioning of many CM-2 and CM-5 machines, the best way to protect the substantial investment in CM Fortran codes is to port the codes to High Performance Fortran (HPF) on highly parallel systems. HPF is very similar to CM Fortran and thus represents a natural transition. Conversion issues involved in porting CM Fortran codes on the CM-5 to HPF are presented. In particular, the differences between data distribution directives and the CM Fortran Utility Routines Library, as well as the equivalent functionality in the HPF Library are discussed. Several CM Fortran codes (Cannon algorithm for matrix-matrix multiplication, Linear solver Ax=b, 1-D convolution for 2-D datasets, Laplace's Equation solver, and Direct Simulation Monte Carlo (DSMC) codes have been ported to Subset HPF on the IBM SP2 and the SGI Power Challenge. Speedup ratios versus number of processors for the Linear solver and DSMC code are presented.

Saini, Subhash↗

Hardware Implementation of Lossless Adaptive and Scalable Hyperspectral Data Compression for Space

On-board lossless hyperspectral data compression reduces data volume in order to meet NASA and DoD limited downlink capabilities. The technique also improves signature extraction, object recognition and feature classification capabilities by providing exact reconstructed data on constrained downlink resources. At JPL a novel, adaptive and predictive technique for lossless compression of hyperspectral data was recently developed. This technique uses an adaptive filtering method and achieves a combination of low complexity and compression effectiveness that far exceeds state-of-the-art techniques currently in use. The JPL-developed 'Fast Lossless' algorithm requires no training data or other specific information about the nature of the spectral bands for a fixed instrument dynamic range. It is of low computational complexity and thus well-suited for implementation in hardware. A modified form of the algorithm that is better suited for data from pushbroom instruments is generally appropriate for flight implementation. A scalable field programmable gate array (FPGA) hardware implementation was developed. The FPGA implementation achieves a throughput performance of 58 Msamples/sec, which can be increased to over 100 Msamples/sec in a parallel implementation that uses twice the hardware resources This paper describes the hardware implementation of the 'Modified Fast Lossless' compression algorithm on an FPGA. The FPGA implementation targets the current state-of-the-art FPGAs (Xilinx Virtex IV and V families) and compresses one sample every clock cycle to provide a fast and practical real-time solution for space applications.

FPGA implementation↗

Sensor Co-design for $\textit{smartpixels}$

Pixel tracking detectors at upcoming collider experiments will see unprecedented charged-particle densities. Real-time data reduction on the detector will enable higher granularity and faster readout, possibly enabling the use of the pixel detector in the first level of the trigger for a hadron collider. This data reduction can be accomplished with a neural network (NN) in the readout chip bonded with the sensor that recognizes and rejects tracks with low transverse momentum (p$_T$) based on the geometrical shape of the charge deposition (``cluster''). To design a viable detector for deployment at an experiment, the dependence of the NN as a function of the sensor geometry, external magnetic field, and irradiation must be understood. In this paper, we present first studies of the efficiency and data reduction for planar pixel sensors exploring these parameters. A smaller sensor pitch in the bending direction improves the p$_T$ discrimination, but a larger pitch can be partially compensated with detector depth. An external magnetic field parallel to the sensor plane induces Lorentz drift of the electron-hole pairs produced by the charged particle, broadening the cluster and improving the network performance. The absence of the external field diminishes the background rejection compared to the baseline by $\mathcal{O}$(10%). Any accumulated radiation damage also changes the cluster shape, reducing the signal efficiency compared to the baseline by $\sim$ 30 - 60%, but nearly all of the performance can be recovered through retraining of the network and updating the weights. Finally, the impact of noise was investigated, and retraining the network on noise-injected datasets was found to maintain performance within 6% of the baseline network trained and evaluated on noiseless data.

Shekar, Danush [Illinois U., Chicago]↗

CORRLA-RS

The CORRLA-RS package provides a suite of statistical methods for sampling multidimensional distributions and to conduct sensitivity and correlation analysis of large scale data in the Rust programming language. The software provides a unique solution to multidimensional constrained sampling problems utilizing a combination of parallelized Markov Chain Monte Carlo methods and traditional rejection sampling. The sensitivity and correlation analysis methods are backed by a high performance randomized singular value decomposition implementation which enables datasets larger than the random access memory (RAM) size to be analyzed. Additionally, CORRLA-RS implements the active subspace identification method using a KD-Tree and the randomized singular value decomposition acting in concert.

Gurecky, William [Oak Ridge National Laboratory (O↗

SAR processing on the MPP

The processing of synthetic aperture radar (SAR) signals using the massively parallel processor (MPP) is discussed. The fast Fourier transform convolution procedures employed in the algorithms are described. The MPP architecture comprises an array unit (ARU) which processes arrays of data; an array control unit which controls the operation of the ARU and performs scalar arithmetic; a program and data management unit which controls the flow of data; and a unique staging memory (SM) which buffers and permutes data. The ARU contains a 128 by 128 array of bit-serial processing elements (PE). Two-by-four surarrays of PE's are packaged in a custom VLSI HCMOS chip. The staging memory is a large multidimensional-access memory which buffers and permutes data flowing with the system. Efficient SAR processing is achieved via ARU communication paths and SM data manipulation. Real time processing capability can be realized via a multiple ARU, multiple SM configuration.

Batcher, K. E.↗

Government-to-government cooperation in space station development

A memoranda of understanding was recently signed between the United States (NASA) and three international Space Station partners - Canada, European Space Agency (ESA), and Japan. The international partners are performing parallel Phase B preliminary design studies, concurrent with the U.S., on their proposed elements/systems for possible integration and operation with the U.S. Space Station System complex. During the 21-month Space Station Phase B study, a large amount of technical interface data will have to be transferred between the U.S. and the international partners. Scheduled bilateral technical coordination meetings will also be held. The coordination and large number of interfaces required to integrate the international requirements into the Space Station require a clean interface management organizational structure and operation procedures to accomplish the integration task. The international coordination management organizational structure, management tools, and communications network are discussed including the proposed international elements/systems being studied by the international partners.

Nassiff, S. H.↗

A system for remote measurements of the wind stress over the ocean

The DISSTRESS system for remote measurements of the surface wind stress over the ocean from ships and buoys is described. It is fully digital, utilizing the inertial dissipation technique. Parallel processing allows anemometer data to be filtered in natural frequency space; that is, the filter cutoffs shift linearly with the mean wind speed of the data to be filtered. The construction of the digital Butterworth bandpass filters is presented in detail. The performance of the system is evaluated by analyzing the results from 28 days of operation during the Frontal Air-Sea Interaction Experiment. The mean wind speed is checked, the anemometer response function is established, and drag coefficients are compared to previous studies. The capability of the system is demonstrated by continuous time series of the friction velocity computed every 20 min. The conclusion is that the surface wind stress can be measured more reliably and accurately (20 percent) with this system than from anemometer wind speeds and a bulk formula.

Large, William G.↗

Applications of massively parallel computers in telemetry processing

Telemetry processing refers to the reconstruction of full resolution raw instrumentation data with artifacts, of space and ground recording and transmission, removed. Being the first processing phase of satellite data, this process is also referred to as level-zero processing. This study is aimed at investigating the use of massively parallel computing technology in providing level-zero processing to spaceflights that adhere to the recommendations of the Consultative Committee on Space Data Systems (CCSDS). The workload characteristics, of level-zero processing, are used to identify processing requirements in high-performance computing systems. An example of level-zero functions on a SIMD MPP, such as the MasPar, is discussed. The requirements in this paper are based in part on the Earth Observing System (EOS) Data and Operation System (EDOS).

El-Ghazawi, Tarek A.↗

An Atmospheric General Circulation Model with Chemistry for the CRAY T3E: Design, Performance Optimization and Coupling to an Ocean Model

The design, implementation and performance optimization on the CRAY T3E of an atmospheric general circulation model (AGCM) which includes the transport of, and chemical reactions among, an arbitrary number of constituents is reviewed. The parallel implementation is based on a two-dimensional (longitude and latitude) data domain decomposition. Initial optimization efforts centered on minimizing the impact of substantial static and weakly-dynamic load imbalances among processors through load redistribution schemes. Recent optimization efforts have centered on single-node optimization. Strategies employed include loop unrolling, both manually and through the compiler, the use of an optimized assembler-code library for special function calls, and restructuring of parts of the code to improve data locality. Data exchanges and synchronizations involved in coupling different data-distributed models can account for a significant fraction of the running time. Therefore, the required scattering and gathering of data must be optimized. In systems such as the T3E, there is much more aggregate bandwidth in the total system than in any particular processor. This suggests a distributed design. The design and implementation of a such distributed 'Data Broker' as a means to efficiently couple the components of our climate system model is described.

Farrara, John D.↗

From Oxygen Generation to Metals Production: In Situ Resource Utilization by Molten Oxide Electrolysis

For the exploration of other bodies in the solar system, electrochemical processing is arguably the most versatile technology for conversion of local resources into usable commodities: by electrolysis one can, in principle, produce (1) breathable oxygen, (2) silicon for the fabrication of solar cells, (3) various reactive metals for use as electrodes in advanced storage batteries, and (4) structural metals such as steel and aluminum. Even so, to date there has been no sustained effort to develop such processes, in part due to the inadequacy of the database. The objective here is to identify chemistries capable of sustaining molten oxide electrolysis in the cited applications and to examine the behavior of laboratory-scale cells designed to generate oxygen and to produce metal. The basic research includes the study of the underlying high-temperature physical chemistry of oxide melts representative of lunar regolith and of Martian soil. To move beyond empirical approaches to process development, the thermodynamic and transport properties of oxide melts are being studied to help set the limits of composition and temperature for the processing trials conducted in laboratory-scale electrolysis cells. The goal of this investigation is to deliver a working prototype cell that can use lunar regolith and Martian soil to produce breathable oxygen along with metal by-product. Additionally, the process can be generalized to permit adaptation to accommodate different feedstock chemistries, such as those that will be encountered on other bodies in the solar system. The expected results of this research include: (1) the identification of appropriate electrolyte chemistries; (2) the selection of candidate anode and cathode materials compatible with electrolytes named above; and (3) performance data from a laboratory-scale cell producing oxygen and metal. On the strength of these results it should be possible to assess the technical viability of molten oxide electrolysis for in situ resource utilization on the Moon and Mars. In parallel, there may be commercial applications here on earth, such as new green technologies for metals extraction and for treatment of hazardous waste, e.g., fixing heavy metals.

Khetpal, Deepak↗

An Accelerated Clip Algorithm for Unstructured Meshes: A Batch-Driven Approach

The clip technique is a popular method for visualizing complex structures and phenomena within 3D unstructured meshes. Meshes can be clipped by specifying a scalar isovalue to produce an output unstructured mesh with its external surface as the isovalue. Similar to isocontouring, the clipping process relies on scalar data associated with the mesh points, including scalar data generated by implicit functions such as planes, boxes, and spheres, which facilitates the visualization of results interior to the grid. In this paper, we introduce a novel batch-driven parallel algorithm based on a sequential clip algorithm designed for high-quality results in partial volume extraction. Our algorithm comprises five passes, each progressively processing data to generate the resulting clipped unstructured mesh. The novelty lies in the use of fixed-size batches of points and cells, which enable rapid workload trimming and parallel processing, leading to a significantly improved memory footprint and run-time performance compared to the original version. On a 32-core CPU, the proposed batch-driven parallel algorithm demonstrates a run-time speed-up of up to 32.6x and a memory footprint reduction of up to 4.37x compared to the existing sequential algorithm. The software is currently available under an open-source license in the VTK visualization system.

Tsalikis, Spiros↗

Performance Assessment of OVERFLOW on Distributed Computing Environment

The aerodynamic computer code, OVERFLOW, with a multi-zone overset grid feature, has been parallelized to enhance its performance on distributed and shared memory paradigms. Practical application benchmarks have been set to assess the efficiency of code's parallelism on high-performance architectures. The code's performance has also been experimented with in the context of the distributed computing paradigm on distant computer resources using the Information Power Grid (IPG) toolkit, Globus. Two parallel versions of the code, namely OVERFLOW-MPI and -MLP, have developed around the natural coarse grained parallelism inherent in a multi-zonal domain decomposition paradigm. The algorithm invokes a strategy that forms a number of groups, each consisting of a zone, a cluster of zones and/or a partition of a large zone. Each group can be thought of as a process with one or multithreads assigned to it and that all groups run in parallel. The -MPI version of the code uses explicit message-passing based on the standard MPI library for sending and receiving interzonal boundary data across processors. The -MLP version employs no message-passing paradigm; the boundary data is transferred through the shared memory. The -MPI code is suited for both distributed and shared memory architectures, while the -MLP code can only be used on shared memory platforms. The IPG applications are implemented by the -MPI code using the Globus toolkit. While a computational task is distributed across multiple computer resources, the parallelism can be explored on each resource alone. Performance studies are achieved with some practical aerodynamic problems with complex geometries, consisting of 2.5 up to 33 million grid points and a large number of zonal blocks. The computations were executed primarily on SGI Origin 2000 multiprocessors and on the Cray T3E. OVERFLOW's IPG applications are carried out on NASA homogeneous metacomputing machines located at three sites, Ames, Langley and Glenn. Plans for the future will exploit the distributed parallel computing capability on various homogeneous and heterogeneous resources and large scale benchmarks. Alternative IPG toolkits will be used along with sophisticated zonal grouping strategies to minimize the communication time across the computer resources.

Djomehri, M. Jahed↗

Results of the Second U.S. Manned Suborbital Space Flight, July 21, 1961

This document presents the results of the second United States manned suborbital space flight. The data and flight description presented form a continuation of the information provided at an open conference held under the auspices of the National Aeronautics and Space Administration, in cooperation with the National Institutes of Health and the National Academy of Sciences, at the U.S. Department of State Auditorium on June 6, 1961. The papers presented herein generally parallel the presentations of the first report and were prepared by the personnel of the NASA Manned Spacecraft Center in collaboration with personnel from other government agencies, participating industry, and universities. The second successful manned suborbital space flight on July 21, 1961, in which Astronaut Virgil I. Grissom was the pilot was another step in the progressive research, development, and training program leading to the study of man's capabilities in a space environment during manned orbital flight. Data and operational experiences gained from this flight were in agreement with and supplemented the knowledge obtained from the first suborbital flight of May 5, 1961, piloted by Astronaut Alan B. Shepard, Jr. The two recent manned suborbital flights, coupled with the unmanned research and development flights, have provided valuable engineering nd scientific data on which the program can progress. The successful active participation of the pilots, in much the same way as in the development and testing of high performance aircraft, has. greatly increased our confidence in giving man a significant role in future space flight activities. It is the purpose of this report to continue the practice of providing data to the scientific community interested in activities of this nature. Brief descriptions are presented of the Project Mercury spacecraft and flight plan. Papers are provided which parallel the presentations of data published for the first suborbital space flight. Additional information is given relating to the operational aspects of the medical support activities for the two manned suborbital space flights.

Source record↗

Parallel quantum computing simulations via quantum accelerator platform virtualization

Quantum circuit execution is a central task in quantum computation. Due to inherent quantum-mechanical constraints, quantum computing workflows often involve a considerable number of independent measurements over a large set of slightly different quantum circuits. Here we discuss a simple model for parallelizing such quantum circuit executions that is based on introducing a large array of virtual quantum processing units (mapped to HPC nodes in our case) as a parallel quantum computing platform. Implemented within the XACC framework, the model can readily take advantage of its backend-agnostic features, enabling parallel quantum computing/simulation over any target backend supported by XACC. We illustrate the performance of this approach by demonstrating strong scaling in two pertinent domain science problems, namely in computing the gradients for the multi-contracted variational quantum eigensolver and in data-driven quantum circuit learning, where we vary the number of qubits and the number of circuit layers. Here, the latter simulation leverages the cuQuantum library to run efficiently on GPU-accelerated HPC platforms.

97 MATHEMATICS AND COMPUTING↗

Scaling Ultrahigh-Resolution E3SM Land Model for Leadership-Class Supercomputers

This paper presents advancements in scaling the ultrahigh-resolution E3SM Land Model (uELM) for deployment on leadership-class supercomputers, addressing the increased demand for km-scale Earth system modeling. By focusing on km-scale ELM simulations, we enhance predictive capabilities for climate interactions, facilitating improved responses to climate change impacts on energy systems, agriculture, and water resources. Our approach leverages innovative software architecture optimizations, sophisticated data handling techniques, and advanced parallel processing, achieving strong scalability on two leadership supercomputers (2400 nodes (105,600 cores) on Summit, and 1200 nodes (76,800 cores) on Frontier). Results from extensive scalability assessments on the Summit and Frontier also demonstrate outstanding I/O performance (close to 400 GB/s write throughput) and the model's ability to efficiently handle increasing computational demands. This study not only establishes uELM's capability for high-resolution simulations over vast geographical domains, but also sets a foundation for future Earth system modeling breakthroughs.

Wang, Dali [ORNL] (ORCID:0000000168065108)↗

Parametric analysis of RF communications and tracking systems for manned space stations

System performance, system interface compatibility, and system management and operations are analyzed for external and internal communications provided by the modular space station. Mathematical models are utilized to evaluate performances of the various communication links between the MSS and the tracking and data relay satellites, the shuttle, and the ground stations. Communication control design requirements are determined from overall MSS control design and checkout requirements are developed in a related systems study. Parallel design efforts consider equipment configurations for the external communication assembly baseband and the internal communication assembly to accommodate signal transfer and control. The recommended baseline design considers also interfaces between equipment groups and the critical functional and signal characteristics of these interfaces. The data processing assembly has direct communication circuits to the external communication assembly and the control interface is provided via a digital data bus and a remote acquisition and control unit.

Source record↗

Battery Reinitialization of the Photovoltaic Module of the International Space Station

The photovoltaic (PV) module on the International Space Station (ISS) has been operating since November 2000 and supporting electric power demands of the ISS and its crew of three. The PV module contains photovoltaic arrays that convert solar energy to electrical power and an integrated equipment assembly (IEA) that houses electrical hardware and batteries for electric power regulation and storage. Each PV module contains two independent power channels for fault tolerance. Each power channel contains three batteries in parallel to meet its performance requirements and for fault tolerance. Each battery consists of 76 Ni-Hydrogen (Ni-H2) cells in series. These 76 cells are contained in two orbital replaceable units (ORU) that are connected in series. On-orbit data are monitored and trended to ensure that all hardware is operating normally. Review of on-orbit data showed that while five batteries are operating very well, one is showing signs of mismatched ORUs. The cell pressure in the two ORUs differs by an amount that exceeds the recommended range. The reason for this abnormal behavior may be that the two ORUs have different use history. An assessment was performed and it was determined that capacity of this battery would be limited by the lower pressure ORU. Steps are being taken to reduce this pressure differential before battery capacity drops to the point of affecting its ability to meet performance requirements. As a first step, a battery reinitialization procedure was developed to reduce this pressure differential. The procedure was successfully carried out on-orbit and the pressure differential was reduced to the recommended range. This paper describes the battery performance and the consequences of mismatched ORUs that make a battery. The paper also describes the reinitialization procedure, how it was performed on orbit, and battery performance after the reinitialization. On-orbit data monitoring and trending is an ongoing activity and it will continue as ISS assembly progresses.

Hajela, Gyan↗

The Goes-R Geostationary Lightning Mapper (GLM): Algorithm and Instrument Status

The Geostationary Operational Environmental Satellite (GOES-R) is the next series to follow the existing GOES system currently operating over the Western Hemisphere. Superior spacecraft and instrument technology will support expanded detection of environmental phenomena, resulting in more timely and accurate forecasts and warnings. Advancements over current GOES capabilities include a new capability for total lightning detection (cloud and cloud-to-ground flashes) from the Geostationary Lightning Mapper (GLM), and improved capability for the Advanced Baseline Imager (ABI). The Geostationary Lighting Mapper (GLM) will map total lightning activity (in-cloud and cloud-to-ground lighting flashes) continuously day and night with near-uniform spatial resolution of 8 km with a product refresh rate of less than 20 sec over the Americas and adjacent oceanic regions. This will aid in forecasting severe storms and tornado activity, and convective weather impacts on aviation safety and efficiency. In parallel with the instrument development (a prototype and 4 flight models), a GOES-R Risk Reduction Team and Algorithm Working Group Lightning Applications Team have begun to develop the Level 2 algorithms, cal/val performance monitoring tools, and new applications. Proxy total lightning data from the NASA Lightning Imaging Sensor on the Tropical Rainfall Measuring Mission (TRMM) satellite and regional test beds are being used to develop the pre-launch algorithms and applications, and also improve our knowledge of thunderstorm initiation and evolution. A joint field campaign with Brazilian researchers in 2010-2011 will produce concurrent observations from a VHF lightning mapping array, Meteosat multi-band imagery, Tropical Rainfall Measuring Mission (TRMM) Lightning Imaging Sensor (LIS) overpasses, and related ground and in-situ lightning and meteorological measurements in the vicinity of Sao Paulo. These data will provide a new comprehensive proxy data set for algorithm and application development.

Goodman, Steven J.↗