Search NASA⌕ Search

SEARCH · Search NASA

Results for “Parallel Processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 955 records · Page 53

Analysis of DE-1 PWI electric field data

The measurement of low frequency electric field oscillations may be accomplished with the Plasma Wave Instrument (PWI) on DE 1. Oscillations at a frequency around 1 Hz are below the range of the conventional plasma wave receivers, but they can be detected by using a special processing of the quasi-static electric field data. With this processing it is also possible to determine if the electric field oscillations are predominately parallel or perpendicular to the ambient magnetic field. The quasi-static electric field in the DE 1 spin/orbit plane is measured with a long-wire 'double probe'. This antenna is perpendicular to the satellite spin axis, which in turn is approximately perpendicular to the geomagnetic field in the polar magnetosphere. The electric field data are digitally sampled at a frequency of 16 Hz. The measured electric field signal, which has had phase reversals introduced by the rotating antenna, is multiplied by the sine of the rotation angle between the antenna and the magnetic field. This is called the 'perpendicular' signal. The measured time series is also multiplied with the cosine of the angle to produce a separate 'parallel' signal. These two separate time series are then processed to determine the frequency power spectrum.

Weimer, Daniel↗

A model for simulation and processing of radar images

A model for recording, processing, presentation, and analysis of radar images in digital form is presented. The observed image is represented as having two random components, one which models the variation due to the coherent addition of electromagnetic energy scattered from different objects in the illuminated areas. This component is referred to as fading. The other component is a representation of the terrain variation which can be described as the actual signal which the radar is attempting to measure. The combination of these two components provides a description of radar images as being the output of a linear space-variant filter operating on the product of the fading and terrain random processes. In addition, the model is applied to a digital image processing problem using the design and implementation of enhancement scene. Finally, parallel approaches are being employed as possible means of solving other processing problems such as SAR image map-matching, data compression, and pattern recognition.

Stiles, J. A.↗

A Parallel Kinetic Model for Surface and Bulk Charge Storage in ε -MnO 2 Pseudocapacitors

Pseudocapacitive materials such as manganese dioxide (MnO 2 ) are attractive for energy storage applications due to their ability to combine the fast kinetics of capacitors with the higher energy density of battery-type systems. However, the electrochemical behavior of MnO 2 remains difficult to interpret mechanistically, in part because existing models often fail to distinguish between surface-based redox processes and bulk intercalation mechanisms. In this work, we develop a physics-based model that represents MnO 2 pseudocapacitance as a linear combination of two independent, parallel electrochemical processes: (i) the surface or near-surface redox storage and (ii) lithium ion intercalation into the bulk material. These two processes are treated with distinct kinetic and thermodynamic parameters and are assumed to proceed independently. The total measured current is assumed to be the sum of these two partial currents. We validate the model using rate-dependent cyclic voltammetry experiments, demonstrating that it captures key trends and provides physically interpretable parameters reflecting the relative contributions of surface and bulk processes. By enabling a clear separation between these mechanisms, the model offers a useful framework for analyzing pseudocapacitive materials and can guide the rational design of high-performance energy storage electrodes.

Energy - Storage↗

An Advanced Simulation Framework for Parallel Discrete-Event Simulation

Discrete-event simulation (DEVS) users have long been faced with a three-way trade-off of balancing execution time, model fidelity, and number of objects simulated. Because of the limits of computer processing power the analyst is often forced to settle for less than desired performances in one or more of these areas.

parallel processing technologies DEVS (PDEVS) ssto↗

Remote Objects Message Exchange (ROME)

The performance of a single program running on a single processor is limited by the character of the processor. Moreover, the cost and difficulty of developing and sustaining programs tend to increase as their size and complexity increase. Clearly there ought to be some advantage in partitioning powerful application software in relatively small and simple components that can run in parallel on multiple processors; the software should run faster and it should be cheaper and easier to deploy. Remote Objects Message Exchange (ROME) is an attempt to provide a single relatively simple, universally available abstraction for data communication among C++ objects. It aims to enable the C++ application developer to specify objects' interactions with other objects wholly in terms of the application domain, without concern for details of interprocess communication. Every ROME-compliant object is conceptually a network peer of every other, as if each one were (for example) a separate UNIX process.

processors application software data communication↗

On-Board AC Charging Topology Integrated with Electric Vehicle Motor Drive System

On-board AC charging is a convenient and widely adopted method for recharging electric vehicles (EVs) directly from standard alternating current (AC) power sources. This paper presents a novel topology for AC charging of EVs that utilizes EV 3-phase electric machine windings as the input inductors, thus eliminating the requirement for bulky grid interfacing inductors and resulting in a compact and cost-effective integrated motor drive and charger system. The proposed approach leverages the motor windings and parallel operating half-bridge inverter during the charging process by interconnecting the inverter phases with the motor windings in a mechanically interleaved and electrically paralleled manner. The implementation of this unique and innovative idea, achieved through precise control and arrangement of the motor winding as a series inductor, successfully eliminates the possibility of unintended motion of the electric machine during the charging process.

ADVANCED PROPULSION SYSTEMS↗

Parallel Implementation of the Recursive Approximation of an Unsupervised Hierarchical Segmentation Algorithm

The hierarchical image segmentation algorithm (referred to as HSEG) is a hybrid of hierarchical step-wise optimization (HSWO) and constrained spectral clustering that produces a hierarchical set of image segmentations. HSWO is an iterative approach to region grooving segmentation in which the optimal image segmentation is found at N(sub R) regions, given a segmentation at N(sub R+1) regions. HSEG's addition of constrained spectral clustering makes it a computationally intensive algorithm, for all but, the smallest of images. To counteract this, a computationally efficient recursive approximation of HSEG (called RHSEG) has been devised. Further improvements in processing speed are obtained through a parallel implementation of RHSEG. This chapter describes this parallel implementation and demonstrates its computational efficiency on a Landsat Thematic Mapper test scene.

Tilton, James C.↗

Optimal message log reclamation for independent checkpointing

Independent (uncoordinated) check pointing for parallel and distributed systems allows maximum process autonomy but suffers from possible domino effects and the associated storage space overhead for maintaining multiple checkpoints and message logs. In most research on check pointing and recovery, it was assumed that only the checkpoints and message logs older than the global recovery line can be discarded. It is shown how recovery line transformation and decomposition can be applied to the problem of efficiently identifying all discardable message logs, thereby achieving optimal garbage collection. Communication trace-driven simulation for several parallel programs is used to show the benefits of the proposed algorithm for message log reclamation.

Wang, Yi-Min↗

Optimal message log reclamation for independent checkpointing

Independent (uncoordinated) check pointing for parallel and distributed systems allows maximum process autonomy but suffers from possible domino effects and the associated storage space overhead for maintaining multiple checkpoints and message logs. In most research on check pointing and recovery, it was assumed that only the checkpoints and message logs older than the global recovery line can be discarded. It is shown how recovery line transformation and decomposition can be applied to the problem of efficiently identifying all discardable message logs, thereby achieving optimal garbage collection. Communication trace-driven simulation for several parallel programs is used to show the benefits of the proposed algorithm for message log reclamation.

Wang, Yi-Min↗

On the application of under-decimated filter banks

Maximally decimated filter banks have been extensively studied in the past. A filter bank is said to be under-decimated if the number of channels is more than the decimation ratio in the subbands. A maximally decimated filter bank is well known for its application in subband coding. Another application of maximally decimated filter banks is in block filtering. Convolution through block filtering has the advantages that parallelism is increased and data are processed at a lower rate. However, the computational complexity is comparable to that of direct convolution. More recently, another type of filter bank convolver has been developed. In this scheme, the convolution is performed in the subbands. Quantization and bit allocation of subband signals are based on signal variance, as in subband coding. Consequently, for a fixed rate, the result of convolution is more accurate than is direct convolution. This type of filter bank convolver also enjoys the advantages of block filtering, parallelism, and a lower working rate. Nevertheless, like block filtering, there is no computational saving. In this article, under-decimated systems are introduced to solve the problem. The new system is decimated only by half the number of channels. Two types of filter banks can be used in the under-decimated system: the discrete Fourier transform (DFT) filter banks and the cosine modulated filter banks. They are well known for their low complexity. In both cases, the system is approximately alias free, and the overall response is equivalent to a tunable multilevel filter. Properties of the DFT filter banks and the cosine modulated filter banks can be exploited to simultaneously achieve parallelism, computational saving, and a lower working rate. Furthermore, for both systems, the implementation cost of the analysis or synthesis bank is comparable to that of one prototype filter plus some low-complexity modulation matrices. The individual analysis and synthesis filters have complex coefficients in the DFT filter banks but have real coefficients in the cosine modulated filter banks.

Lin, Y.-P.↗

SPROC: A multiple-processor DSP IC

A large, single-chip, multiple-processor, digital signal processing (DSP) integrated circuit (IC) fabricated in HP-Cmos34 is presented. The innovative architecture is best suited for analog and real-time systems characterized by both parallel signal data flows and concurrent logic processing. The IC is supported by a powerful development system that transforms graphical signal flow graphs into production-ready systems in minutes. Automatic compiler partitioning of tasks among four on-chip processors gives the IC the signal processing power of several conventional DSP chips.

Davis, R.↗

Implementation and Testing of VLBI Software Correlation at the USNO

The Washington Correlator (WACO) at the U.S. Naval Observatory (USNO) is a dedicated VLBI processor based on dedicated hardware of ASIC design. The WACO is currently over 10 years old and is nearing the end of its expected lifetime. Plans for implementation and testing of software correlation at the USNO are currently being considered. The VLBI correlation process is, by its very nature, well suited to a parallelized computing environment. Commercial off-the-shelf computer hardware has advanced in processing power to the point where software correlation is now both economically and technologically feasible. The advantages of software correlation are manifold but include flexibility, scalability, and easy adaptability to changing environments and requirements. We discuss our experience with and plans for use of software correlation at USNO with emphasis on the use of the DiFX software correlator.

Fey, Alan↗

Massive parallelism in the future of science

Massive parallelism appears in three domains of action of concern to scientists, where it produces collective action that is not possible from any individual agent's behavior. In the domain of data parallelism, computers comprising very large numbers of processing agents, one for each data item in the result will be designed. These agents collectively can solve problems thousands of times faster than current supercomputers. In the domain of distributed parallelism, computations comprising large numbers of resource attached to the world network will be designed. The network will support computations far beyond the power of any one machine. In the domain of people parallelism collaborations among large groups of scientists around the world who participate in projects that endure well past the sojourns of individuals within them will be designed. Computing and telecommunications technology will support the large, long projects that will characterize big science by the turn of the century. Scientists must become masters in these three domains during the coming decade.

Denning, Peter J.↗

Legacy Code Modernization

Over the past decade, high performance computing has evolved rapidly; systems based on commodity microprocessors have been introduced in quick succession from at least seven vendors/families. Porting codes to every new architecture is a difficult problem; in particular, here at NASA, there are many large CFD applications that are very costly to port to new machines by hand. The LCM ("Legacy Code Modernization") Project is the development of an integrated parallelization environment (IPE) which performs the automated mapping of legacy CFD (Fortran) applications to state-of-the-art high performance computers. While most projects to port codes focus on the parallelization of the code, we consider porting to be an iterative process consisting of several steps: 1) code cleanup, 2) serial optimization,3) parallelization, 4) performance monitoring and visualization, 5) intelligent tools for automated tuning using performance prediction and 6) machine specific optimization. The approach for building this parallelization environment is to build the components for each of the steps simultaneously and then integrate them together. The demonstration will exhibit our latest research in building this environment: 1. Parallelizing tools and compiler evaluation. 2. Code cleanup and serial optimization using automated scripts 3. Development of a code generator for performance prediction 4. Automated partitioning 5. Automated insertion of directives. These demonstrations will exhibit the effectiveness of an automated approach for all the steps involved with porting and tuning a legacy code application for a new architecture.

Hribar, Michelle R.↗

Powder Bed Fusion Laser Beam Metals Additive Manufacturing: Process Monitoring Approaches for Qualification and Certification

The use of in-situ process monitoring is of interest to lower the cost of inspection for the qualification of powder bed fusion laser beam metal (PBF-LB/M) additively manufactured (AM) parts. Precise monitoring of the PBF-LB/M AM build process constitutes a multi-scale and multi-discipline task. There are several significant challenges to the in-situ approach: the synchronization of sensor signals to process steps; the physical interpretation and classification of sensor signals; managing very large datasets; and comparing the inputs with the observed monitoring signals. At NASA Langley Research Center, a configurable architecture additive testbed has been developed to monitor the build process with synchronized sensors. The philosophy and method adopted for the synchronization of the cameras with laser power and position throughout a complex PBF-LB/M AM build will be described. The synchronized in-situ monitoring signals are compared with ex-situ nondestructive inspection, x-ray computed tomography (XCT). Such comparisons permit a better understanding of how the sequential process actions of LPBF-AM can affect build quality. The multi-scale and complex process of printing additively manufactured (AM) parts can have unexpected, but predictable, build conditions that result in material microstructure variability. This presentation will describe an additive manufacturing model-based process metric (AM-PM) computational method that is a fully parallel reduced order modeling approach developed to evaluate the evolution of AM processes. This method couples the known sequence of the AM process with a physically informed nearest neighbors’ calculation to map the conditions of a part-scale build. The result is a map of the build that is derived directly from build files or in-situ process monitoring sensors. The methodology of the approach will be described and mapped to the porosity observed from XCT for a complex PBF-LB/M build. Such comparative results develop understanding of how the sequential process actions can affect the PBF-LB/M AM build quality and microstructure variability.

Laser Powder Bed Fusion↗

RAM-Based parallel-output controller

Selected bit strings in serial-data link are extracted for processing. Controller is programmable interface between serial-data link and peripherals that accept parallel data. It can be used to drive displays, printers, plotters, digital-to-analog converters, and parallel-output ports.

Niswander, J. K.↗

Parallelization of a Six Degree of Freedom Entry Vehicle Trajectory Simulation Using OpenMP and OpenACC

The art and science of writing parallelized software, using methods such as Open Multi-Processing (OpenMP) and Open Accelerators (OpenACC), is dominated by computer scientists. Engineers and non-computer scientists looking to apply these techniques to their project applications face a steep learning curve, especially when looking to adapt their original single threaded software to run multi-threaded on graphics processing units (GPUs). There are significant changes in mindset that must occur; such as how to manage memory, the organization of instructions, and the use of if statements (also known as branching). The purpose of this work is twofold: 1) to demonstrate the applicability of parallelized coding methodologies, OpenMP and OpenACC, to tasks outside of the typical large scale matrix mathematics; and 2) to discuss, from an engineer’s perspective, the lessons learned from parallelizing software using these computer science techniques. This work applies OpenMP, on both multi-core central processing units (CPUs) and Intel® Xeon Phi™ 7210, and OpenACC on GPUs. These parallelization techniques are used to tackle the simulation of thousands of entry vehicle trajectories through the integration of six degree of freedom (DoF) equations of motion (EoM). The forces and moments acting on the entry vehicle, and used by the EoM, are estimated using multiple models of varying levels of complexity. Several benchmark comparisons are made on the execution of six DoF trajectory simulation: single thread Intel® Xeon® E5-2670 CPU, multi-thread CPU using OpenMP, multi-thread Xeon Phi™ 7210 using OpenMP, and multi-thread NVIDIA® Tesla® K40 GPU using OpenACC. These benchmarks are run on the Pleiades Supercomputer Cluster at the National Aeronautics and Space Administration (NASA) Ames Research Center (ARC), and a Xeon Phi™ 7210 node at NASA Langley Research Center (LaRC).

Green, Justin S.↗