Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallel processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,639 records · Page 91

Design and Implementation of a Parallel Multivariate Ensemble Kalman Filter for the Poseidon Ocean General Circulation Model

A multivariate ensemble Kalman filter (MvEnKF) implemented on a massively parallel computer architecture has been implemented for the Poseidon ocean circulation model and tested with a Pacific Basin model configuration. There are about two million prognostic state-vector variables. Parallelism for the data assimilation step is achieved by regionalization of the background-error covariances that are calculated from the phase-space distribution of the ensemble. Each processing element (PE) collects elements of a matrix measurement functional from nearby PEs. To avoid the introduction of spurious long-range covariances associated with finite ensemble sizes, the background-error covariances are given compact support by means of a Hadamard (element by element) product with a three-dimensional canonical correlation function. The methodology and the MvEnKF configuration are discussed. It is shown that the regionalization of the background covariances; has a negligible impact on the quality of the analyses. The parallel algorithm is very efficient for large numbers of observations but does not scale well beyond 100 PEs at the current model resolution. On a platform with distributed memory, memory rather than speed is the limiting factor.

Keppenne, Christian L.↗

A high-order language for a system of closely coupled processing elements

The research reported in this paper was occasioned by the requirements on part of the Real-Time Digital Simulator (RTDS) project under way at NASA Lewis Research Center. The RTDS simulation scheme employs a network of CPUs running lock-step cycles in the parallel computations of jet airplane simulations. Their need for a high order language (HOL) that would allow non-experts to write simulation applications and that could be implemented on a possibly varying network can best be fulfilled by using the programming language Ada. We describe how the simulation problems can be modeled in Ada, how to map a single, multi-processing Ada program into code for individual processors, regardless of network reconfiguration, and why some Ada language features are particulary well-suited to network simulations.

Feyock, S.↗

A parallel-pipeline architecture of the fast polynomial transform for computing a two-dimensional cyclic convolution

It is pointed out that the two-dimensional cyclic convolution is a useful tool for many two-dimensional digital signal processing applications. Two important applications are related to spaceborne high-resolution synthetic aperture radar (SAR) processing and image processing. Nussbaumer and Quandalle (1978) showed that a radix-2 polynomial transform analogous to the conventional radix-2 FFT algorithm can be used to compute a two-dimensional cyclic convolution. On the basis of results reported by Arambepola and Rayner (1979), a radix-2 polynomial transform can be defined to compute a multidimensional cyclic convolution. Truong et al. (1981) used the considered ideas together with the Chinese Theorem to further reduce the complexity of the radix-2 fast polynomial transform (FPT). Reed et al. (1981) demonstrated that such a new FPT algorithm is significantly faster than the FFT algorithm for computing a two-dimensional convolution. In the present investigation, a parallel-pipeline architecture is considered for implementing the FPT developed by Truong et al.

Truong, T. K.↗

RLGBS: Reinforcement Learning-Guided Beam Search for process optimization in a paper machine dryer section

Paper drying is responsible for over two-thirds of energy consumption in the U.S. pulp and paper industry, presenting significant potential for energy savings through optimization of process parameters. Current approaches often assume fixed operating conditions, neglecting dynamic ambient and process variations that limit achievable savings and real-world applicability. To this end, we develop a physics-based simulation environment for a paper machine dryer section and propose a reinforcement learning (RL) framework to minimize overall energy consumption by optimizing drying process parameters under diverse operating conditions. To mitigate overdrying and numerical instabilities caused by suboptimal local RL actions, we introduce Reinforcement Learning-Guided Beam Search (RLGBS), which explores multiple action sequences in parallel using beam search. Instead of making step-by-step decisions, RLGBS prioritizes solutions based on cumulative probability, reducing the impact of individual suboptimal actions. Experiments demonstrate that RLGBS achieves consistent energy savings under unseen operating conditions not encountered during training, outperforming conventional RL methods. While validated in drying optimization, this framework is broadly applicable to other RL-based industrial process control problems.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Constraint treatment techniques and parallel algorithms for multibody dynamic analysis

Computational procedures for kinematic and dynamic analysis of three-dimensional multibody dynamic (MBD) systems are developed from the differential-algebraic equations (DAE's) viewpoint. Constraint violations during the time integration process are minimized and penalty constraint stabilization techniques and partitioning schemes are developed. The governing equations of motion, a two-stage staggered explicit-implicit numerical algorithm, are treated which takes advantage of a partitioned solution procedure. A robust and parallelizable integration algorithm is developed. This algorithm uses a two-stage staggered central difference algorithm to integrate the translational coordinates and the angular velocities. The angular orientations of bodies in MBD systems are then obtained by using an implicit algorithm via the kinematic relationship between Euler parameters and angular velocities. It is shown that the combination of the present solution procedures yields a computationally more accurate solution. To speed up the computational procedures, parallel implementation of the present constraint treatment techniques, the two-stage staggered explicit-implicit numerical algorithm was efficiently carried out. The DAE's and the constraint treatment techniques were transformed into arrowhead matrices to which Schur complement form was derived. By fully exploiting the sparse matrix structural analysis techniques, a parallel preconditioned conjugate gradient numerical algorithm is used to solve the systems equations written in Schur complement form. A software testbed was designed and implemented in both sequential and parallel computers. This testbed was used to demonstrate the robustness and efficiency of the constraint treatment techniques, the accuracy of the two-stage staggered explicit-implicit numerical algorithm, and the speed up of the Schur-complement-based parallel preconditioned conjugate gradient algorithm on a parallel computer.

Chiou, Jin-Chern↗

Parallel digital modem using multirate digital filter banks

A new class of architectures for an all-digital modem is presented in this report. This architecture, referred to as the parallel receiver (PRX), is based on employing multirate digital filter banks (DFB's) to demodulate, track, and detect the received symbol stream. The resulting architecture is derived, and specifications are outlined for designing the DFB for the PRX. The key feature of this approach is a lower processing rate then either the Nyquist rate or the symbol rate, without any degradation in the symbol error rate. Due to the freedom in choosing the processing rate, the designer is able to arbitrarily select and use digital components, independent of the speed of the integrated circuit technology. PRX architecture is particularly suited for high data rate applications, and due to the modular structure of the parallel signal path, expansion to even higher data rates is accommodated with each. Applications of the PRX would include gigabit satellite channels, multiple spacecraft, optical links, interactive cable-TV, telemedicine, code division multiple access (CDMA) communications, and others.

Sadr, Ramin↗

Multidisciplinary systems optimization by linear decomposition

In a typical design process major decisions are made sequentially. An illustrated example is given for an aircraft design in which the aerodynamic shape is usually decided first, then the airframe is sized for strength and so forth. An analogous sequence could be laid out for any other major industrial product, for instance, a ship. The loops in the discipline boxes symbolize iterative design improvements carried out within the confines of a single engineering discipline, or subsystem. The loops spanning several boxes depict multidisciplinary design improvement iterations. Omitted for graphical simplicity is parallelism of the disciplinary subtasks. The parallelism is important in order to develop a broad workfront necessary to shorten the design time. If all the intradisciplinary and interdisciplinary iterations were carried out to convergence, the process could yield a numerically optimal design. However, it usually stops short of that because of time and money limitations. This is especially true for the interdisciplinary iterations.

Sobieski, J.↗

Scalable training of trustworthy and energy-efficient predictive graph foundation models for atomistic materials modeling: a case study with HydraGNN

We present our work on developing and training scalable, trustworthy, and energy-efficient predictive graph foundation models (GFMs) using HydraGNN, a multi-headed graph convolutional neural network architecture. HydraGNN expands the boundaries of graph neural network (GNN) computations in both training scale and data diversity. It abstracts over message passing algorithms, allowing both reproduction of and comparison across algorithmic innovations that define nearest-neighbor convolution in GNNs. This work discusses a series of optimizations that have allowed scaling up the GFMs training to tens of thousands of GPUs on datasets consisting of hundreds of millions of graphs. Our GFMs use multitask learning (MTL) to simultaneously learn graph-level and node-level properties of atomistic structures, such as energy and atomic forces. Using over 154 million atomistic structures for training, we illustrate the performance of our approach along with the lessons learned on two state-of-the-art US Department of Energy (US-DOE) supercomputers, namely the Perlmutter petascale system at the National Energy Research Scientific Computing Center and the Frontier exascale system at Oak Ridge Leadership Computing Facility. The HydraGNN architecture enables the GFM to achieve near-linear strong scaling performance using more than 2000 GPUs on Perlmutter and 16,000 GPUs on Frontier.

97 MATHEMATICS AND COMPUTING↗

A graphite crystal polarimeter for stellar X-ray astronomy.

The first crystal X-ray polarimeter to be used for X-ray astronomy is described. Polarization is measured by modulation of the X rays diffracted at an average 45 deg glancing angle from large, curved graphite crystal panels as these rotate about an axis parallel to the incident X-ray flux. Arrangement of the crystal panels, the design of the detector, and the signal-processing circuitry were optimized to minimize systematic effects produced by off-axis pointing of the rocket and cosmic ray induced events. The in-flight performance of the instrument in relation to the observed background signal is discussed.

Weisskopf, M. C.↗

Rain scavenging of solid rocket exhaust clouds

An explicit model for cloud microphysics was developed for application to the problem of co-condensation/vaporization of HCl and H2O in the presence of Al2O3 particulate nuclei. Validity of the explicit model relative to the implicit model, which has been customarily applied to atmospheric cloud studies, was demonstrated by parallel computations of H2O condensation upon (NH4)2 SO4 nuclei. A mesoscale predictive model designed to account for the impact of wet processes on atmospheric dynamics is also under development. Input data specifying the equilibrium state of HC1 and H2O vapors in contact with aqueous HC1 solutions were found to be limited, particularly in respect to temperature range.

Dingle, A. N.↗

Peripheral processors for high-speed simulation

This paper describes some of the results of a study directed to the specification and procurement of a new cockpit simulator for an advanced class of helicopters. A part of the study was the definition of a challenging benchmark problem, and detailed analyses of it were made to assess the suitability of a variety of simulation techniques. The analyses showed that a particularly cost-effective approach to the attainment of adequate speed for this extremely demanding application is to employ a large minicomputer acting as host and controller for a special-purpose digital peripheral processor. Various realizations of such peripheral processors, all employing state-of-the-art electronic circuitry and a high degree of parallelism and pipelining, are available or under development. The types of peripheral processors array processors, simulation-oriented processors, and arrays of processing elements - are analyzed and compared. They are particularly promising approaches which should be suitable for high-speed simulations of all kinds, the cockpit simulator being a case in point.

Karplus, W. J.↗

Commercialization of materials processing in space

The primary motivation of the Materials Processing in Space program is the scientific and commercial utilization of the effects of the unique environments of space on material processes. The reduction or elimination of the pervasive influences of gravity on Earth-based process mechanisms affords opportunities for understanding and improving ground-based processing or producing select materials in space which, typically, would be of low volume, high value commercial interest. Additionally, the unlimited, if not 'hard' vacuum of space affords equally interesting influences on material processes. To evolve the commercialization of Materials Processing in Space, the program seeks to establish and demonstrate the scientific/technological precepts for analyzing and using the space environment and, in parallel, to establish the legal and management mechanisms to implement commercial ventures.

Zoller, L. K.↗

Flow in a differentially rotated cylindrical drop at low Reynolds number

A liquid drop held captive between parallel disks that are differentially rotated is a model for the swirling flows induced by crystal rotation in the floating-zone process for growing semiconductor materials. An asymptotic analysis for a cylindrical drop is presented that elucidates the structure of the axisymmetric cellular motions caused by disk rotation at low Reynolds number. Variations of meniscus shape induced by these flows are described in the limit of small capillary number. Most cellular flow fields break the bifurcation point that corresponds to the Plateau-Rayleigh limit for the length of a static drop into two disjoint shape families and lower the maximum stable drop length. This effect is studied by a singular bifurcation analysis.

Harriott, G. M.↗

Unconditionally stable concurrent procedures for transient finite-element analysis

A family of algorithms was outlined which would appear to be particularly well-suited for implementation in a parallel environment. This is due to the fact that for any partition of the mesh each subdomain in the partition can be processed over a time step simultaneously and independently of the rest. The method eliminates the need for assembling and factorizing large global arrays while retaining the unconditional stability properties of the algorithms used at the local level. To critically appraise the proposed methodology, two limiting cases were considered: element-by-element mesh partitions, and coarse mesh partitions. It was concluded that while the proposed methodology can be useful in sequential machines, it would appear to be promising as it bears on computation. It should also be emphasized that extensions of the method to nonlinear problems are possible.

Ortiz, Michael↗

Analysis of fault-tolerant neurocontrol architectures

The fault-tolerance of analog parallel distributed implementations of a multivariable aircraft neurocontroller is analyzed by simulating weight and neuron failures in a simplified scheme of analog processing based on the functional architecture of the ETANN chip (Electrically Trainable Artificial Neural Network). The neural information processing is found to be only partially distributed throughout the set of weights of the neurocontroller synthesized with the backpropagation algorithm. Although the degree of distribution of the neural processing, and consequently the fault-tolerance of the neurocontroller, could be enhanced using Locally Distributed Weight and Neuron Approaches, a satisfactory level of fault-tolerance could only be obtained by retraining the degrated VLSI neurocontroller. The possibility of maintaining neurocontrol performance and stability in the presence of single weight of neuron failures was demonstrated through an automated retraining procedure of the neurocontroller based on a pre-programmed choice and sequence of the training parameters.

Troudet, T.↗

Research Studies on Advanced Optical Module/Head Designs for Optical Data Storage

Preprints are presented from the recent 1992 Optical Data Storage meeting in San Jose. The papers are divided into the following topical areas: Magneto-optical media (Modeling/design and fabrication/characterization/testing); Optical heads (holographic optical elements); and Optical heads (integrated optics). Some representative titles are as follow: Diffraction analysis and evaluation of several focus and track error detection schemes for magneto-optical disk systems; Proposal for massively parallel data storage system; Transfer function characteristics of super resolving systems; Modeling and measurement of a micro-optic beam deflector; Oxidation processes in magneto-optic and related materials; and A modal analysis of lamellar diffraction gratings in conical mountings.

Source record↗

Periodic Application of Concurrent Error Detection in Processor Array Architectures

Processor arrays can provide an attractive architecture for some applications. Featuring modularity, regular interconnection and high parallelism, such arrays are well-suited for VLSI/WSI implementations, and applications with high computational requirements, such as real-time signal processing. Preserving the integrity of results can be of paramount importance for certain applications. In these cases, fault tolerance should be used to ensure reliable delivery of a system's service. One aspect of fault tolerance is the detection of errors caused by faults. Concurrent error detection (CED) techniques offer the advantage that transient and intermittent faults may be detected with greater probability than with off-line diagnostic tests. Applying time-redundant CED techniques can reduce hardware redundancy costs. However, most time-redundant CED techniques degrade a system's performance.

Chen, Paul Peichuan↗

High-performance passive microwave survey on Josephson Junctions

The quasi-optical generations of images of objects with their internal structure in millimeter (MM) and submillimeter (SMM) bands is one of prime problems of modern radioelectronics. The main advantage of passive MM imaging systems in comparison with visible and infrared (IR) systems is small attenuation of signals in fog, cloud, smoke, dust and other obscurants. However, at a panoramic scanning of space the observation time lengthens and thereby the information processing rate becomes restricted so that single-channel system cannot image in real time. Therefore we must use many radiometers in parallel to reduce the observation time. Such system must contain receiving sensors as pixels in multibeam antenna. The use of Josephson Junctions (JJ) for this purpose together with the cryoelectronic devices like GaAs FET (field effect transistors) or SQUIDS for signal amplifications after JJ is of particular interest in this case.

Denisov, A. G.↗