Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallel integration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29

DFT algorithms for bit-serial GaAs array processor architectures

Systems and Processes Engineering Corporation (SPEC) has developed an innovative array processor architecture for computing Fourier transforms and other commonly used signal processing algorithms. This architecture is designed to extract the highest possible array performance from state-of-the-art GaAs technology. SPEC's architectural design includes a high performance RISC processor implemented in GaAs, along with a Floating Point Coprocessor and a unique Array Communications Coprocessor, also implemented in GaAs technology. Together, these data processors represent the latest in technology, both from an architectural and implementation viewpoint. SPEC has examined numerous algorithms and parallel processing architectures to determine the optimum array processor architecture. SPEC has developed an array processor architecture with integral communications ability to provide maximum node connectivity. The Array Communications Coprocessor embeds communications operations directly in the core of the processor architecture. A Floating Point Coprocessor architecture has been defined that utilizes Bit-Serial arithmetic units, operating at very high frequency, to perform floating point operations. These Bit-Serial devices reduce the device integration level and complexity to a level compatible with state-of-the-art GaAs device technology.

Mcmillan, Gary B.↗

Group implicit concurrent algorithms in nonlinear structural dynamics

During the 70's and 80's, considerable effort was devoted to developing efficient and reliable time stepping procedures for transient structural analysis. Mathematically, the equations governing this type of problems are generally stiff, i.e., they exhibit a wide spectrum in the linear range. The algorithms best suited to this type of applications are those which accurately integrate the low frequency content of the response without necessitating the resolution of the high frequency modes. This means that the algorithms must be unconditionally stable, which in turn rules out explicit integration. The most exciting possibility in the algorithms development area in recent years has been the advent of parallel computers with multiprocessing capabilities. So, this work is mainly concerned with the development of parallel algorithms in the area of structural dynamics. A primary objective is to devise unconditionally stable and accurate time stepping procedures which lend themselves to an efficient implementation in concurrent machines. Some features of the new computer architecture are summarized. A brief survey of current efforts in the area is presented. A new class of concurrent procedures, or Group Implicit algorithms is introduced and analyzed. The numerical simulation shows that GI algorithms hold considerable promise for application in coarse grain as well as medium grain parallel computers.

Ortiz, M.↗

Data fusion with artificial neural networks (ANN) for classification of earth surface from microwave satellite measurements

A data fusion system with artificial neural networks (ANN) is used for fast and accurate classification of five earth surface conditions and surface changes, based on seven SSMI multichannel microwave satellite measurements. The measurements include brightness temperatures at 19, 22, 37, and 85 GHz at both H and V polarizations (only V at 22 GHz). The seven channel measurements are processed through a convolution computation such that all measurements are located at same grid. Five surface classes including non-scattering surface, precipitation over land, over ocean, snow, and desert are identified from ground-truth observations. The system processes sensory data in three consecutive phases: (1) pre-processing to extract feature vectors and enhance separability among detected classes; (2) preliminary classification of Earth surface patterns using two separate and parallely acting classifiers: back-propagation neural network and binary decision tree classifiers; and (3) data fusion of results from preliminary classifiers to obtain the optimal performance in overall classification. Both the binary decision tree classifier and the fusion processing centers are implemented by neural network architectures. The fusion system configuration is a hierarchical neural network architecture, in which each functional neural net will handle different processing phases in a pipelined fashion. There is a total of around 13,500 samples for this analysis, of which 4 percent are used as the training set and 96 percent as the testing set. After training, this classification system is able to bring up the detection accuracy to 94 percent compared with 88 percent for back-propagation artificial neural networks and 80 percent for binary decision tree classifiers. The neural network data fusion classification is currently under progress to be integrated in an image processing system at NOAA and to be implemented in a prototype of a massively parallel and dynamically reconfigurable Modular Neural Ring (MNR).

Lure, Y. M. Fleming↗

Advanced automation of a prototypic thermal control system for Space Station

Viewgraphs on an advanced automation of a prototypic thermal control system for space station are presented. The Thermal Expert System (TEXSYS) was initiated in 1986 as a cooperative project between ARC and JCS as a way to leverage on-going work at both centers. JSC contributed Thermal Control System (TCS) hardware and control software, TCS operational expertise, and integration expertise. ARC contributed expert system and display expertise. The first years of the project were dedicated to parallel development of expert system tools, displays, interface software, and TCS technology and procedures by a total of four organizations.

Dominick, Jeff↗

Effects of leading and trailing edge flaps on the aerodynamics of airfoil/vortex interactions

A numerical procedure has been developed for predicting the two-dimensional parallel interaction between a free convecting vortex and a NACA 0012 airfoil having leading and trailing edge integral-type flaps. Special emphasis is placed on the unsteady flap motion effects which result in alleviating the interaction at subcritical and supercritical onset flows. The numerical procedure described here is based on the implicit finite-difference solutions to the unsteady two-dimensional full potential equation. Vortex-induced effects are computed using the Biot-Savart Law with allowance for a finite core radius. The vortex-induced velocities at the surface of the airfoil are incorporated into the potential flow model via the use of the velocity transpiration approach. Flap motion effects are also modeled using the transpiration approach. For subcritical interactions, our results indicate that trailing edge flaps can be used to alleviate the impulsive loads experienced by the airfoil. For supercritical interactions, our results demonstrate the necessity of using a leading edge flap, rather than a trailing edge flap, to alleviate the interaction. Results for various time-dependent flap motions and their effect on the predicted temporal sectional loads, differential pressures, and the free vortex trajectories are presented

Hassan, Ahmed A.↗

Design and Performance Analysis of a Massively Parallel Atmospheric General Circulation Model

In the 1990's computer manufacturers are increasingly turning to the development of parallel processor machines to meet the high performance needs of their customers. Simultaneously, atmospheric scientists study weather and climate phenomena ranging from hurricanes to El Nino to global warming that require increasingly fine resolution models. Here, implementation of a parallel atmospheric general circulation model (GCM) which exploits the power of massively parallel machines is described. Using the horizontal data domain decomposition methodology, this FORTRAN 90 model is able to integrate a 0.6 deg. longitude by 0.5 deg. latitude problem at a rate of 19 Gigaflops on 512 processors of a Cray T3E 600; corresponding to 280 seconds of wall-clock time per simulated model day. At this resolution, the model has 64 times as many degrees of freedom and performs 400 times as many floating point operations per simulated day as the model it replaces.

Schaffer, Daniel S.↗

Multiple Differential-Amplifier MMICs Embedded in Waveguides

Compact amplifier assemblies of a type now being developed for operation at frequencies of hundreds of gigahertz comprise multiple amplifier units in parallel arrangements to increase power and/or cascade arrangements to increase gains. Each amplifier unit is a monolithic microwave integrated circuit (MMIC) implementation of a pair of amplifiers in differential (in contradistinction to single-ended) configuration. Heretofore, in cascading amplifiers to increase gain, it has been common practice to interconnect the amplifiers by use of wires and/or thin films on substrates. This practice has not yielded satisfactory results at frequencies greater than 200 Hz, in each case, for either or both of two reasons: Wire bonds introduce large discontinuities. Because the interconnections are typically tens of wavelengths long, any impedance mismatches give rise to ripples in the gain-vs.-frequency response, which degrade the performance of the cascade.

Kangaslahti, Pekka↗

A Discrete Component Low-Noise Preamplifier Readout for a Linear (1x16) SiC Photodiode Array

A compact, low-noise and inexpensive preamplifier circuit has been designed and fabricated to optimally readout a common cathode (1x16) channel 4H-SiC Schottky photodiode array for use in ultraviolet experiments. The readout uses an operational amplifier with 10 pF capacitor in the feedback loop in parallel with a low leakage switch for each of the channels. This circuit configuration allows for reiterative sample, integrate and reset. A sampling technique is given to remove Johnson noise, enabling a femtoampere level readout noise performance. Commercial-off-the-shelf acquisition electronics are used to digitize the preamplifier analogue signals. The data logging acquisition electronics has a different integration circuit, which allows the bandwidth and gain to be independently adjusted. Using this readout, photoresponse measurements across the array between spectral wavelengths 200 nm and 370 nm are made to establish the array pixels external quantum efficiency, current responsivity and noise equivalent power.

low-noise preamplifier↗

Implementing distributed operations : a comparison of two Deep Space Missions

Two very different deep space exploration missions—Mars Exploration Rover and Cassini—have made use of distributed operations for their science teams. In the case of MER, the distributed operations capability was implemented only after the prime mission was completed, as the rovers continued to operate well in excess of their expected mission lifetimes; Cassini, designed for a prime mission of four years, had planned for distributed operations from its inception. The rapid command turnaround timeline of MER, as well as many of the operations features implemented to support it, have proven to be conducive to distributed operations. These features include: a single science team leader during the tactical operations timeline, highly integrated science and engineering teams, processes and file structures designed to permit multiple team members to work in parallel to deliver sequencing products, web-based spacecraft status and planning reports for team-wide access, and near-elimination of paper products from the operations process.

Larsen, Barbara↗

Cholla-MHD: An Exascale-capable Magnetohydrodynamic Extension to the Cholla Astrophysical Simulation Code

Abstract We present an extension of the massively parallel, GPU native, astrophysical hydrodynamics code Cholla to magnetohydrodynamics (MHD). Cholla solves the ideal MHD equations in their Eulerian form on a static Cartesian mesh utilizing the Van Leer + constrained transport integrator, the HLLD Riemann solver, and reconstruction methods at second and third order. Cholla’s MHD module can perform ≈260 million cell updates per GPU-second on an NVIDIA A100 while using the HLLD Riemann solver and second order reconstruction. The inherently parallel nature of GPUs combined with increased memory in new hardware allows Cholla’s MHD module to perform simulations with resolutions ∼500 3 cells on a single high-end GPU (e.g., an NVIDIA A100 with 80 GB of memory). We employ GPU direct Message Passing Interface to attain excellent weak scaling on the exascale supercomputer Frontier, while using 74,088 GPUs and simulating a total grid size of over 7.2 trillion cells. A suite of test problems highlights the accuracy of Cholla’s MHD module and demonstrates that zero magnetic divergence in solutions is maintained to round off error. We also present new testing and CI tools using GoogleTest, GitHub Actions, and Jenkins that have made development more robust and accurate and ensure reliability in the future.

Astronomy & Astrophysics↗

Application of the NASA Multiscale Analysis Tool: Multiscale Integration and Interoperability

The NASA Multiscale Analysis Tool (NASMAT) was developed recently to allow a wide variety of multiscale analysis problems to be effectively and efficiently solved. The architecture of NASMAT was established specifically to enable parallelized, “plug-and-play” functionality to reduce the complexity associated with adding new features to the code in the future and to allow end users to rapidly implement and evaluate user-defined capabilities. Additionally, the tool utilizes recursive data structures and subroutines to allow for an arbitrary number of length scales when performing multiscale analyses of heterogeneous materials. These features permit the rapid integration of user-defined capabilities (e.g., a material model, micromechanics approach, or failure theory) at all stages within a NASMAT calculation while leveraging built-in techniques where needed. Additionally, these features allow NASMAT to both be called from an external program as well as call an external program. This paper specifically focuses on the multiscale integration and interoperability of NASMAT with other analysis techniques through an illustrative, multiscale analysis of a 3D woven polymer matrix composite (PMC).

NASMAT↗

Methods for design and evaluation of integrated hardware-software systems for concurrent computation

Research activities and publications are briefly summarized. The major tasks reviewed are: (1) VAX implementation of the PISCES parallel programming environment; (2) Apollo workstation network implementation of the PISCES environment; (3) FLEX implementation of the PISCES environment; (4) sparse matrix iterative solver in PSICES Fortran; (5) image processing application of PISCES; and (6) a formal model of concurrent computation being developed.

Pratt, T. W.↗

The surface integral approach to radarclinometry

A radar image of the Lake Champlain West quadrangle in the Adirondack Mountains of the U.S. is synthesized and used to test the surface integral approach to radarclinometry. It is shown that the surface integral approach to radarclinometry possesses an inherent instability that can be avoided only if the radar reflectance function possesses a shallow slope over the range of operation and if terrain slopes are bounded to prevent their being either parallel or perpendicular to the Poynting vector of the radar irradiance. It is found that the noise associated with real SAR systems makes this instability worse. It is concluded that the value of the surface integral approach to radarclinometry shows little promise.

Wildey, Robert L.↗

Developing the Science Basis for Understanding Polymer Encapsulant Degradation Mechanisms: DuraMAT 2.0 Final Project Report

Polymeric encapsulants are essential materials in photovoltaic modules, protecting sensitive electronics from the environment while providing mechanical integrity to the multilayered assembly. However, these polymeric materials are susceptible to degradation processes driven by the ingress of environmental species, ultraviolet radiation, thermal stresses, and mechanical loading. In this study, we employ a combined atomistic simulation and accelerated aging experimental approach to study the molecular-scale mechanisms of encapsulant degradation. Classical molecular dynamics simulations quantify the diffusion of environmental and degradation species through the polymer matrix, producing composition-specific diffusion coefficients. Reactive simulations characterize activation energy barriers and reaction rate constants for key chemical pathways. In parallel, thermal-desorption analyses coupled with mass spectrometry monitor the emergence and concentration profiles of degradation products under controlled stressor conditions. By integrating simulation and experiment, we establish quantitative correlations between polymer composition, species diffusivity, and chemical reactivity. We anticipate that these relations and quantitative values could serve as high-fidelity inputs to reaction-diffusion models, enabling physics-informed lifetime predictions and guiding the design of more durable encapsulant materials for solar energy applications.

36 MATERIALS SCIENCE↗

Design of a Low-Light-Level Image Sensor with On-Chip Sigma-Delta Analog-to- Digital Conversion

The design and projected performance of a low-light-level active-pixel-sensor (APS) chip with semi-parallel analog-to-digital (A/D) conversion is presented. The individual elements have been fabricated and tested using MOSIS* 2 micrometer CMOS technology, although the integrated system has not yet been fabricated. The imager consists of a 128 x 128 array of active pixels at a 50 micrometer pitch. Each column of pixels shares a 10-bit A/D converter based on first-order oversampled sigma-delta (Sigma-Delta) modulation. The 10-bit outputs of each converter are multiplexed and read out through a single set of outputs. A semi-parallel architecture is chosen to achieve 30 frames/second operation even at low light levels. The sensor is designed for less than 12 e^- rms noise performance.

Mendis, Sunetra K.↗

GSRP/David Marshall: Fully Automated Cartesian Grid CFD Application for MDO in High Speed Flows

With the renewed interest in Cartesian gridding methodologies for the ease and speed of gridding complex geometries in addition to the simplicity of the control volumes used in the computations, it has become important to investigate ways of extending the existing Cartesian grid solver functionalities. This includes developing methods of modeling the viscous effects in order to utilize Cartesian grids solvers for accurate drag predictions and addressing the issues related to the distributed memory parallelization of Cartesian solvers. This research presents advances in two areas of interest in Cartesian grid solvers, viscous effects modeling and MPI parallelization. The development of viscous effects modeling using solely Cartesian grids has been hampered by the widely varying control volume sizes associated with the mesh refinement and the cut cells associated with the solid surface. This problem is being addressed by using physically based modeling techniques to update the state vectors of the cut cells and removing them from the finite volume integration scheme. This work is performed on a new Cartesian grid solver, NASCART-GT, with modifications to its cut cell functionality. The development of MPI parallelization addresses issues associated with utilizing Cartesian solvers on distributed memory parallel environments. This work is performed on an existing Cartesian grid solver, CART3D, with modifications to its parallelization methodology.

Source record↗

Parallel processors and nonlinear structural dynamics algorithms and software

Techniques are discussed for the implementation and improvement of vectorization and concurrency in nonlinear explicit structural finite element codes. In explicit integration methods, the computation of the element internal force vector consumes the bulk of the computer time. The program can be efficiently vectorized by subdividing the elements into blocks and executing all computations in vector mode. The structuring of elements into blocks also provides a convenient way to implement concurrency by creating tasks which can be assigned to available processors for evaluation. The techniques were implemented in a 3-D nonlinear program with one-point quadrature shell elements. Concurrency and vectorization were first implemented in a single time step version of the program. Techniques were developed to minimize processor idle time and to select the optimal vector length. A comparison of run times between the program executed in scalar, serial mode and the fully vectorized code executed concurrently using eight processors shows speed-ups of over 25. Conjugate gradient methods for solving nonlinear algebraic equations are also readily adapted to a parallel environment. A new technique for improving convergence properties of conjugate gradients in nonlinear problems is developed in conjunction with other techniques such as diagonal scaling. A significant reduction in the number of iterations required for convergence is shown for a statically loaded rigid bar suspended by three equally spaced springs.

Belytschko, Ted↗

Path Planning: Differential Dynamic Programming and Model Predictive Path Integral Control on VTOL Aircraft

This paper explores two optimal control approaches, widely used in robotics, to establish their viability as real-time trajectory planners for vehicle configurations envisioned for the emerging aviation sector of Urban Air Mobility (UAM). Differential Dynamic Programming (DDP) enables planning over highly nonlinear dynamics using second-order approximations along a nominal trajectory, and displays quadratic convergence to a local solution. Model Predictive Path Integral (MPPI) is a stochastic sampling-based algorithm that can optimize for general cost criteria, including potentially highly nonlinear formulations, and supports parallel computation through the use of modern GPU hardware. In this work, DDP and MPPI were implemented using model predictive control (MPC), and the results indicate they are able to successfully transition the aircraft over different flight envelopes and generate trajectories unique to UAM vehicles.

Differential Dynamic Programming↗