Search NASA⌕ Search

SEARCH · Search NASA

Results for “Parallel in time”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30

Parallelized Quadrupole Simulations of Thermographic Responses of Composites

Thermography has been shown to be a viable technique for inspection of composites. Model inversion of the thermography data requires a fast method for performing the forward problem. Viable numerical methods for the thermal response forward problem are finite element, finite difference and the quadrupole method. Normally both the finite element and finite difference methods solve for the thermal response in the time domain which limits one’s ability to increase the speed of the simulation by parallelization. In contrast, the quadrupole method solves for the Laplace transform of the thermal response. One of the features of the Laplace transform methodology is the solution at any discrete time is independent of the solution at all other times. Therefore, it is easy to separate into a set of independent calculations with each of the times of interest being performed in parallel. Additionally, the numeric inversion of the Laplace transform typically involves numerically solving for the Laplace transform at multiple Laplace frequencies. Each of those solutions are also independent of solutions at other frequencies and can be calculated in parallel. By parallelization of this method, it is possible to perform the simulations of three-dimensional configurations in seconds. When the input stimulus for thermal response is a delta function heat flux (a reasonable approximation for flash heating), the thermal response is smooth. For this case, it is possible to accurately estimate the thermal response at any time within a given time interval from a set of simulations separated by exponentially increasing time steps. From these simulations, it is possible to accurately interpolate to find the response at intermediate times by a spline interpolation of the logarithm of time versus logarithm of temperature. The thermal response with exponential time stepping is shown to produce values for the thermal response which are within 1% of values within the time interval. The simulations are compared to finite element simulations of the same inspection configurations. The simulations are also compared to the thermographic measurements on composites where shape and depth of the delaminations are obtained from other inspection methods.

Thermography↗

Electrochemical leaching of spent LIBs: Kinetics, novel reactor, and modeling

The use of electrons as main reagent for the recovery and recycling of critical metals from spent lithium-ion batteries (LIBs) is a process electrification strategy that can be used to close the life-cycle loop of LIBs through more sustainable methods. Electrochemical leaching, a process that uses a reductant that is constantly regenerated electrochemically for the leaching of lithium-ion battery black mass (LIBBM), has shown high extraction efficiencies and sustainable scores. However, slow kinetics, reactor design challenges and lack of deeper understanding of the underlying processes are barriers to the optimization, scale-up, and market adoption of this technology. In this paper, a kinetic study and mathematical model for dissolving LIBBM is presented to better understand the underlying mechanisms aiming to reduce the processing time and make predictions for future design and scale-up. The effect of acid and electrochemically mediated reductant concentrations, LIBBM loading, and cathode/reactor designs were explored. As a result, the leaching time was reduced from 7h to under 1h at a pulp density of 73 g/L, without external heating. A novel reactor with parallel baffle electrodes (PBE) was developed, which significantly reduced the leaching time by improving convection in a stirred slurry electrochemical reactor. Dimensionless numbers were deduced from an unsteady state model, which can be used in dimensional analysis for future process design and scale-up.

25 ENERGY STORAGE↗

Distributed Parallel Processing and Dynamic Load Balancing Techniques for Multidisciplinary High Speed Aircraft Design

Multidisciplinary design optimization (MDO) for large-scale engineering problems poses many challenges (e.g., the design of an efficient concurrent paradigm for global optimization based on disciplinary analyses, expensive computations over vast data sets, etc.) This work focuses on the application of distributed schemes for massively parallel architectures to MDO problems, as a tool for reducing computation time and solving larger problems. The specific problem considered here is configuration optimization of a high speed civil transport (HSCT), and the efficient parallelization of the embedded paradigm for reasonable design space identification. Two distributed dynamic load balancing techniques (random polling and global round robin with message combining) and two necessary termination detection schemes (global task count and token passing) were implemented and evaluated in terms of effectiveness and scalability to large problem sizes and a thousand processors. The effect of certain parameters on execution time was also inspected. Empirical results demonstrated stable performance and effectiveness for all schemes, and the parametric study showed that the selected algorithmic parameters have a negligible effect on performance.

Krasteva, Denitza T.↗

Evidence of Multiple Reconnection Lines at the Magnetopause from Cusp Observations

Recent global hybrid simulations investigated the formation of flux transfer events (FTEs) and their convection and interaction with the cusp. Based on these simulations, we have analyzed several Polar cusp crossings in the Northern Hemisphere to search for the signature of such FTEs in the energy distribution of downward precipitating ions: precipitating ion beams at different energies parallel to the ambient magnetic field and overlapping in time. Overlapping ion distributions in the cusp are usually attributed to a combination of variable ion acceleration during the magnetopause crossing together with the time-of-flight effect from the entry point to the observing satellite. Most "step up" ion cusp structures (steps in the ion energy dispersions) only overlap for the populations with large pitch angles and not for the parallel streaming populations. Such cusp structures are the signatures predicted by the pulsed reconnection model, where the reconnection rate at the magnetopause decreased to zero, physically separating convecting flux tubes and their parallel streaming ions. However, several Polar cusp events discussed in this study also show an energy overlap for parallel-streaming precipitating ions. This condition might be caused by reopening an already reconnected field line, forming a magnetic island (flux rope) at the magnetopause similar to that reported in global MHD and Hybrid simulations

flux transfer events↗

Assignment Of Finite Elements To Parallel Processors

Elements assigned approximately optimally to subdomains. Mapping algorithm based on simulated-annealing concept used to minimize approximate time required to perform finite-element computation on hypercube computer or other network of parallel data processors. Mapping algorithm needed when shape of domain complicated or otherwise not obvious what allocation of elements to subdomains minimizes cost of computation.

Salama, Moktar A.↗

Unsteady turbomachinery flow simulations on massively parallel architectures

The accurate numerical simulation of unsteady, three-dimensional viscous flow in turbomachines is computationally very intensive, requiring prohibitively large amounts of computer time on current vector supercomputers. In recent years, computer systems based on massively parallel architectures have been developed that offer the promise of meeting the computational power requirements of such large-scale simulations. However, a rethinking of existing algorithms and methodology is required in order to fully harness the computational power of such architectures. In this paper the capabilities of the Connection Machine (CM-2) in predicting unsteady flows in turbomachines are evaluated. The implementation on the CM-2 of an implicit, time-accurate, zonal algorithm for the Navier-Stokes equations in two dimensions is described. Programming issues and modifications made to the original algorithm (developed for vector, pipelined supercomputers) in order to improve performance on the CM-2 are outlined. Algorithm performance is evaluated and compared with a functionally equivalent code for the CRAY-YMP.

Madavan, N. K.↗

CFD Research, Parallel Computation and Aerodynamic Optimization

During the last five years, CFD has matured substantially. Pure CFD research remains to be done, but much of the focus has shifted to integration of CFD into the design process. The work under these cooperative agreements reflects this trend. The recent work, and work which is planned, is designed to enhance the competitiveness of the US aerospace industry. CFD and optimization approaches are being developed and tested, so that the industry can better choose which methods to adopt in their design processes. The range of computer architectures has been dramatically broadened, as the assumption that only huge vector supercomputers could be useful has faded. Today, researchers and industry can trade off time, cost, and availability, choosing vector supercomputers, scalable parallel architectures, networked workstations, or heterogenous combinations of these to complete required computations efficiently.

Ryan, James S.↗

High performance flight simulation at NASA Langley

The use of real-time simulation at the NASA facility is reviewed specifically with regard to hardware, software, and the use of a fiberoptic-based digital simulation network. The network hardware includes supercomputers that support 32- and 64-bit scalar, vector, and parallel processing technologies. The software include drivers, real-time supervisors, and routines for site-configuration management and scheduling. Performance specifications include: (1) benchmark solution at 165 sec for a single CPU; (2) a transfer rate of 24 million bits/s; and (3) time-critical system responsiveness of less than 35 msec. Simulation applications include the Differential Maneuvering Simulator, Transport Systems Research Vehicle simulations, and the Visual Motion Simulator. NASA is shown to be in the final stages of developing a high-performance computing system for the real-time simulation of complex high-performance aircraft.

Cleveland, Jeff I., II↗

Relative molecular orientation can impact the onset of plasticity in molecular crystals

Abstract Creating or moving dislocations is the first step to dissipating mechanical energy via plastic deformation under contact loading. In molecular crystals there is both a lattice that defines crystal orientation and a relative orientation of the basis of the molecules. We define a normalization parameter which relates strain at yield, the hardness of the bulk crystal, and a distance parameter analogous to a Burgers vector that nominally predicts the relative ease of initiating plasticity in this broad class of materials. Analyzing the yield behavior of 10 different molecular crystals of varying space groups shows the inter-molecular orientation predicts the experimentally observed applied stress needed to nucleate dislocations. When molecules are oriented ‘parallel’ relative to one another the normalized maximum shear stress at the onset of plasticity is on the order of 3–5 times lower than when molecules within the crystal are ‘anti-parallel’, and molecules with a more equiaxed shape fall in between these bounds. This provides an initial indication of a structural feature which predicts the relative ease of initiating plasticity during contact loading in molecular crystals.

36 MATERIALS SCIENCE↗

Virtual Time III, Part 3: Throttling and Message Cancellation

This is Part 3 of a trio of papers that unify in a natural way the two historically distinct parallel discrete event synchronization paradigms, optimistic and conservative, combining the best properties of both into a single framework called Unified Virtual Time (UVT). In this part, we survey the synchronization effects that can be achieved by restricting to corner cases the relationships permitted among the control variables, GVT, CVT, TVT, and LVT, which were defined in Part 1. Here we also survey various throttling policies from the literature and describe how they can be implemented in UVT by controlling the value of TVT, including policies that can take advantage of rollback in addition to LP blocking. A significant result is a new category of efficient and higher precision throttling algorithms for optimistic execution that are based on optimistic lookahead, defined in a way that is symmetric to what we now call the conservative lookahead information that is traditionally used for conservative synchronization. Finally, we present a novel algorithm allowing the choice between lazy and aggressive cancellation to be made on a message-by-message basis using either external logic expressed in the model code, or policy code internal to the simulator, or a mixture of both.

throttling↗

Image sensor with high dynamic range linear output

Designs and operational methods to increase the dynamic range of image sensors and APS devices in particular by achieving more than one integration times for each pixel thereof. An APS system with more than one column-parallel signal chains for readout are described for maintaining a high frame rate in readout. Each active pixel is sampled for multiple times during a single frame readout, thus resulting in multiple integration times. The operation methods can also be used to obtain multiple integration times for each pixel with an APS design having a single column-parallel signal chain for readout. Furthermore, analog-to-digital conversion of high speed and high resolution can be implemented.

Yadid-Pecht, Orly↗

Parallelized reliability estimation of reconfigurable computer networks

A parallelized system, ASSURE, for computing the reliability of embedded avionics flight control systems which are able to reconfigure themselves in the event of failure is described. ASSURE accepts a grammar that describes a reliability semi-Markov state-space. From this it creates a parallel program that simultaneously generates and analyzes the state-space, placing upper and lower bounds on the probability of system failure. ASSURE is implemented on a 32-node Intel iPSC/860, and has achieved high processor efficiencies on real problems. Through a combination of improved algorithms, exploitation of parallelism, and use of an advanced microprocessor architecture, ASSURE has reduced the execution time on substantial problems by a factor of one thousand over previous workstation implementations. Furthermore, ASSURE's parallel execution rate on the iPSC/860 is an order of magnitude faster than its serial execution rate on a Cray-2 supercomputer. While dynamic load balancing is necessary for ASSURE's good performance, it is needed only infrequently; the particular method of load balancing used does not substantially affect performance.

Nicol, David M.↗

Massive parallelism in the future of science

Massive parallelism appears in three domains of action of concern to scientists, where it produces collective action that is not possible from any individual agent's behavior. In the domain of data parallelism, computers comprising very large numbers of processing agents, one for each data item in the result will be designed. These agents collectively can solve problems thousands of times faster than current supercomputers. In the domain of distributed parallelism, computations comprising large numbers of resource attached to the world network will be designed. The network will support computations far beyond the power of any one machine. In the domain of people parallelism collaborations among large groups of scientists around the world who participate in projects that endure well past the sojourns of individuals within them will be designed. Computing and telecommunications technology will support the large, long projects that will characterize big science by the turn of the century. Scientists must become masters in these three domains during the coming decade.

Denning, Peter J.↗

Transputer based control system for MTLRS

The Modular Transportable Laser Ranging Systems (MTLRS-1 and MTLRS-2) have been designed in the early eighties and have been in operation very successfully since 1984. The original design of the electronic control system was based on the philosophy of parallel processing, but these ideas could at that time only be implemented to a very limited extent. This present system utilizes two MOTOROLA 6800 8-bit processors slaved to a HP A-600 micro-computer. These processors support the telescope tracking system and the data-acquisition/formatting, respectively. Nevertheless, the overall design still is largely hardware oriented. Because the system is now some nine years old, aging of components increases the risk of malfunctioning and some components or units are outdated and not available anymore. The control system for MTLRS is now being re-designed completely, based on the original philosophy of parallel processing, making use of contemporary advanced electronics and processor technology. The new design aims at the requirements for Satellite Laser Ranging (SLR) in the nineties, making use of the extensive operational experience obtained with the two transportable systems.

Vermaat, Erik↗

PC-CUBE: A Personal Computer Based Hypercube

PC-CUBE is an ensemble of IBM PCs or close compatibles connected in the hypercube topology with ordinary computer cables. Communication occurs at the rate of 115.2 K-band via the RS-232 serial links. Available for PC-CUBE is the Crystalline Operating System III (CrOS III), Mercury Operating System, CUBIX and PLOTIX which are parallel I/O and graphics libraries. A CrOS performance monitor was developed to facilitate the measurement of communication and computation time of a program and their effects on performance. Also available are CXLISP, a parallel version of the XLISP interpreter; GRAFIX, some graphics routines for the EGA and CGA; and a general execution profiler for determining execution time spent by program subroutines. PC-CUBE provides a programming environment similar to all hypercube systems running CrOS III, Mercury and CUBIX. In addition, every node (personal computer) has its own graphics display monitor and storage devices. These allow data to be displayed or stored at every processor, which has much instructional value and enables easier debugging of applications. Some application programs which are taken from the book Solving Problems on Concurrent Processors (Fox 88) were implemented with graphics enhancement on PC-CUBE. The applications range from solving the Mandelbrot set, Laplace equation, wave equation, long range force interaction, to WaTor, an ecological simulation.

Ho, Alex↗

Aerothermal loads analysis for high speed flow over a quilted surface configuration

Attention is given to hypersonic laminar flow over a quilted surface configuration that simulates an array of Space Shuttle Thermal Protection System panels bowed in a spherical shape as a result of thermal gradient through the panel thickness. Pressure and heating loads to the surface are determined. The flow field over the configuration was mathematically modeled by means of time-dependent, three-dimensional conservation of mass, momentum, and energy equations. A boundary mapping technique was then used to obtain a rectangular, parallel piped computational domain, and an explicit MacCormack (1972) explicit time-split predictor corrector finite difference algorithm was used to obtain steady state solutions. Total integrated heating loads vary linearly with bowed height when this value does not exceed the local boundary layer thickness.

Olsen, G. C.↗

Parallel algorithms and archtectures for computational structural mechanics

The determination of the fundamental (lowest) natural vibration frequencies and associated mode shapes is a key step used to uncover and correct potential failures or problem areas in most complex structures. However, the computation time taken by finite element codes to evaluate these natural frequencies is significant, often the most computationally intensive part of structural analysis calculations. There is continuing need to reduce this computation time. This study addresses this need by developing methods for parallel computation.

Patrick, Merrell↗