Search NASA⌕ Search

SEARCH · Search NASA

Results for “Parallel Programming”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,009 records · Page 56

Expression of Novel Gene Products Upregulated by Disuse is Normalized by an Osteogenic Mechanical Stimulus: Evidence for the Molecular Basis of a Low Level Biomechanical Countermeasure for Osteoporosis?

The National Research Council's report entitled: A Strategy for Space Biology and Medical Science, highlighted several areas of fundamental scientific investigation which must be addressed to make long-term space exploration not only feasible, but safe. This "Goldberg Strategy," as well as several subsequent reports published by the NRC's Space Studies Board (e.g., Assessment of Programs in Space Biology and Medicine, Smith et. al., 1991), suggests that the principal hurdle to man's extended presence in space is the osteopenia which parallels reduced gravity. Ironically, the most significant risk to the skeleton may only be realized on return to normal gravitational fields, and full recovery of bone mass may never occur. Effective counter-measures to this microgravity induced bone loss are thus essential. Considering the similarities of space and aging induced osteopenia, an indisputable benefit of such a prophylaxis would be its potential as a treatment for the bone loss which plagues over 25 million people in the U.S. The osteogenic potential of mechanical strain is strongly frequency dependent, with sensitivity increasing up through at least 60 Hz (cycles per second). One hundred seconds per day of a 1 Hz cyclic loading will inhibit disuse osteopenia only if sufficient in magnitude to engender 1000 microstrain (mu(epsilon)) in the tissue. When loading is applied at 30 Hz, however, mechanical strains on the order of 5O mu(epsilon) (approx. 1% of the peak strains which occur in bone during vigorous functional activity), can stimulate bone formation in a duration dependent manner. In longer term animal studies, strains of less than 10 mu(epsilon), induced non-invasively via a whole body vibration, will stimulate bone formation on the surfaces of trabeculae, increase bone density, and improve strength. Finally, preliminary results from a double blind prospective clinical trial shows promise in inhibiting the bone loss which parallels the menopause. Based on these observations, we propose that these high frequency, low magnitude, mechanical strains effectively serve as a "surrogate" for musculoskeletal ground reaction forces, and thus represent an ideal countermeasure to the osteopenia which parallels microgravity conditions. The specific goal of this NASA funded work is to identify genes in bone upregulated by disuse, and to determine the efficacy of an osteogenic mechanical stimulus to downregulate their expression.

Rubin, C.↗

Compute Server Performance Results

Parallel-vector supercomputers have been the workhorses of high performance computing. As expectations of future computing needs have risen faster than projected vector supercomputer performance, much work has been done investigating the feasibility of using Massively Parallel Processor systems as supercomputers. An even more recent development is the availability of high performance workstations which have the potential, when clustered together, to replace parallel-vector systems. We present a systematic comparison of floating point performance and price-performance for various compute server systems. A suite of highly vectorized programs was run on systems including traditional vector systems such as the Cray C90, and RISC workstations such as the IBM RS/6000 590 and the SGI R8000. The C90 system delivers 460 million floating point operations per second (FLOPS), the highest single processor rate of any vendor. However, if the price-performance ration (PPR) is considered to be most important, then the IBM and SGI processors are superior to the C90 processors. Even without code tuning, the IBM and SGI PPR's of 260 and 220 FLOPS per dollar exceed the C90 PPR of 160 FLOPS per dollar when running our highly vectorized suite,

Stockdale, I. E.↗

Performance of Nanotube-Based Ceramic Composites: Modeling and Experiment

The excellent mechanical properties of carbon-nanotubes are driving research into the creation of new strong, tough nanocomposite systems. In this program, our initial work presented the first evidence of toughening mechanisms operating in carbon-nanotube- reinforced ceramic composites using a highly-ordered array of parallel multiwall carbon-nanotubes (CNTs) in an alumina matrix. Nanoindentation introduced controlled cracks and the damage was examined by SEM. These nanocomposites exhibit the three hallmarks of toughening in micron-scale fiber composites: crack deflection at the CNT/matrix interface; crack bridging by CNTs; and CNT pullout on the fracture surfaces. Furthermore, for certain geometries a new mechanism of nanotube collapse in shear bands was found, suggesting that these materials can have multiaxial damage tolerance. The quantitative indentation data and computational models were used to determine the multiwall CNT axial Young's modulus as 200-570 GPa, depending on the nanotube geometry and quality.

Curtin, W. A.↗

Eta Carinae - A Demanding Mistress

Over the past 15 years, a number of observers and modelers have increasingly focused on this massive system that is approaching its end stage, a supernova? a hypernova? When? The discovery by Augusto Damineli that Eta Carinae had a 5.5-year period proved timely as the newly-installed STIS was primed to observe its properties in the visible and ultraviolet. Initial observations occurred on January 1998, and through multiple programs, including the multi-cycle Hubble Treasury program, have sampled changes across two cycles. Now a multi-cycle program, focused on mapping variations in the extended wind-wind collision zones through early 2015, will test 3-D models of the interacting winds. In parallel, studies have been accomplished in X-rays with RXTE and CHANDRA, now in the far infrared with Herschel and from the ground with VLT. Each new observation is helping to peel back the veil of mystery on this massive binary system, but also opening up more questions to be answered. Timely inclusion of laboratory studies and models have greatly enhanced the observational results. We will summarize the latest results including submitted papers and very recent results with Herschel.

Gull, Theodore R.↗

Global Swath and Gridded Data Tiling

This software generates cylindrically projected tiles of swath-based or gridded satellite data for the purpose of dynamically generating high-resolution global images covering various time periods, scaling ranges, and colors called "tiles." It reconstructs a global image given a set of tiles covering a particular time range, scaling values, and a color table. The program is configurable in terms of tile size, spatial resolution, format of input data, location of input data (local or distributed), number of processes run in parallel, and data conditioning.

Thompson, Charles K.↗

Implicit Coupling Approach for Simulation of Charring Carbon Ablators

This study demonstrates that coupling of a material thermal response code and a flow solver with nonequilibrium gas/surface interaction for simulation of charring carbon ablators can be performed using an implicit approach. The material thermal response code used in this study is the three-dimensional version of Fully Implicit Ablation and Thermal response program, which predicts charring material thermal response and shape change on hypersonic space vehicles. The flow code solves the reacting Navier-Stokes equations using Data Parallel Line Relaxation method. Coupling between the material response and flow codes is performed by solving the surface mass balance in flow solver and the surface energy balance in material response code. Thus, the material surface recession is predicted in flow code, and the surface temperature and pyrolysis gas injection rate are computed in material response code. It is demonstrated that the time-lagged explicit approach is sufficient for simulations at low surface heating conditions, in which the surface ablation rate is not a strong function of the surface temperature. At elevated surface heating conditions, the implicit approach has to be taken, because the carbon ablation rate becomes a stiff function of the surface temperature, and thus the explicit approach appears to be inappropriate resulting in severe numerical oscillations of predicted surface temperature. Implicit coupling for simulation of arc-jet models is performed, and the predictions are compared with measured data. Implicit coupling for trajectory based simulation of Stardust fore-body heat shield is also conducted. The predicted stagnation point total recession is compared with that predicted using the chemical equilibrium surface assumption

Ablation↗

Cable Tester Box

Cables are very important electrical devices that carry power and signals across multiple instruments. Any fault in a cable can easily result in a catastrophic outcome. Therefore, verifying that all cables are built to spec is a very important part of Electrical Integration Procedures. Currently, there are two methods used in lab for verifying cable connectivity. (1) Using a Break-Out Box and an ohmmeter this method is time-consuming but effective for custom cables and (2) Commercial Automated Cable Tester Boxes this method is fast, but to test custom cables often requires pre-programmed configuration files, and cables used on spacecraft are often uniquely designed for specific purposes. The idea is to develop a semi-automatic continuity tester that reduces human effort in cable testing, speeds up the electrical integration process, and ensures system safety. The JPL-Cable Tester Box is developed to check every single possible electrical connection in a cable in parallel. This system indicates connectivity through LED (light emitting diode) circuits. Users can choose to test any pin/shell (test node) with a single push of a button, and any other nodes that are shorted to the test node, even if they are in the same connector, will light up with the test node. The JPL-Cable Tester Boxes offers the following advantages: 1. Easy to use: The architecture is simple enough that it only takes 5 minutes for anyone to learn how operate the Cable Tester Box. No pre-programming and calibration are required, since this box only checks continuity. 2. Fast: The cable tester box checks all the possible electrical connections in parallel at a push of a button. If a cable normally takes half an hour to test, using the Cable Tester Box will improve the speed to as little as 60 seconds to complete. 3. Versatile: Multiple cable tester boxes can be used together. As long as all the boxes share the same electrical potential, any number of connectors can be tested together.

Lee, Jason H.↗

DSN Beowulf Cluster-Based VLBI Correlator

The NASA Deep Space Network (DSN) requires a broadband VLBI (very long baseline interferometry) correlator to process data routinely taken as part of the VLBI source Catalogue Maintenance and Enhancement task (CAT M&E) and the Time and Earth Motion Precision Observations task (TEMPO). The data provided by these measurements are a crucial ingredient in the formation of precision deep-space navigation models. In addition, a VLBI correlator is needed to provide support for other VLBI related activities for both internal and external customers. The JPL VLBI Correlator (JVC) was designed, developed, and delivered to the DSN as a successor to the legacy Block II Correlator. The JVC is a full-capability VLBI correlator that uses software processes running on multiple computers to cross-correlate two-antenna broadband noise data. Components of this new system (see Figure 1) consist of Linux PCs integrated into a Beowulf Cluster, an existing Mark5 data storage system, a RAID array, an existing software correlator package (SoftC) originally developed for Delta DOR Navigation processing, and various custom- developed software processes and scripts. Parallel processing on the JVC is achieved by assigning slave nodes of the Beowulf cluster to process separate scans in parallel until all scans have been processed. Due to the single stream sequential playback of the Mark5 data, some ramp-up time is required before all nodes can have access to required scan data. Core functions of each processing step are accomplished using optimized C programs. The coordination and execution of these programs across the cluster is accomplished using Pearl scripts, PostgreSQL commands, and a handful of miscellaneous system utilities. Mark5 data modules are loaded on Mark5 Data systems playback units, one per station. Data processing is started when the operator scans the Mark5 systems and runs a script that reads various configuration files and then creates an experiment-dependent status database used to delegate parallel tasks between nodes and storage areas (see Figure 2). This script forks into three processes: extract, translate, and correlate. Each of these processes iterates on available scan data and updates the status database as the work for each scan is completed. The extract process coordinates and monitors the transfer of data from each of the Mark5s to the Beowulf RAID storage systems. The translate process monitors and executes the data conversion processes on available scan files, and writes the translated files to the slave nodes. The correlate process monitors the execution of SoftC correlation processes on the slave nodes for scans that have completed translation. A comparison of the JVC and the legacy Block II correlator outputs reveals they are well within a formal error, and that the data are comparable with respect to their use in flight navigation. The processing speed of the JVC is improved over the Block II correlator by a factor of 4, largely due to the elimination of the reel-to-reel tape drives used in the Block II correlator.

Rogstad, Stephen P.↗

CG-Kit: Code Generation Toolkit for performant and maintainable variants of source code applied to Flash-X hydrodynamics simulations

CG-Kit is a new Code Generation tool-Kit that we have developed as a part of the solution for portability and maintainability for multiphysics computing applications. The development of CG-Kit is rooted in the urgent need created by the shifting landscape of high-performance computing platforms and the algorithmic complexities of a particular large-scale multiphysics application: Flash-X. To efficiently use computing resources on a heterogeneous node, an application must have a map of computation to resources and a mechanism to move the data and computation to the resources according to the map. Most existing performance portability solutions are focussed on abstracting the expression of computations so that a unified source code can be specialized to run on different resources. However, such an approach is insufficient for a code like Flash-X, which has a multitude of code components that can be assembled in various permutations and combinations to form different instances of applications. Similar challenges apply to any code that has composability, where a single specified way of apportioning work among devices may not be optimal. Additionally, use cases arise where the optimal control flow of computation may differ for different devices while the underlying numerics remain identical. This combination leads to unique challenges including handling an existing large code base in Fortran and/or C/C++, subdivision of code into a great variety of units supporting a wide range of physics and numerical methods, different parallelization techniques for distributed and shared memory systems and accelerator devices, and heterogeneity of computing platforms requiring coexisting variants of parallel algorithms. All of these challenges demand that scientific software developers apply existing knowledge about domain applications, algorithms, and computing platforms to determine custom abstractions and granularity for code generation. There is a critical lack of tools to tackle those problems. CG-Kit is designed to fill this gap by providing a user with the ability to express their desired control flow and computation-to-resource map in the form a pseudocode-like recipe. It consists of standalone tools that can be combined into highly specific and, we argue, highly effective portability and maintainability toolchains. Here we present the design of our new tools: parametrized source trees, control flow graphs, and recipes. The tools are implemented in Python. They are agnostic to the programming language of the source code targeted for code generation. In conclusion, we demonstrate the capabilities of the toolkit with two examples, first, multithreaded variants of the basic AXPY operation, and second, variants of parallel algorithms within a hydrodynamics solver, called Spark, from Flash-X that operates on block-structured adaptive meshes.

Algorithmic portability↗

Low-Temperature Plasma-Based Metrology of Lithium-Ion Battery Electrode Materials (CRADA Final Report)

As part of the Cyclotron Road program, SirenOpt Inc. evaluated its low-temperature plasma-based metrology sensor prototype for measuring multiple critical properties of lithium-ion battery electrode materials in parallel and in real-time. Cost-effective, minimal-waste manufacturing of high-performance battery electrode materials will be vital for achieving society’s net-zero carbon emission goals. Because existing electrode metrology sensors cannot operate within most sections of manufacturing lines, manufacturers often complete hundreds of processing steps before they can test their products and detect problems. When manufacturers perform these offline tests, they typically only test a small portion of the manufactured products. Current electrode manufacturing thus often yields many low-quality products, or off-spec products that must be thrown away all together. For example, at least 6% of the total lithium-ion battery manufacturing cost (i.e., over $250 million/year for the average gigafactory) is devoted to processing defective electrodes that are not scrapped until performance tests are failed during late-stage quality control checks. Electrode variability also leads manufacturers to build extra cells into battery packs to reduce the risk of poor performance. For example, many electric vehicle (EV) manufacturers include up to 10% more cells than needed, which substantially increases the cost and weight of the final EV product. The SirenOpt sensor can potentially enable early detection of poorly manufactured electrodes and allow them to be removed earlier from manufacturing lines, which can save battery manufacturers (hundreds of) millions of dollars per year. The sensor can further be used to improve product quality by accelerating R&D and process optimization, improving quality control, and enabling real-time process control. Overall, a real-time, in-situ metrology strategy can create unprecedented opportunities for implementation of smart manufacturing practices and advanced quality and process control solutions to realize higher battery electrode throughput and performance.

25 ENERGY STORAGE↗

Study of solid rocket motor for a space shuttle booster

The study of solid rocket motors for a space shuttle booster was directed toward definition of a parallel-burn shuttle booster using two 156-in.-dia solid rocket motors. The study effort was organized into the following major task areas: system studies, preliminary design, program planning, and program costing.

Source record↗

NASA management of the Space Shuttle Program

The management system and management technology described have been developed to meet stringent cost and schedule constraints of the Space Shuttle Program. Management of resources available to this program requires control and motivation of a large number of efficient creative personnel trained in various technical specialties. This must be done while keeping track of numerous parallel, yet interdependent activities involving different functions, organizations, and products all moving together in accordance with intricate plans for budgets, schedules, performance, and interaction. Some techniques developed to identify problems at an early stage and seek immediate solutions are examined.

Peters, F.↗

Fluid dynamics applications of the Illiac IV computer

The Illiac IV is a parallel-structure computer with computing power an order of magnitude greater than that of conventional computers. It can be used for experimental tasks in fluid dynamics which can be simulated more economically, for simulating flows that cannot be studied by experiment, and for combining computer and experimental simulations. The architecture of Illiac IV is described, and the use of its parallel operation is demonstrated on the example of its solution of the one-dimensional wave equation. For fluid dynamics problems, a special FORTRAN-like vector programming language was devised, called CFD language. Two applications are described in detail: (1) the determination of the flowfield around the space shuttle, and (2) the computation of transonic turbulent separated flow past a thick biconvex airfoil.

Maccormack, R. W.↗

Improved piston ring materials for 650 deg C service

A program to develop piston ring material systems which will operate at 650C was performed. In this program, two candidate high temperature piston ring substrate materials, Carpenter 709-2 and 440B, were hot formed into the piston ring shape and subsequently evaluated. In a parallel development effort ceramic and metallic piston ring coating materials were applied to cast iron rings by various processing techniques and then subjected to thermal shock and wear evaluation. Finally, promising candidate coatings were applied to the most thermally stable hot formed substrate. The results of evaluation tests of the hot formed substrate show that Carpenter 709-2 has greater thermal stability than 440B. Of the candidate coatings, plasma transferred arc (PTA) applied tungsten carbide and molybdenum based systems exhibit the greatest resistance to thermal shock. For the ceramic based systems, thermal shock resistance was improved by bond coat grading. Wear testing was conducted to 650C (1202F). For ceramic systems, the alumina/titania/zirconia/yttria composition showed highest wear resistance. For the PTA applied systems, the tungsten carbide based system showed highest wear resistance.

Bjorndahl, W. D.↗

Fast Fourier Transform algorithm design and tradeoffs

The Fast Fourier Transform (FFT) is a mainstay of certain numerical techniques for solving fluid dynamics problems. The Connection Machine CM-2 is the target for an investigation into the design of multidimensional Single Instruction Stream/Multiple Data (SIMD) parallel FFT algorithms for high performance. Critical algorithm design issues are discussed, necessary machine performance measurements are identified and made, and the performance of the developed FFT programs are measured. Fast Fourier Transform programs are compared to the currently best Cray-2 FFT program.

Kamin, Ray A., III↗

Parallelization of an Object-Oriented Unstructured Aeroacoustics Solver

A computational aeroacoustics code based on the discontinuous Galerkin method is ported to several parallel platforms using MPI. The discontinuous Galerkin method is a compact high-order method that retains its accuracy and robustness on non-smooth unstructured meshes. In its semi-discrete form, the discontinuous Galerkin method can be combined with explicit time marching methods making it well suited to time accurate computations. The compact nature of the discontinuous Galerkin method also makes it well suited for distributed memory parallel platforms. The original serial code was written using an object-oriented approach and was previously optimized for cache-based machines. The port to parallel platforms was achieved simply by treating partition boundaries as a type of boundary condition. Code modifications were minimal because boundary conditions were abstractions in the original program. Scalability results are presented for the SCI Origin, IBM SP2, and clusters of SGI and Sun workstations. Slightly superlinear speedup is achieved on a fixed-size problem on the Origin, due to cache effects.

Baggag, Abdelkader↗

Constraints and Opportunities in GCM Model Development

Over the past 30 years climate models have evolved from relatively simple representations of a few atmospheric processes to complex multi-disciplinary system models which incorporate physics from bottom of the ocean to the mesopause and are used for seasonal to multi-million year timescales. Computer infrastructure over that period has gone from punchcard mainframes to modern parallel clusters. Constraints of working within an ever evolving research code mean that most software changes must be incremental so as not to disrupt scientific throughput. Unfortunately, programming methodologies have generally not kept pace with these challenges, and existing implementations now present a heavy and growing burden on further model development as well as limiting flexibility and reliability. Opportunely, advances in software engineering from other disciplines (e.g. the commercial software industry) as well as new generations of powerful development tools can be incorporated by the model developers to incrementally and systematically improve underlying implementations and reverse the long term trend of increasing development overhead. However, these methodologies cannot be applied blindly, but rather must be carefully tailored to the unique characteristics of scientific software development. We will discuss the need for close integration of software engineers and climate scientists to find the optimal processes for climate modeling.

Schmidt, Gavin↗

Applications of Automation Methods for Nonlinear Fracture Test Analysis

Using automated and standardized computer tools to calculate the pertinent test result values has several advantages such as: 1. allowing high-fidelity solutions to complex nonlinear phenomena that would be impractical to express in written equation form, 2. eliminating errors associated with the interpretation and programing of analysis procedures from the text of test standards, 3. lessening the need for expertise in the areas of solid mechanics, fracture mechanics, numerical methods, and/or finite element modeling, to achieve sound results, 4. and providing one computer tool and/or one set of solutions for all users for a more "standardized" answer. In summary, this approach allows a non-expert with rudimentary training to get the best practical solution based on the latest understanding with minimum difficulty.Other existing ASTM standards that cover complicated phenomena use standard computer programs: 1. ASTM C1340/C1340M-10- Standard Practice for Estimation of Heat Gain or Loss Through Ceilings Under Attics Containing Radiant Barriers by Use of a Computer Program 2. ASTM F 2815 - Standard Practice for Chemical Permeation through Protective Clothing Materials: Testing Data Analysis by Use of a Computer Program 3. ASTM E2807 - Standard Specification for 3D Imaging Data Exchange, Version 1.0 The verification, validation, and round-robin processes required of a computer tool closely parallel the methods that are used to ensure the solution validity for equations included in test standard. The use of automated analysis tools allows the creation and practical implementation of advanced fracture mechanics test standards that capture the physics of a nonlinear fracture mechanics problem without adding undue burden or expense to the user. The presented approach forms a bridge between the equation-based fracture testing standards of today and the next generation of standards solving complex problems through analysis automation.

Allen, Phillip A.↗