Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallel programming”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 955 records · Page 53

On a Simplified Approach to Achieve Parallel Performance and Portability Across CPU and GPU Architectures

This paper presents software advances to easily exploit computer architectures consisting of a multi-core CPU and CPU+GPU to accelerate diverse types of high-performance computing (HPC) applications using a single code implementation. The paper describes and demonstrates the performance of the open-source C++ matrix and array (MATAR) library that uniquely offers: (1) a straightforward syntax for programming productivity, (2) usable data structures for data-oriented programming (DOP) for performance, and (3) a simple interface to the open-source C++ Kokkos library for portability and memory management across CPUs and GPUs. The portability across architectures with a single code implementation is achieved by automatically switching between diverse fine-grained parallelism backends (e.g., CUDA, HIP, OpenMP, pthreads, etc.) at compile time. The MATAR library solves many longstanding challenges associated with easily writing software that can run in parallel on any computer architecture. This work benefits projects seeking to write new C++ codes while also addressing the challenges of quickly making existing Fortran codes performant and portable over modern computer architectures with minimal syntactical changes from Fortran to C++. We demonstrate the feasibility of readily writing new C++ codes and modernizing existing codes with MATAR to be performant, parallel, and portable across diverse computer architectures.

97 MATHEMATICS AND COMPUTING↗

Spur, helical, and spiral bevel transmission life modeling

A computer program, TLIFE, which estimates the life, dynamic capacity, and reliability of aircraft transmissions, is presented. The program enables comparisons of transmission service life at the design stage for optimization. A variety of transmissions may be analyzed including: spur, helical, and spiral bevel reductions as well as series combinations of these reductions. The basic spur and helical reductions include: single mesh, compound, and parallel path plus revert star and planetary gear trains. A variety of straddle and overhung bearing configurations on the gear shafts are possible as is the use of a ring gear for the output. The spiral bevel reductions include single and dual input drives with arbitrary shaft angles. The program is written in FORTRAN 77 and has been executed both in the personal computer DOS environment and on UNIX workstations. The analysis may be performed in either the SI metric or the English inch system of units. The reliability and life analysis is based on the two-parameter Weibull distribution lives of the component gears and bearings. The program output file describes the overall transmission and each constituent transmission, its components, and their locations, capacities, and loads. Primary output is the dynamic capacity and 90-percent reliability and mean lives of the unit transmissions and the overall system which can be used to estimate service overhaul frequency requirements. Two examples are presented to illustrate the information available for single element and series transmissions.

Savage, Michael↗

CARA Devolution ESMO Pilot Program

Information regarding the current effort ongoing with CARA to determine the steps needed to have a mission move away from being supported by CARA, otherwise referred to as Devolving/Devolution. Information regarding the working group established to undertake this effort, documents agreed that needed to be written and plans for a parallel operations period. The communication between ESMO and CARA during the Parallel Operations period will be as the following: – A daily email will be sent from ESMO to CARA describing each HIE that is being actively planned and worked, and the action that ESMO is planning on taking – A weekly email will be sent from ESMO to CARA listing the progress toward achieving the parallel operations success criteria – The CARA team will remain on distribution for ESMO HIE communications and CAM invitations associated with DAMs/RMMs during the parallel operations period – At a minimum, 2 members of the CARA team will be in attendance at all ESMO planning and review discussions – A monthly email will be sent from ESMO to CARA with the HIE metrics • During parallel operations, ESMO will contact NASA Headquarters to report any redDevolution Working Group laid out and completed all steps needed to begin parallel operations • Parallel Ops for ESMO began on 11/26 (CARA in shadow mode only) • Any lessons learned from parallel ops will be folded back into documents/templates • Hope to achieve all success criteria < 6 months • Decision for permanent devolution will be based on completion of parallel ops and other future documentation creation (CARA Standard and Handbook, etc) and reviews (ORR) event which is less than 2 days from TCA, including the proposed remediation, on an as-needed daily basis flow using SpaceTrack to send and receive data

ESMO pilot devolution program↗

Progressive Hedging Decomposition for Solutions of Large-Scale Process Family Design Problems

In previous work, we have introduced a mathematical model for solving a discretized version of the process family design problem. This involves two sets of decision variables. One set selects which unit module designs are included in the process platform out of a candidate set of options; the other set determines which of these unit module designs are assigned to each variant. In this work, we exploit a parallelized Progressive Hedging (PH) algorithm to solve even larger scale design problems. PH is a well-known algorithm traditionally used to solve stochastic programming problems. While our problem is not a two-stage stochastic programming problem, the structure is similar, and it can be directly mapped to the PH approach, which we employ here to solve this deterministic optimization problem. We decompose our problem by process variant. We treat the platform unit module design variables as first-stage and the assignment of unit module designs to variants as second-stage, solving the problem using mpi-sppy. We demonstrate this approach on case studies of CC, water desalination, and refrigeration.

Stinchfield, Georgia↗

Progress in developing ultrathin solar cell blanket technology

A program was conducted to develop technologies for welding interconnects to three types of 50-micron-thick, 2 by 2-cm solar cells. Parallel-gap resistance welding was used for interconnect attachment. Weld schedules were independently developed for each of the three cell types and were coincidentally identical. Six 48-cell modules were assembled with 50-micron (nominal) thick cells, frosted fused-silica covers, silver-plated Invar interconnectors, and four different substrate designs. Three modules (one for each cell type) have single-layer Kapton (50-micron-thick) substrates. The other three modules each have a different substrate (Kapton-Kevlar-Kapton, Kapton-graphite-Kapton, and Kapton-graphite-aluminum honeycomb-graphite). All six modules were subjected to 4112 thermal cycles from -175 to 65 C (corresponding to over 40 years of simulated geosynchronous orbit thermal cycling) and experienced only negligible electrical degradation (1.1 percent average of six 48-cell modules).

Patterson, R. E.↗

Optimization of V-groove radiator configuration

In the design of spacecraft radiators intended to provide cryogenic cooling, it is important to minimize the space occupied by radiator shielding while satisfying the performance requirements. This study develops and tests the first step toward optimizing radiator shield configurations and examines the design of the highly effective V-groove radiator concept. This investigation, which makes use of a special purpose Monte Carlo/ray tracing computer program, directly identifies the two shield configuration which minimizes radiator footprint for a given temperature drop. Furthermore, multiple shield configurations may be analyzed through a combination of analysis and program statistics. The results presented demonstrate the sensitivity of intershield temperature drop to shield angle and/or offset and provide a comparison of angled and parallel shield configurations. Good agreement with experimental data is shown for a multiple shield configuration.

Schember, Helene R.↗

Internal fluid mechanics research on supercomputers for aerospace propulsion systems

The Internal Fluid Mechanics Division of the NASA Lewis Research Center is combining the key elements of computational fluid dynamics, aerothermodynamic experiments, and advanced computational technology to bring internal computational fluid mechanics (ICFM) to a state of practical application for aerospace propulsion systems. The strategies used to achieve this goal are to: (1) pursue an understanding of flow physics, surface heat transfer, and combustion via analysis and fundamental experiments, (2) incorporate improved understanding of these phenomena into verified 3-D CFD codes, and (3) utilize state-of-the-art computational technology to enhance experimental and CFD research. Presented is an overview of the ICFM program in high-speed propulsion, including work in inlets, turbomachinery, and chemical reacting flows. Ongoing efforts to integrate new computer technologies, such as parallel computing and artificial intelligence, into high-speed aeropropulsion research are described.

Miller, Brent A.↗

Numerical solutions of the compressible 3-D boundary-layer equations for aerospace configurations with emphasis on LFC

The application of stability theory in Laminar Flow Control (LFC) research requires that density and velocity profiles be specified throughout the viscous flow field of interest. These profile values must be as numerically accurate as possible and free of any numerically induced oscillations. Guidelines for the present research project are presented: develop an efficient and accurate procedure for solving the 3-D boundary layer equation for aerospace configurations; develop an interface program to couple selected 3-D inviscid programs that span the subsonic to hypersonic Mach number range; and document and release software to the LFC community. The interface program was found to be a dependable approach for developing a user friendly procedure for generating the boundary-layer grid and transforming an inviscid solution from a relatively coarse grid to a sufficiently fine boundary-layer grid. The boundary-layer program was shown to be fourth-order accurate in the direction normal to the wall boundary and second-order accurate in planes parallel to the boundary. The fourth-order accuracy allows accurate calculations with as few as one-fifth the number of grid points required for conventional second-order schemes.

Harris, Julius E.↗

Model atmosphere analysis of selected luminous B stars

The general scientific goal of this program has been to determine whether the atmospheric structure of the B-type stars can be represented by the current generation of plane parallel, line-blanketed, LTE stellar atmosphere models sufficiently well to allow accurate effective temperatures and surface gravities to be deduced. The B stars cover a wide range of temperature and luminosity. For the hottest such stars (with T approximately 30,000 K) the applicability of the models may be compromised by departures from LTE in the stellar atmospheres ('non-LTE effects'). At the highest luminosities (the B 'super giants'), the models may be invalidated by departures from plane parallel geometry. Thus we seek to identify the temperature and luminosity range within which these effects are unimportant and where the models may be relied upon.

Fitzpatrick, Edward L.↗

A Comparative Propulsion System Analysis for the High-Speed Civil Transport

Six of the candidate propulsion systems for the High-Speed Civil Transport are the turbojet, turbine bypass engine, mixed flow turbofan, variable cycle engine, Flade engine, and the inverting flow valve engine. A comparison of these propulsion systems by NASA's Glenn Research Center, paralleling studies within the aircraft industry, is presented. This report describes the Glenn Aeropropulsion Analysis Office's contribution to the High-Speed Research Program's 1993 and 1994 propulsion system selections. A parametric investigation of each propulsion cycle's primary design variables is analytically performed. Performance, weight, and geometric data are calculated for each engine. The resulting engines are then evaluated on two airframer-derived supersonic commercial aircraft for a 5000 nautical mile, Mach 2.4 cruise design mission. The effects of takeoff noise, cruise emissions, and cycle design rules are examined.

Berton, Jeffrey J.↗

The UDF05 Follow-up of the HUDF: I. The Faint-End Slope of the Lyman-Break Galaxy Population at zeta approx. 5

We present the UDF05 project, a HST Large Program of deep ACS (F606W, F775W, F850LP, and NICMOS (Fll0W, Fl60W) imaging of three fields, two of which coincide with the NICP1-4 NICMOS parallel observations of the Hubble Ultra Deep Field (HUDF). In this first paper we use the ACS data for the NICP12 field, as well as the original HUDF ACS data, to measure the UV Luminosity Function (LF) of z approximately 5 Lyman Break Galaxies (LBGs) down to very faint levels. Specifically, based on a V - i, i - z selection criterion, we identify a sample of 101 and 133 candidate z approximately 5 galaxies down to z(sub 850) = 28.5 and 29.25 magnitudes in the NICP12 field and in the HUDF, respectively. Using an extensive set of Monte Carlo simulations we derive corrections for observational biases and selection effects, and construct the rest-frame 1400 Angstroms LBG LF over the range M(sub 1400) = [-22.2, -17.1], i.e. down to approximately 0.04 L(sub *) at z = 5. We show that: (i) Different assumptions for the SED distribution of the LBG population, dust properties and intergalactic absorption result in a 25% variation in the number density of LBGs at z = 5 (ii) Under consistent assumptions for dust properties and intergalactic absorption, the HUDF is about 30% under-dense in z = 5 LBGs relative to the NICP12 field, a variation which is well explained by cosmic variance; (iii) The faint-end slope of the LF is independent of the specific assumptions for the input physical parameters, and has a value of alpha approximately -1.6, similar to the faint-end slope of the LF that has been measured for LBGs at z = 3 and z = 6. Our study therefore supports no variation in the faint-end of the LBG LF over the whole redshift range z = 3 to z = 6. The comparison with theoretical predictions suggests that (a,) the majority of the stars in the z = 5 LBG population are produced with a Top-Heavy IMF in merger-driven starbursts, and that (b) possibly, either the fraction of stellar mass produced in starburst, or the fraction of high mass stars in the bursts is increased towards the bright end of the LF.

Oesch, P. A.↗

Partial Overhaul and Initial Parallel Optimization of KINETICS, a Coupled Dynamics and Chemistry Atmosphere Model

KINETICS is a coupled dynamics and chemistry atmosphere model that is data intensive and computationally demanding. The potential performance gain from using a supercomputer motivates the adaptation from a serial version to a parallelized one. Although the initial parallelization had been done, bottlenecks caused by an abundance of communication calls between processors led to an unfavorable drop in performance. Before starting on the parallel optimization process, a partial overhaul was required because a large emphasis was placed on streamlining the code for user convenience and revising the program to accommodate the new supercomputers at Caltech and JPL. After the first round of optimizations, the partial runtime was reduced by a factor of 23; however, performance gains are dependent on the size of the data, the number of processors requested, and the computer used.

KINETICS↗

Performance Enhancement of a Computational Persistent Homology Package

In recent years, persistent homology has become an attractive method for data analysis. It captures topological features, such as connected components, holes, voids, etc., from a point cloud by finding out when these features appear and disappear in the filtration sequence. In this project, we focus on improving the performance of Eirene, a fancy computational persistent homology package. Eirene is a 5000-line opensource software implemented by using the dynamic programming language Julia. We use the Julia profiling tools to identify the performance bottlenecks and develop different methods to manage the bottlenecks, including the parallelization of some time-consuming functions on the multicore/manycore hardware. The empirical results show that the performance can be greatly improved.

Profiling↗

Assessment of Cislunar Staging Orbits to Support the Artemis III Lunar Surface Mission

Since NASA’s selection of an L2 9:2 lunar synodic resonant Near Rectilinear Halo Orbit (NRHO) as the baseline for the Gateway Program, the agency has worked to mature its understanding of this orbit and its use for the Artemis III, IV, and V missions. In parallel with these efforts, NASA has investigated alternative staging orbits to perform the Artemis III lunar surface landing mission and compared those options to the baseline NRHO. This paper evaluates a number of alternative orbits on their feasibility and favorability and compares them to the agency baseline NRHO.

Artemis↗

Substructure analysis using NICE/SPAR and applications of force to linear and nonlinear structures

Parallel computing studies are presented for a variety of structural analysis problems. Included are the substructure planar analysis of rectangular panels with and without a hole, the static analysis of space mast, using NICE/SPAR and FORCE, and substructure analysis of plane rigid-jointed frames using FORCE. The computations are carried out on the Flex/32 MultiComputer using one to eighteen processors. The NICE/SPAR runstream samples are documented for the panel problem. For the substructure analysis of plane frames, a computer program is developed to demonstrate the effectiveness of a substructuring technique when FORCE is enforced. Ongoing research activities for an elasto-plastic stability analysis problem using FORCE, and stability analysis of the focus problem using NICE/SPAR are briefly summarized. Speedup curves for the panel, the mast, and the frame problems provide a basic understanding of the effectiveness of parallel computing procedures utilized or developed, within the domain of the parameters considered. Although the speedup curves obtained exhibit various levels of computational efficiency, they clearly demonstrate the excellent promise which parallel computing holds for the structural analysis problem. Source code is given for the elasto-plastic stability problem and the FORCE program.

Razzaq, Zia↗

Computer-Aided Parallelizer and Optimizer

The Computer-Aided Parallelizer and Optimizer (CAPO) automates the insertion of compiler directives (see figure) to facilitate parallel processing on Shared Memory Parallel (SMP) machines. While CAPO currently is integrated seamlessly into CAPTools (developed at the University of Greenwich, now marketed as ParaWise), CAPO was independently developed at Ames Research Center as one of the components for the Legacy Code Modernization (LCM) project. The current version takes serial FORTRAN programs, performs interprocedural data dependence analysis, and generates OpenMP directives. Due to the widely supported OpenMP standard, the generated OpenMP codes have the potential to run on a wide range of SMP machines. CAPO relies on accurate interprocedural data dependence information currently provided by CAPTools. Compiler directives are generated through identification of parallel loops in the outermost level, construction of parallel regions around parallel loops and optimization of parallel regions, and insertion of directives with automatic identification of private, reduction, induction, and shared variables. Attempts also have been made to identify potential pipeline parallelism (implemented with point-to-point synchronization). Although directives are generated automatically, user interaction with the tool is still important for producing good parallel codes. A comprehensive graphical user interface is included for users to interact with the parallelization process.

Jin, Haoqiang↗

USRA/RIACS

The Research Institute for Advanced Computer Science (RIACS) was established by the Universities Space Research Association (USRA) at the NASA Ames Research Center (ARC) on June 6, 1983. RIACS is privately operated by USRA, a consortium of universities with research programs in the aerospace sciences, under a cooperative agreement with NASA. The primary mission of RIACS is to provide research and expertise in computer science and scientific computing to support the scientific missions of NASA ARC. The research carried out at RIACS must change its emphasis from year to year in response to NASA ARC's changing needs and technological opportunities. A flexible scientific staff is provided through a university faculty visitor program, a post doctoral program, and a student visitor program. Not only does this provide appropriate expertise but it also introduces scientists outside of NASA to NASA problems. A small group of core RIACS staff provides continuity and interacts with an ARC technical monitor and scientific advisory group to determine the RIACS mission. RIACS activities are reviewed and monitored by a USRA advisory council and ARC technical monitor. Research at RIACS is currently being done in the following areas: (1) parallel computing; (2) advanced methods for scientific computing; (3) learning systems; (4) high performance networks and technology; and (5) graphics, visualization, and virtual environments. In the past year, parallel compiler techniques and adaptive numerical methods for flows in complicated geometries were identified as important problems to investigate for ARC's involvement in the Computational Grand Challenges of the next decade. We concluded a summer student visitors program during this six months. We had six visiting graduate students that worked on projects over the summer and presented seminars on their work at the conclusion of their visits. RIACS technical reports are usually preprints of manuscripts that have been submitted to research journals or conference proceedings. A list of these reports for the period July 1, 1992 through December 31, 1992 is provided.

Oliger, Joseph↗

Parallel runway requirement analysis study. Volume 1: The analysis

The correlation of increased flight delays with the level of aviation activity is well recognized. A main contributor to these flight delays has been the capacity of airports. Though new airport and runway construction would significantly increase airport capacity, few programs of this type are currently underway, let alone planned, because of the high cost associated with such endeavors. Therefore, it is necessary to achieve the most efficient and cost effective use of existing fixed airport resources through better planning and control of traffic flows. In fact, during the past few years the FAA has initiated such an airport capacity program designed to provide additional capacity at existing airports. Some of the improvements that that program has generated thus far have been based on new Air Traffic Control procedures, terminal automation, additional Instrument Landing Systems, improved controller display aids, and improved utilization of multiple runways/Instrument Meteorological Conditions (IMC) approach procedures. A useful element to understanding potential operational capacity enhancements at high demand airports has been the development and use of an analysis tool called The PLAND_BLUNDER (PLB) Simulation Model. The objective for building this simulation was to develop a parametric model that could be used for analysis in determining the minimum safety level of parallel runway operations for various parameters representing the airplane, navigation, surveillance, and ATC system performance. This simulation is useful as: a quick and economical evaluation of existing environments that are experiencing IMC delays, an efficient way to study and validate proposed procedure modifications, an aid in evaluating requirements for new airports or new runways in old airports, a simple, parametric investigation of a wide range of issues and approaches, an ability to tradeoff air and ground technology and procedures contributions, and a way of considering probable blunder mechanisms and range of blunder scenarios. This study describes the steps of building the simulation and considers the input parameters, assumptions and limitations, and available outputs. Validation results and sensitivity analysis are addressed as well as outlining some IMC and Visual Meteorological Conditions (VMC) approaches to parallel runways. Also, present and future applicable technologies (e.g., Digital Autoland Systems, Traffic Collision and Avoidance System II, Enhanced Situational Awareness System, Global Positioning Systems for Landing, etc.) are assessed and recommendations made.

Ebrahimi, Yaghoob S.↗