Search NASA⌕ Search

SEARCH · Search NASA

Results for “Parallel Performance Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Performance Modeling and Measurement of Parallelized Code for Distributed Shared Memory Multiprocessors

This paper presents a model to evaluate the performance and overhead of parallelizing sequential code using compiler directives for multiprocessing on distributed shared memory (DSM) systems. With increasing popularity of shared address space architectures, it is essential to understand their performance impact on programs that benefit from shared memory multiprocessing. We present a simple model to characterize the performance of programs that are parallelized using compiler directives for shared memory multiprocessing. We parallelized the sequential implementation of NAS benchmarks using native Fortran77 compiler directives for an Origin2000, which is a DSM system based on a cache-coherent Non Uniform Memory Access (ccNUMA) architecture. We report measurement based performance of these parallelized benchmarks from four perspectives: efficacy of parallelization process; scalability; parallelization overhead; and comparison with hand-parallelized and -optimized version of the same benchmarks. Our results indicate that sequential programs can conveniently be parallelized for DSM systems using compiler directives but realizing performance gains as predicted by the performance model depends primarily on minimizing architecture-specific data locality overhead.

Waheed, Abdul↗

Single- and Multiple-Objective Optimization with Differential Evolution and Neural Networks

Genetic and evolutionary algorithms have been applied to solve numerous problems in engineering design where they have been used primarily as optimization procedures. These methods have an advantage over conventional gradient-based search procedures became they are capable of finding global optima of multi-modal functions and searching design spaces with disjoint feasible regions. They are also robust in the presence of noisy data. Another desirable feature of these methods is that they can efficiently use distributed and parallel computing resources since multiple function evaluations (flow simulations in aerodynamics design) can be performed simultaneously and independently on ultiple processors. For these reasons genetic and evolutionary algorithms are being used more frequently in design optimization. Examples include airfoil and wing design and compressor and turbine airfoil design. They are also finding increasing use in multiple-objective and multidisciplinary optimization. This lecture will focus on an evolutionary method that is a relatively new member to the general class of evolutionary methods called differential evolution (DE). This method is easy to use and program and it requires relatively few user-specified constants. These constants are easily determined for a wide class of problems. Fine-tuning the constants will off course yield the solution to the optimization problem at hand more rapidly. DE can be efficiently implemented on parallel computers and can be used for continuous, discrete and mixed discrete/continuous optimization problems. It does not require the objective function to be continuous and is noise tolerant. DE and applications to single and multiple-objective optimization will be included in the presentation and lecture notes. A method for aerodynamic design optimization that is based on neural networks will also be included as a part of this lecture. The method offers advantages over traditional optimization methods. It is more flexible than other methods in dealing with design in the context of both steady and unsteady flows, partial and complete data sets, combined experimental and numerical data, inclusion of various constraints and rules of thumb, and other issues that characterize the aerodynamic design process. Neural networks provide a natural framework within which a succession of numerical solutions of increasing fidelity, incorporating more realistic flow physics, can be represented and utilized for optimization. Neural networks also offer an excellent framework for multiple-objective and multi-disciplinary design optimization. Simulation tools from various disciplines can be integrated within this framework and rapid trade-off studies involving one or many disciplines can be performed. The prospect of combining neural network based optimization methods and evolutionary algorithms to obtain a hybrid method with the best properties of both methods will be included in this presentation. Achieving solution diversity and accurate convergence to the exact Pareto front in multiple objective optimization usually requires a significant computational effort with evolutionary algorithms. In this lecture we will also explore the possibility of using neural networks to obtain estimates of the Pareto optimal front using non-dominated solutions generated by DE as training data. Neural network estimators have the potential advantage of reducing the number of function evaluations required to obtain solution accuracy and diversity, thus reducing cost to design.

Rai, Man Mohan↗

VAN-DAMME: GPU-accelerated and symmetry-assisted quantum optimal control of multi-qubit systems

We present an open-source software package, VAN-DAMME (Versatile Approaches to Numerically Design, Accelerate, and Manipulate Magnetic Excitations), for massively-parallelized quantum optimal control (QOC) calculations of multi-qubit systems. To enable large QOC calculations, the VAN-DAMME software package utilizes symmetry-based techniques with custom GPU-enhanced algorithms. This combined approach allows for the simultaneous computation of hundreds of matrix exponential propagators that efficiently leverage the intra-GPU parallelism found in high-performance GPUs. In addition, to maximize the computational efficiency of the VAN-DAMME code, we carried out several extensive tests on data layout, computational complexity, memory requirements, and performance. These extensive analyses allowed us to develop computationally efficient approaches for evaluating complex-valued matrix exponential propagators based on Padé approximants. To assess the computational performance of our GPU-accelerated VAN-DAMME code, we carried out QOC calculations of systems containing 10 - 15 qubits, which showed that our GPU implementation is 18.4× faster than the corresponding CPU implementation. Our GPU-accelerated enhancements allow efficient calculations of multi-qubit systems, which can be used for the efficient implementation of QOC applications across multiple domains.

97 MATHEMATICS AND COMPUTING↗

Comprehensive Design Reliability Activities for Aerospace Propulsion Systems

This technical publication describes the methodology, model, software tool, input data, and analysis result that support aerospace design reliability studies. The focus of these activities is on propulsion systems mechanical design reliability. The goal of these activities is to support design from a reliability perspective. Paralleling performance analyses in schedule and method, this requires the proper use of metrics in a validated reliability model useful for design, sensitivity, and trade studies. Design reliability analysis in this view is one of several critical design functions. A design reliability method is detailed and two example analyses are provided-one qualitative and the other quantitative. The use of aerospace and commercial data sources for quantification is discussed and sources listed. A tool that was developed to support both types of analyses is presented. Finally, special topics discussed include the development of design criteria, issues of reliability quantification, quality control, and reliability verification.

Christenson, R. L.↗

Nozomi Cis-Lunar Phase Orbit Determination

Japan's Institute of Space and Astronautical Science (ISAS) launched Nozomi, its first mission to the planet Mars using the newly developed M-V launch vehicle on July 3, 1998. Scientific objectives of the mission are to study the structure and dynamics of the Martian upper atmosphere and its interaction with the solar wind. Nozomi is a cooperative mission between ISAS and the National Aeronautics and Space Administration (NASA). The NASA contribution includes navigation and tracking services provided by the Jet Propulsion Laboratory (JPL). The spacecraft also serves as an engineering demonstration of basic technology for planetary exploration. One of the new technologies was a unique trajectory, developed by ISAS, which used solar gravitational perturbations at the weak stability boundary as an aid to achieve an Earth-Mars transfer orbit. This trajectory saves approximately 120 m/s of Delta V compared to direct hyperbolic insertion and is considered an enabling technology for the mission. Nozomi was the first spacecraft to employ this trajectory and provided on-orbit validation of the technique. The trajectory was achieved by initially placing the spacecraft in a highly elliptical cis-lunar phasing orbit. Six maneuvers were performed during this period to correct injection errors and target an outbound lunar swingby in September 1998. The gravity assist from the lunar swingby raised apogee to the vicinity of the weak stability boundary. After three more targeting maneuvers, Nozomi performed an inbound lunar swingby followed immediately by a powered Earth swingby in late December 1998. A 420 m/s Trans Mars Insertion (TMI) burn at the final Earth periapsis was intended to place the spacecraft on a heliocentric trajectory leading to Mars orbit insertion in October 1999. Orbit determination for Nozomi is performed in parallel by both ISAS and the Multi-Mission Navigation (MMNAV) group at JPL. This was an advantage for the mission because each group would generate solutions based on data collected from their respective tracking networks. Spacecraft events, such as sequence uplinks and maneuvers, were generally scheduled during passes at the Usuda tracking station in Japan. As a result, maneuver design and reconstruction was derived from MMNAV solutions based on JPL tracking data obtained immediately prior to or following maneuvers. Data was also exchanged between ISAS and MMNAV so orbit determination could be performed on joint data sets in support of critical targeting late in the cis-lunar phase. In this paper, information regarding the MMNAV orbit determination effort for the first six months of the mission is presented. The spacecraft trajectory is characterized first, followed by a discussion of the orbit determination estimation procedure and models. Results from selected orbit solutions are presented and compared against reconstructed trajectories. One area of emphasis in this paper is orbit determination in the vicinity of the weak stability boundary. Precise navigation was necessary to target the second lunar swingby and the powered Earth swingby. Delivery accuracy of 150 m was required for these critical encounters, but a number of factors contributed to the general degradation of orbit determination accuracy. This included the fact that the spacecraft was at apogee, at a range of 1.7 million km and moving at less than I km/sec perpendicular to the line of sight. Nozomi was also close to zero degrees declination where there are known limitations on orbit determination performance. Finally, S-band tracking data was acquired through the Nozomi backup low gain antenna. This antenna is offset from the axis of this spin stabilized spacecraft and superimposed large signatures in the Doppler and range data. These difficulties were overcome by combining long data arcs, spanning several maneuvers, with a high fidelity solar pressure model. The model included a physically accurate representation of the spacecraft structure and a high time resolution orientation model. Observation modeling included the removal of the spin induced Doppler bias, spin signature and per pass correction of range calibration errors applied for data leading up to critical events. As a result, all orbit determination goals were met. A second area of emphasis in this paper is the JPL tracking and orbit determination effort in support of the TMI maneuver. TMI occurred out of contact with ground stations and the JPL Goldstone tracking complex had the first pass following the bum. As a result, MMNAV had the responsibility to make a rapid assessment of the maneuver performance. MMNAV made the determination that a 100 m/s under bum had occurred and promptly informed ISAS via voice lines. ISAS immediately began preparations for a correction maneuver (TMIc), which had to be performed during the next Usuda pass. The near real time assessment by MMNAV provided accurate antenna frequency and pointing updates for the spacecraft acquisition at Usuda and the close coordination between the two agencies enabled the design and successful execution of the TMc maneuver. Propellant consumption during the correction burn dictated that the mission be redesigned. ISAS developed a new plan which adds 3 full solar orbits, two Earth swingbys and one lunar swingby with arrival at Mars in January 2004. The final Mars orbit will still enable the mission to achieve all of its science objectives.

Ryne, Mark↗

ChatMPI: LLM-Driven MPI Code Generation for HPC Workloads

The Message Passing Interface (MPI) standard plays a crucial role in enabling scientific applications for parallel computing and is an essential component in high-performance computing (HPC). However, implementing MPI code manually—especially applying a proper domain decomposition and communication pattern—is a challenging and error-prone task. We present ChatMPI, an AI assistant for MPI parallelization of sequential C codes. In our analysis, we focus on testing six essential HPC workloads, which are based on Basic Linear Algebra Subprograms levels 1, 2, and 3 as well as sparse, stencil, and iterative operations. We analyze the process of creating ChatMPI by using the ChatHPC library. This lightweight large language model (LLM)–based infrastructure enables HPC experts to efficiently create and supervise trustworthy AI capabilities for critical HPC software tasks. We study the data required for training (fine-tuning) ChatMPI to generate parallel codes that not only use MPI syntax correctly but also apply HPC techniques to reduce memory communication and maximize performance by using proper work decomposition. With a relatively small training dataset composed of a few dozen prompts and fewer than 15 minutes of fine-tuning on one node equipped with two NVIDIA H100 GPUs, ChatMPI elevates trustworthiness for MPI code generation of current LLMs (e.g., Code Llama, ChatGPT-4o and ChatGPT 5). Additionally, we evaluate the performance of the MPI codes generated by ChatMPI in comparison with the ones generated by ChatGPT-4o and ChatGPT-5. The codes generated by ChatMPI provide up to a 4 × boost in performance by using better problem decomposition, communication patterns, and HPC techniques (e.g., communication avoiding).

Valero Lara, Pedro [ORNL] (ORCID:0000000214794310)↗

Software Engineering Support of the Third Round of Scientific Grand Challenge Investigations: An Earth Modeling System Software Framework Strawman Design that Integrates Cactus and UCLA/UCB Distributed Data Broker

One of the most significant challenges in large-scale climate modeling, as well as in high-performance computing in other scientific fields, is that of effectively integrating many software models from multiple contributors. A software framework facilitates the integration task. both in the development and runtime stages of the simulation. Effective software frameworks reduce the programming burden for the investigators, freeing them to focus more on the science and less on the parallel communication implementation, while maintaining high performance across numerous supercomputer and workstation architectures. This document proposes a strawman framework design for the climate community based on the integration of Cactus, from the relativistic physics community, and UCLA/UCB Distributed Data Broker (DDB) from the climate community. This design is the result of an extensive survey of climate models and frameworks in the climate community as well as frameworks from many other scientific communities. The design addresses fundamental development and runtime needs using Cactus, a framework with interfaces for FORTRAN and C-based languages, and high-performance model communication needs using DDB. This document also specifically explores object-oriented design issues in the context of climate modeling as well as climate modeling issues in terms of object-oriented design.

Talbot, Bryan↗

The Use of Linear Feature Detection to Investigate Thematic Mapper Data Performance and Processing

Geometric and radiometric characteristics of Thematic Mapper data are investigated through analysis of linear features in the data. A linear feature is defined as two close, parallel and opposite edges. Examples in remotely sensed data are such features as rivers and roads. The geometric and radiometric precision TM data is sufficient to allow accurate measurement of linear feature widths. Results also confirm a 28.5m ground IFOV as specified prior to launch. The increase dimensionality of the TM data as compared with MSS data allows the possibility of independent verification of results by using data from several bands.

Gurney, C. M.↗

A dual-processor multi-frequency implementation of the FINDS algorithm

This report presents a parallel processing implementation of the FINDS (Fault Inferring Nonlinear Detection System) algorithm on a dual processor configured target flight computer. First, a filter initialization scheme is presented which allows the no-fail filter (NFF) states to be initialized using the first iteration of the flight data. A modified failure isolation strategy, compatible with the new failure detection strategy reported earlier, is discussed and the performance of the new FDI algorithm is analyzed using flight recorded data from the NASA ATOPS B-737 aircraft in a Microwave Landing System (MLS) environment. The results show that low level MLS, IMU, and IAS sensor failures are detected and isolated instantaneously, while accelerometer and rate gyro failures continue to take comparatively longer to detect and isolate. The parallel implementation is accomplished by partitioning the FINDS algorithm into two parts: one based on the translational dynamics and the other based on the rotational kinematics. Finally, a multi-rate implementation of the algorithm is presented yielding significantly low execution times with acceptable estimation and FDI performance.

Godiwala, Pankaj M.↗

Direct I/O for RNTuple Columnar Data

RNTuple is the new columnar data format designed as the successor to ROOT’s TTree format. It allows to make use of modern hardware capabilities and is expected to be used in production by the LHC experiments during the HL-LHC. In this paper, we discuss the usage of Direct I/O to fully exploit modern SSDs, especially in the context of the recent addition of parallel RNTuple writing. We describe the alignment requirements imposed by Direct I/O and approaches to meet them for columnar data formats. Finally, we discuss performance results for both writing and reading, in synthetic benchmarks as well as real-world applications.

Hahnfeld, Jonas [CERN; Goethe U., Frankfurt (main)↗

PETSc/TAO Users Manual Revision 3.22

This manual describes the use of the Portable, Extensible Toolkit for Scientific Computation (PETSc) and the Toolkit for Advanced Optimization (TAO) for the numerical solution of partial differential equations (PDEs) and related problems on high-performance computers. PETSc/TAO is a suite of data structures and routines that provide the building blocks for implementing large-scale application codes on parallel (and serial) computers. PETSc uses the MPI standard for all distributed memory communication.

97 MATHEMATICS AND COMPUTING↗

PETSc/TAO Users Manual Revision 3.23

This manual describes the use of the Portable, Extensible Toolkit for Scientific Computation (PETSc) and the Toolkit for Advanced Optimization (TAO) for the numerical solution of partial differential equations (PDEs) and related problems on high-performance computers. PETSc/TAO is a suite of data structures and routines that provide the building blocks for implementing large-scale application codes on parallel (and serial) computers. PETSc uses the MPI standard for all distributed memory communication.

97 MATHEMATICS AND COMPUTING↗

PETSc/TAO Users Manual Revision 3.24

This manual describes the use of the Portable, Extensible Toolkit for Scientific Computation (PETSc) and the Toolkit for Advanced Optimization (TAO) for the numerical solution of partial differential equations (PDEs) and related problems on high-performance computers. PETSc/TAO is a suite of data structures and routines that provide the building blocks for implementing large-scale application codes on parallel (and serial) computers. PETSc uses the MPI standard for all distributed memory communication.

97 MATHEMATICS AND COMPUTING↗

PETSc/TAO Users Manual Revision 3.25

This manual describes the use of the Portable, Extensible Toolkit for Scientific Computation (PETSc) and the Toolkit for Advanced Optimization (TAO) for the numerical solution of partial differential equations (PDEs) and related problems on high-performance computers. PETSc/TAO is a suite of data structures and routines that provide the building blocks for implementing large-scale application codes on parallel (and serial) computers. PETSc uses the MPI standard for all distributed memory communication.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Dual throat thruster cold flow analysis

The concept was evaluated with cold flow (nitrogen gas) testing and through analysis for application as a tripropellant engine for single-stage-to-orbit type missions. Three modes of operation were tested and analyzed: (1) Mode 1 Series Burn, (2) Mode 1 Parallel Burn, and (3) Mode 2. Primary emphasis was placed on the Mode 2 plume attachment aerodynamics and performance. The conclusions from the test data analysis are as follows: (1) the concept is aerodynamically feasible, (2) the performance loss is as low as 0.5 percent, (3) the loss is minimized by an optimum nozzle spacing corresponding to an AF-ATS ratio of about 1.5 or an Le/Rtp ratio of 3.0 for the dual throat hardware tested, requiring only 4% bleed flow, (4) the Mode 1 and Mode 2 geometry requirements are compatible and pose no significant design problems.

Lundgreen, R. B.↗

Intelligent systems technology infrastructure for integrated systems

Significant advances have occurred during the last decade in intelligent systems technologies (a.k.a. knowledge-based systems, KBS) including research, feasibility demonstrations, and technology implementations in operational environments. Evaluation and simulation data obtained to date in real-time operational environments suggest that cost-effective utilization of intelligent systems technologies can be realized for Automated Rendezvous and Capture applications. The successful implementation of these technologies involve a complex system infrastructure integrating the requirements of transportation, vehicle checkout and health management, and communication systems without compromise to systems reliability and performance. The resources that must be invoked to accomplish these tasks include remote ground operations and control, built-in system fault management and control, and intelligent robotics. To ensure long-term evolution and integration of new validated technologies over the lifetime of the vehicle, system interfaces must also be addressed and integrated into the overall system interface requirements. An approach for defining and evaluating the system infrastructures including the testbed currently being used to support the on-going evaluations for the evolutionary Space Station Freedom Data Management System is presented and discussed. Intelligent system technologies discussed include artificial intelligence (real-time replanning and scheduling), high performance computational elements (parallel processors, photonic processors, and neural networks), real-time fault management and control, and system software development tools for rapid prototyping capabilities.

Lum, Henry, Jr.↗

Optimization of Microelectronic Devices for Sensor Applications

The NASA/JPL goal to reduce payload in future space missions while increasing mission capability demands miniaturization of active and passive sensors, analytical instruments and communication systems among others. Currently, typical system requirements include the detection of particular spectral lines, associated data processing, and communication of the acquired data to other systems. Advances in lithography and deposition methods result in more advanced devices for space application, while the sub-micron resolution currently available opens a vast design space. Though an experimental exploration of this widening design space-searching for optimized performance by repeated fabrication efforts-is unfeasible, it does motivate the development of reliable software design tools. These tools necessitate models based on fundamental physics and mathematics of the device to accurately model effects such as diffraction and scattering in opto-electronic devices, or bandstructure and scattering in heterostructure devices. The software tools must have convenient turn-around times and interfaces that allow effective usage. The first issue is addressed by the application of high-performance computers and the second by the development of graphical user interfaces driven by properly developed data structures. These tools can then be integrated into an optimization environment, and with the available memory capacity and computational speed of high performance parallel platforms, simulation of optimized components can proceed. In this paper, specific applications of the electromagnetic modeling of infrared filtering, as well as heterostructure device design will be presented using genetic algorithm global optimization methods.

Cwik, Tom↗

Simple Models of the Spatial Distribution of Cloud Radiative Properties for Remote Sensing Studies

This project aimed to assess the degree to which estimates of three-dimensional cloud structure can be inferred from a time series of profiles obtained at a point. The work was motivated by the desire to understand the extent to which high-frequency profiles of the atmosphere (e.g. ARM data streams) can be used to assess the magnitude of non-plane parallel transfer of radiation in thc atmosphere. We accomplished this by performing an observing system simulation using a large-eddy simulation and a Monte Carlo radiative transfer model. We define the 3D effect as the part of the radiative transfer that isn't captured by one-dimensional radiative transfer calculations. We assess the magnitude of the 3D effect in small cumulus clouds by using a fine-scale cloud model to simulate many hours of cloudiness over a continental site. We then use a Monte Carlo radiative transfer model to compute the broadband shortwave fluxes at the surface twice, once using the complete three-dimensional radiative transfer F(sup 3D), and once using the ICA F (sup ICA); the difference between them is the 3D effect given.

Source record↗