Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallel processing (computers)”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 955 records · Page 53

SHARP - A multi-mission AI system for spacecraft telemetry monitoring and diagnosis

The Spacecraft Health Automated Reasoning Prototype (SHARP) is a system designed to demonstrate automated health and status analysis for multi-mission spacecraft and ground data systems operations. Telecommunications link analysis of the Voyager II spacecraft is the initial focus for the SHARP system demonstration which will occur during Voyager's encounter with the planet Neptune in August, 1989, in parallel with real-time Voyager operations. The SHARP system combines conventional computer science methodologies with artificial intelligence techniques to produce an effective method for detecting and analyzing potential spacecraft and ground systems problems. The system performs real-time analysis of spacecraft and other related telemetry, and is also capable of examining data in historical context. A brief introduction is given to the spacecraft and ground systems monitoring process at the Jet Propulsion Laboratory. The current method of operation for monitoring the Voyager Telecommunications subsystem is described, and the difficulties associated with the existing technology are highlighted. The approach taken in the SHARP system to overcome the current limitations is also described, as well as both the conventional and artificial intelligence solutions developed in SHARP.

Lawson, Denise L.↗

A novel VLSI processor architecture for supercomputing arrays

Design of the processor element for general purpose massively parallel supercomputing arrays is highly complex and cost ineffective. To overcome this, the architecture and organization of the functional units of the processor element should be such as to suit the diverse computational structures and simplify mapping of complex communication structures of different classes of algorithms. This demands that the computation and communication structures of different class of algorithms be unified. While unifying the different communication structures is a difficult process, analysis of a wide class of algorithms reveals that their computation structures can be expressed in terms of basic IP,IP,OP,CM,R,SM, and MAA operations. The execution of these operations is unified on the PAcube macro-cell array. Based on this PAcube macro-cell array, we present a novel processor element called the GIPOP processor, which has dedicated functional units to perform the above operations. The architecture and organization of these functional units are such to satisfy the two important criteria mentioned above. The structure of the macro-cell and the unification process has led to a very regular and simpler design of the GIPOP processor. The production cost of the GIPOP processor is drastically reduced as it is designed on high performance mask programmable PAcube arrays.

Venkateswaran, N.↗

Comparison of 250 MHz R10K Origin 2000 and 400 MHz Origin 2000 Using NAS Parallel Benchmarks

This report describes results of benchmark tests on Steger, a 250 MHz Origin 2000 system with R10K processors, currently installed at the NASA Ames National Advanced Supercomputing (NAS) facility. For comparison purposes, the tests were also run on Lomax, a 400 MHz Origin 2000 with R12K processors. The BT, LU, and SP application benchmarks in the NAS Parallel Benchmark Suite and the kernel benchmark FT were chosen to measure system performance. Having been written to measure performance on Computational Fluid Dynamics applications, these benchmarks are assumed appropriate to represent the NAS workload. Since the NAS runs both message passing (MPI) and shared-memory, compiler directive type codes, both MPI and OpenMP versions of the benchmarks were used. The MPI versions used were the latest official release of the NAS Parallel Benchmarks, version 2.3. The OpenMP versions used were PBN3b2, a beta version that is in the process of being released. NPB 2.3 and PBN3b2 are technically different benchmarks, and NPB results are not directly comparable to PBN results.

Turney, Raymond D.↗

Comparison of Origin 2000 and Origin 3000 Using NAS Parallel Benchmarks

This report describes results of benchmark tests on the Origin 3000 system currently being installed at the NASA Ames National Advanced Supercomputing facility. This machine will ultimately contain 1024 R14K processors. The first part of the system, installed in November, 2000 and named mendel, is an Origin 3000 with 128 R12K processors. For comparison purposes, the tests were also run on lomax, an Origin 2000 with R12K processors. The BT, LU, and SP application benchmarks in the NAS Parallel Benchmark Suite and the kernel benchmark FT were chosen to determine system performance and measure the impact of changes on the machine as it evolves. Having been written to measure performance on Computational Fluid Dynamics applications, these benchmarks are assumed appropriate to represent the NAS workload. Since the NAS runs both message passing (MPI) and shared-memory, compiler directive type codes, both MPI and OpenMP versions of the benchmarks were used. The MPI versions used were the latest official release of the NAS Parallel Benchmarks, version 2.3. The OpenMP versiqns used were PBN3b2, a beta version that is in the process of being released. NPB 2.3 and PBN 3b2 are technically different benchmarks, and NPB results are not directly comparable to PBN results.

Turney, Raymond D.↗

The Use of Field Programmable Gate Arrays (FPGA) in Small Satellite Communication Systems

This paper will describe the use of digital Field Programmable Gate Arrays (FPGA) to contribute to advancing the state-of-the-art in software defined radio (SDR) transponder design for the emerging SmallSat and CubeSat industry and to provide advances for NASA as described in the TAO5 Communication and Navigation Roadmap (Ref 4). The use of software defined radios (SDR) has been around for a long time. A typical implementation of the SDR is to use a processor and write software to implement all the functions of filtering, carrier recovery, error correction, framing etc. Even with modern high speed and low power digital signal processors, high speed memories, and efficient coding, the compute intensive nature of digital filters, error correcting and other algorithms is too much for modern processors to get efficient use of the available bandwidth to the ground. By using FPGAs, these compute intensive tasks can be done in parallel, pipelined fashion and more efficiently use every clock cycle to significantly increase throughput while maintaining low power. These methods will implement digital radios with significant data rates in the X and Ka bands. Using these state-of-the-art technologies, unprecedented uplink and downlink capabilities can be achieved in a 1/2 U sized telemetry system. Additionally, modern FPGAs have embedded processing systems, such as ARM cores, integrated inside the FPGA allowing mundane tasks such as parameter commanding to occur easily and flexibly. Potential partners include other NASA centers, industry and the DOD. These assets are associated with small satellite demonstration flights, LEO and deep space applications. MSFC currently has an SDR transponder test-bed using Hardware-in-the-Loop techniques to evaluate and improve SDR technologies.

Varnavas, Kosta↗

Recent Advancements in the PATO Material Response Code

Introduction: Predicting the complicated multiphysics phenomena during atmospheric entry requires high-fidelity modeling tools to refine estimates of mission risks during entry. To this end, new capabilities are being added to the Porous-material Analysis Toolbox based on OpenFOAM (PATO) [1,2,3]. PATO is an open-source software for Computational Material Response (CMR) of reactive porous materials submitted to high-temperature environments. The objective of this work is to highlight current efforts to add to and improve upon the modeling capabilities of PATO. These include efforts to loosely couple PATO with other discipline specialized codes including hypersonic Computational Fluid Dynamics (CFD), to assess the interaction effects between pyrolysis gas blowing and the boundary layer, and Computational Solid Mechanics (CSM), to address modeling of mechanical erosion. Other refinements include surface phenomena modeling capabilities to address the effects of silicone-based coatings applied to the TPS during flight preparation, and a unified multiphase solver for a mixed porous-material and plain-fluid domain. Coupling CMR with CFD (CMR/CFD): A loose coupling between PATO and the Data Parallel Line Relaxation (DPLR) [4] CFD code has been achieved by making use of a blowing boundary condition at the heatshield surface available in DPLR. Starting with heat flux estimates with no pyrolysis gas blowing at the surface, blowing gases are computed by the CMR and passed to the CFD such that aerothermal properties of the environment can be recomputed for a new CMR computation. This leads to an iterative process which is supplemented with an estimate of the radiative heat flux using the Nonequilibrium air radiation (NEQAIR) [5] program. The entire iterative process is illustrated in Figure 1. This coupling strategy has been utilized in computing the MSL material response. The goal is to compare the coupled CMR/CFD results with material response results obtained using traditional blowing corrections. Coupling CMS with CMR: A mechanical erosion model is currently being implemented in PATO to account for the additional mass removal induced by high shear conditions. The modeling process at each timestep consists of updating the mechanical properties as a function of temperature and computing the stress tensor and displacement fields of the material. Then, a failure criteria model determines the regions in which the stress exceeds the ultimate strength values resulting in mesh movement to account for mass removal. This model allows the material response simulation to compute the recession due to both oxidation and shear-induced erosion. The model is demonstrated by computing material response of sphere-cone arc jet samples. Surface Modeling Capabilities: NuSil, a silicone-based coating, was sprayed onto the MSL and Mars 2020 heatshields to mitigate shedding of phenolic dust. To better understand the effects of the NuSil coating on the material response, a novel model has been implemented in PATO. In this model, the equilibrium of the charred NuSil surface is modeled as pure silica, and a constant offset, inspired by the classical spallation model, is added to the the char blowing rate and wall enthalpy to reproduce HyMETS experimental results. The model has also been used to estimate the 3D material response of the MSL heatshield [6]. Unified Solver: In addition to the iterative loose coupling approach mentioned above, a multiphase unified solver is being developed to couple the environment (plain-fluid phase) and the porous-material phase. The solver is based on the volume averaged conservation of mass, momentum, and energy for the macroscale with closure models which include microscale effects through effective physicochemical properties. The unified solver has been used to compute flow through a porous plug and solve the Beavers and Joseph problem [7]. Since the strong coupling between phases is inherent to this solver, modeling assumptions present in other coupling methods of material response are mitigated. This strategy also makes it feasible to capture the competition between surface and volume ablation in the same computational domain, which is usually not possible with other coupling approaches.

Thermal Protection Systems↗

Assessment of ESM Readiness Level for Exascale HPC

Advancement of Earth System Models (ESMs) is becoming increasingly challenging due to a confluence of factors including increasing model complexity – to more fully represent the earth system, increasing spatial resolution - to achieve higher accuracy by resolving fine-scale dynamical to physical, biological, and chemical processes and their interaction, increasing ensemble size - to more accurately represent predictive uncertainty, and increased computing requirements – to enable more accurate and timely weather predictions and climate projections for societal benefit. The belief by many that computing will take care of itself is no longer valid given the disruptive changes in HPC that are driving up the cost of computing, increasing the difficulty of using emerging HPC effectively, and exposing limits in parallelism, portability and scalability of the ESM applications themselves.

54 ENVIRONMENTAL SCIENCES↗

VLSI architectures for geometrical mapping problems in high-definition image processing

This paper explores a VLSI architecture for geometrical mapping address computation. The geometric transformation is discussed in the context of plane projective geometry, which invokes a set of basic transformations to be implemented for the general image processing. The homogeneous and 2-dimensional cartesian coordinates are employed to represent the transformations, each of which is implemented via an augmented CORDIC as a processing element. A specific scheme for a processor, which utilizes full-pipelining at the macro-level and parallel constant-factor-redundant arithmetic and full-pipelining at the micro-level, is assessed to produce a single VLSI chip for HDTV applications using state-of-art MOS technology.

Kim, K.↗

Root-Raised Cosine Filter Implementation That Uses Canonical Signed Digits for High-Speed Digital Filter Applications

NASA Lewis Research Center's Space Communications Division has been investigating high-speed digital filters that can operate at a higher speed than those in current use for a digital modulator and demodulator (modem). Using the Canonical Signed Digits (CSD) number representation for filter coefficients is a very effective way to increase the filter's speed while reducing complexity in the digital filter hardware design. This approach is a good alternative to using an expensive parallel-processing design technique or custom, application-specific integrated circuits. Such integrated circuits may not be suitable for applications that require filter speeds faster than what application-specific integrated circuits digital signal processors can offer for a dedicated channel. When a communication channel is a dedicated, multiplication process--a costly, time-consuming process--it can be greatly simplified by a replacement of the filter coefficients with CSD numbers. A computer code written with the MATLAB software package runs the program and generates CSD-represented filter coefficients that are based on minimizing minimum mean square errors. Also, the Alta Group of Cadence's Signal Processing Workstation is used to simulate and analyze the CSD filter responses. The impulse response of the root-raised cosine filter that is used as a base model is defined. From this filter, a set of coefficients is sampled and stored in a file. For the all coefficients, the optimal CSD number for each coefficient is searched on the basis of the minimum-mean-square-errors criterion. Because the distribution of CSD numbers is not uniform, quantization errors tend to be bigger for coefficients greater than 1/2. To offset errors that occur in a region of coefficients between 1/2 to 1 and to better represent fractions with CSD numbers, an extra nonzero digit is allowed for any coefficients exceeding 1/2. This will greatly improve frequency response as well as intersymbol interference at the receiver. The frequency response of a set of collected CSD-represented filter coefficients was compared with the same filter that was conventionally implemented. Analyses show CSD-implemented filters perform as well as conventional filters. Comparison of eye diagrams and bit-error-rate curves between CSD filters and traditionally implemented filters are almost indistinguishable. However, filter complexity was reduced from almost 3.5 to 1 for CSD filters. Complete computer simulation results are available. In the near future, work will focus on building actual working digital filter hardware in a field programmable gate array (FPGA).

Kim, Heechul↗

Automated rendezvous and docking with video imagery

For rendezvous and docking, assessing and tracking relative orientation is necessary within a minimum approach distance. Special target light patterns have previously been considered for use with video sensors for ease of determining relative orientation. A generalization of those approaches is addressed. At certain ranges, the entire structure of the target vehicle constitutes an acceptable target; at closer ranges, substructures will suffice. Acting on the same principle as the human intelligence, these structures can be compared with a memory model to assess the relative orientation and range. Models for comparison are constructed from a CAD facet model and current imagery. This approach requires fast image handling, projection, and comparison techniques which rely on rapidly developing parallel processing technology. Relative orientation and range assessment consists of successful comparison of the perceived target aspect with a known aspect. Generating a known projection from a model within required times, say subsecond times, is only now approaching feasibility. With this capability, rates of comparison used by the human brain can be approached and arbitrary known structures can be compared in reasonable times. Future space programs will have access to powerful computation devices which far exceed even this capability. For example, the possibility will exist to assess unknown structures and then control rendezvous and docking, all at very fast rates. The first step which has the current utility, namely applying this to known structures, is taken.

Rodgers, Mike↗

Smart-Pixel Array Processors Based on Optimal Cellular Neural Networks for Space Sensor Applications

A smart-pixel cellular neural network (CNN) with hardware annealing capability, digitally programmable synaptic weights, and multisensor parallel interface has been under development for advanced space sensor applications. The smart-pixel CNN architecture is a programmable multi-dimensional array of optoelectronic neurons which are locally connected with their local neurons and associated active-pixel sensors. Integration of the neuroprocessor in each processor node of a scalable multiprocessor system offers orders-of-magnitude computing performance enhancements for on-board real-time intelligent multisensor processing and control tasks of advanced small satellites. The smart-pixel CNN operation theory, architecture, design and implementation, and system applications are investigated in detail. The VLSI (Very Large Scale Integration) implementation feasibility was illustrated by a prototype smart-pixel 5x5 neuroprocessor array chip of active dimensions 1380 micron x 746 micron in a 2-micron CMOS technology.

Fang, Wai-Chi↗

Sculpting in cyberspace: Parallel processing the development of new software

Stimulating creativity in problem solving, particularly where software development is involved, is applicable to many disciplines. Metaphorical thinking keeps the problem in focus but in a different light, jarring people out of their mental ruts and sparking fresh insights. It forces the mind to stretch to find patterns between dissimilar concepts, in the hope of discovering unusual ideas in odd associations (Technology Review January 1993, p. 37). With a background in Engineering and Visual Design from MIT, I have for the past 30 years pursued a career as a sculptor of interdisciplinary monumental artworks that bridge the fields of science, engineering and art. Since 1979, I have pioneered the application of computer simulation to solve the complex problems associated with these projects. A recent project for the roof of the Carnegie Science Center in Pittsburgh made particular use of the metaphoric creativity technique described above. The problem-solving process led to the creation of hybrid software combining scientific, architectural and engineering visualization techniques. David Steich, a Doctoral Candidate in Electrical Engineering at Penn State, was commissioned to develop special software that enabled me to create innovative free-form sculpture. This paper explores the process of inventing the software through a detailed analysis of the interaction between an artist and a computer programmer.

Fisher, Rob↗

Multi-task Parallelism for Robust Pre-training of Graph Foundation Models on Multi-source, Multi-fidelity Atomistic Modeling Data

Graph foundation models using graph neural networks promise sustainable, efficient atomistic modeling. To tackle challenges of processing multi-source, multi-fidelity data during pre-training, recent studies employ multi-task learning, in which shared message passing layers initially process input atomistic structures regardless of source, then route them to multiple decoding heads that predict data-specific outputs. This approach stabilizes pre-training and enhances a model’s transferability to unexplored chemical regions. Preliminary results on approximately four million structures are encouraging, yet questions remain about generalizability to larger, more diverse datasets and scalability on supercomputers. We propose a multi-task parallelism method that distributes each head across computing resources with GPU acceleration. Implemented in the open-source HydraGNN architecture, our method was trained on over 24 million structures from five datasets and tested on the Perlmutter, Aurora, and Frontier supercomputers, demonstrating efficient scaling on all three highly heterogeneous super-computing architectures.

Lupo Pasini, Massimiliano [ORNL] (ORCID:0000000249↗

Rapid Commissioning of Large Machine Tools Using Finite Element-Based Correction of Geometric Errors

Large computer numerical control (CNC) machine tools derive their stiffness from monolithic cast iron bases or weldments that are sometimes integral to machine motion systems like box ways or guideways. However, the sheer size of castings and even floor flatness deviations result in dimensional errors in these systems, which manifest as machine motion errors. Typical geometric alignment processes rely on an iterative approach, where measurements are taken to assess alignment (straightness, squareness, and parallelism), followed by adjustment of the machine supports (fixators or leveling pads), which can take weeks even for an experienced operator. Conversely, a novel method is proposed to shorten the correction time by eliminating the trial-and-error process in favor of a more deterministic approach guided by a finite element (FE) method. A feasibility study is conducted on a CNC polymer hybrid machine, with a steel weldment frame, supported by six leveling pads. An FE model of the frame is utilized to obtain recommended leveling pad adjustments, based on measurement of machine errors taken using a laser tracker. After a single adjustment cycle, measurements reveal that geometric errors of the machine tool are reduced from 2.22 mm of flatness deviation to 0.32 mm, achieving an 85.6% reduction. Furthermore, the entire process including measurement, adjustment, and assessment is completed in just 6 h by two operators who are not professional service engineers. In conclusion, this methodology demonstrates feasibility for scaling up, especially to large, high-precision CNC machine tools with bases mounted by fixators, offering the capability for bidirectional adjustment.

42 ENGINEERING↗

ChatMPI: LLM-Driven MPI Code Generation for HPC Workloads

The Message Passing Interface (MPI) standard plays a crucial role in enabling scientific applications for parallel computing and is an essential component in high-performance computing (HPC). However, implementing MPI code manually—especially applying a proper domain decomposition and communication pattern—is a challenging and error-prone task. We present ChatMPI, an AI assistant for MPI parallelization of sequential C codes. In our analysis, we focus on testing six essential HPC workloads, which are based on Basic Linear Algebra Subprograms levels 1, 2, and 3 as well as sparse, stencil, and iterative operations. We analyze the process of creating ChatMPI by using the ChatHPC library. This lightweight large language model (LLM)–based infrastructure enables HPC experts to efficiently create and supervise trustworthy AI capabilities for critical HPC software tasks. We study the data required for training (fine-tuning) ChatMPI to generate parallel codes that not only use MPI syntax correctly but also apply HPC techniques to reduce memory communication and maximize performance by using proper work decomposition. With a relatively small training dataset composed of a few dozen prompts and fewer than 15 minutes of fine-tuning on one node equipped with two NVIDIA H100 GPUs, ChatMPI elevates trustworthiness for MPI code generation of current LLMs (e.g., Code Llama, ChatGPT-4o and ChatGPT 5). Additionally, we evaluate the performance of the MPI codes generated by ChatMPI in comparison with the ones generated by ChatGPT-4o and ChatGPT-5. The codes generated by ChatMPI provide up to a 4 × boost in performance by using better problem decomposition, communication patterns, and HPC techniques (e.g., communication avoiding).

Valero Lara, Pedro [ORNL] (ORCID:0000000214794310)↗

Advanced information processing system: The Army Fault-Tolerant Architecture detailed design overview

The Army Avionics Research and Development Activity (AVRADA) is pursuing programs that would enable effective and efficient management of large amounts of situational data that occurs during tactical rotorcraft missions. The Computer Aided Low Altitude Night Helicopter Flight Program has identified automated Terrain Following/Terrain Avoidance, Nap of the Earth (TF/TA, NOE) operation as key enabling technology for advanced tactical rotorcraft to enhance mission survivability and mission effectiveness. The processing of critical information at low altitudes with short reaction times is life-critical and mission-critical necessitating an ultra-reliable/high throughput computing platform for dependable service for flight control, fusion of sensor data, route planning, near-field/far-field navigation, and obstacle avoidance operations. To address these needs the Army Fault Tolerant Architecture (AFTA) is being designed and developed. This computer system is based upon the Fault Tolerant Parallel Processor (FTPP) developed by Charles Stark Draper Labs (CSDL). AFTA is hard real-time, Byzantine, fault-tolerant parallel processor which is programmed in the ADA language. This document describes the results of the Detailed Design (Phase 2 and 3 of a 3-year project) of the AFTA development. This document contains detailed descriptions of the program objectives, the TF/TA NOE application requirements, architecture, hardware design, operating systems design, systems performance measurements and analytical models.

Harper, Richard E.↗

Parallelization of ARC3D with Computer-Aided Tools

A series of efforts have been devoted to investigating methods of porting and parallelizing applications quickly and efficiently for new architectures, such as the SCSI Origin 2000 and Cray T3E. This report presents the parallelization of a CFD application, ARC3D, using the computer-aided tools, Cesspools. Steps of parallelizing this code and requirements of achieving better performance are discussed. The generated parallel version has achieved reasonably well performance, for example, having a speedup of 30 for 36 Cray T3E processors. However, this performance could not be obtained without modification of the original serial code. It is suggested that in many cases improving serial code and performing necessary code transformations are important parts for the automated parallelization process although user intervention in many of these parts are still necessary. Nevertheless, development and improvement of useful software tools, such as Cesspools, can help trim down many tedious parallelization details and improve the processing efficiency.

Jin, Haoqiang↗

New Tools for Automating Arcjet Sample Recession Tracking and Analysis

Arcjet Computer Vision (arcjetCV) has been significantly upgraded to enhance accuracy and performance in tracking material recession and shock-material standoff in test videos. These improvements include integrating new machine learning models, developing a specialized edge detection class, and incorporating a more comprehensive training dataset. These upgrades have refined the software’s ability to automate time-resolved recession tracking, making it more precise and reliable for analyzing complex physical processes. In parallel, a new tool called STARscan (Spatial Targeting and Alignment Rig for Scanning) is being developed to capture detailed 3D surface data before and after testing. By comparing these pre- and post-test scans with arcjetCV’s automated video analysis results, users can achieve a more comprehensive assessment of material recession. This method enables cross-validation of results, improving confidence in the analysis of tested materials. The expanded capabilities of arcjetCV have been successfully demonstrated on videos from various facilities, including the NASA Ames arcjets, UIUC’s PlasmatronX, and the VKI Plasmatron. It has been adopted as a new standard for in-situ recession tracking by the Mars Sample Return Project and Orion. ArcjetCV’s improved efficiency and accuracy are critical for reducing testing uncertainties and validating heatshield material performance under extreme conditions. The software’s user-friendly graphical interface ensures ease of use, enabling seamless processing and precise analysis of arcjet videos, providing deeper insights into material behavior in hypersonic environments. ArcjetCV is now available on both PyPI and Conda, allowing easy installation via "pip install arcjetCV" or through the Conda package manager, ensuring broad accessibility and streamlined deployment for users across various platforms.

Ablation↗