Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallel processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,441 records · Page 80

Advancement of Extreme Environment Additively Manufactured Alloys for Next Generation Space Propulsion Applications

The National Aeronautics and Space Administration (NASA) has been involved in the development and maturation of metal additive manufacturing (AM) for space applications since the late 2000’s. Several efforts focused on the understanding of AM processes through material characterization and testing, standards development, component fabrication, and infusion into propulsion development and flight applications. NASA matured commonly used aerospace alloys from various alloy families (Nickel, Copper, Stainless and Steel, Aluminum, and Titanium-based) through detailed AM process and heat treatment characterization, in addition to mechanical and thermophysical testing. While these alloys are actively used in many propulsion applications, there is a need for ongoing AM optimized alloys using integrated computational materials engineering (ICME) and process development for high performance applications. The applications targeted are liquid rocket engines; advanced propulsion systems; and in-space propulsion with high heat fluxes, high pressure, and/or that use propellants that can degrade alloys (e.g., hydrogen). This paper highlights the characterization and physical properties of the more common AM alloys using laser powder bed fusion (L-PBF) and laser powder directed energy deposition (LP-DED) processes. Additionally, this paper discusses some of the ongoing novel alloy development and maturation using AM for use in these harsh environments, such as GRCop42, GRCop-84, NASA HR-1, GRX-810, and C-103. The results from these processes demonstrated that AM could enable rapid development and ongoing efforts for optimized alloys using ICME, yielding higher performances. These alloys have undergone modeling, fundamental metallurgical evaluations, heat treatment studies, detailed microstructure characterization, and mechanical testing campaigns. This, combined with direct application-specific component fabrication and hot-fire testing, enabled the increase of the Technology Readiness Level (TRL) through high duty-cycle testing. A background and overview of these novel AM-enabled alloys and AM processing developments including metallurgical and mechanical property studies is presented here. The latest advancement in the parallel component development and hot-fire testing and future developments for these alloys is also discussed.

Additive Manufacturing↗

Advancement of Extreme Environment Additively Manufactured Alloys for Next Generation Space Propulsion Applications

The National Aeronautics and Space Administration (NASA) has been involved in the development and maturation of metal additive manufacturing (AM) for space applications since the late 2000’s. Several efforts focused on the understanding of AM processes through material characterization and testing, standards development, component fabrication, and infusion into propulsion development and flight applications. NASA matured commonly used aerospace alloys from various alloy families (Nickel, Copper, Stainless and Steel, Aluminum, and Titanium-based) through detailed AM process and heat treatment characterization, in addition to mechanical and thermophysical testing. While these alloys are actively used in many propulsion applications, there is a need for ongoing AM optimized alloys using integrated computational materials engineering (ICME) and process development for high performance applications. The applications targeted are liquid rocket engines; advanced propulsion systems; and in-space propulsion with high heat fluxes, high pressure, and/or that use propellants that can degrade alloys (e.g., hydrogen). This paper highlights the characterization and physical properties of the more common AM alloys using laser powder bed fusion (L-PBF) and laser powder directed energy deposition (LP-DED) processes. Additionally, this paper discusses some of the ongoing novel alloy development and maturation using AM for use in these harsh environments, such as GRCop-42, GRCop-84, NASA HR-1, GRX-810, and C-103. The results from these processes demonstrated that AM could enable rapid development and ongoing efforts for optimized alloys using ICME, yielding higher performances. These alloys have undergone modeling, fundamental metallurgical evaluations, heat treatment studies, detailed microstructure characterization, and mechanical testing campaigns. This, combined with direct application-specific component fabrication and hot-fire testing, enabled the increase of the Technology Readiness Level (TRL) through high duty-cycle testing. A background and overview of these novel AM-enabled alloys and AM processing developments including metallurgical and mechanical property studies is presented here. The latest advancement in the parallel component development and hot-fire testing and future developments for these alloys is also discussed.

Additive Manufacturing↗

Cross-Cutting Flight Infrastructure Improvements on M2020

Mars2020 (M2020) was formulated as a mission that leveraged as much Mars Science Laboratory (MSL) heritage as possible, while focusing major new development efforts on the original and unique elements needed to accomplish the different mission objectives. Well publicized examples of high profile new developments include precision landing, the sampling and caching system, the specific instrument suite, improved mobility via Autonomous Navigation, and later the addition of the Ingenuity helicopter. Less well known are the refinements to the core flight infrastructure, primarily in the cross-cutting functions of Telecom, Avionics, Data Management, Communications Behaviors, and Parameter Management. These enhancements are introduced predominately via flight software, and represent increases in capability that justified their inclusion in an otherwise heritage-focused project environment.Perseverance’s cross-cutting flight infrastructure improvements fall into and across the following five categories. First is a trimming of the software footprint of infrastructure modules, in order to make room for memory demands elsewhere in the system. Second is the minimization of data volume to be downlinked, through various methods such as the incorporation of new compression options. Third is the maximization of the available downlink bandwidth for data, by curtailing content-less data (fill) and introducing an improved UHF proximity link protocol. Fourth is a reduction in vulnerabilities, through increased file system redundancy, robustness, and software process monitoring. Fifth is an increase in operations efficiency by lowering file system mount times, improving parallelism between simultaneous events, minimizing the time to recover from file system errors, streamlining the purging of obsolete data, and reducing the number of commands to service parameters by a factor of 100.Individually, none of the cross-cutting infrastructure improvements are likely to garner headlines, but collectively they appreciably improve the safety and operability of Perseverance over its predecessor. This paper will describe the improvements, their promise, and where applicable, their actual impact in operations.

Bohannon, Emily↗

FENIX: Towards a Fully Integrated Multiphysics Framework for Plasma Facing Component Modeling

Computational tools have a crucial role to play in accelerating the deployment of fusion as a clean, reliable, abundant, and sustainable energy source. Multiphysics, high-fidelity simulation capabilities can help model, study, and predict intricate interactions between materials performance, plasma exposure, neutron irradiation, and engineering processes. As such, they can assist in the resolution of scientific and engineering challenges underpinning design, construction, and commission of fusion power plants. To address these needs, ongoing efforts are leveraging the Multiphysics Object-Oriented Simulation Environment (MOOSE) framework and delivering new computational tools for the fusion community. These tools inherit crucial attributes from MOOSE. They are open-source, modular, integrated with nuclear industry-standard software quality assurance processes, and enable multiphysics, multi-fidelity, fully integrated, zero- to three-dimensional, and massively parallel simulations. After a short overview of these capabilities, we will present the development of Fusion ENergy Integrated multiphys-X (FENIX), a MOOSE-based application designed to enable plasma facing component design and performance evaluation. Throughout their lifetime, plasma facing components are exposed to extreme thermal loads, repeated thermal shocks, and irradiation by plasma ions, neutral particles, and high-energy neutrons. Consequently, designing a plasma facing component with acceptable lifetime degradation is extremely challenging. FENIX aims to model the multiphysics environment in which plasma facing components evolve to accelerate their design studies. To that end, FENIX couples existing MOOSE capabilities such as heat transfer, thermomechanics, and thermal hydraulics, with tritium transport via the MOOSE-based Tritium Migration Analysis Program, Version 8 (TMAP8), with neutronics via the MOOSE-based high-fidelity neutron-photon transport and fluid dynamics code Cardinal, and finally with Particle-in-Cell plasma simulation capabilities being developed in this project. In this study, we present the current FENIX capabilities and preliminary results of its application to model the Tritium Plasma Experiment set up at Idaho National Laboratory.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Graphics Processing Unit Assisted Thermographic Compositing

Objective: To develop a software application utilizing general purpose graphics processing units (GPUs) for the analysis of large sets of thermographic data. Background: Over the past few years, an increasing effort among scientists and engineers to utilize the GPU in a more general purpose fashion is allowing for supercomputer level results at individual workstations. As data sets grow, the methods to work them grow at an equal, and often great, pace. Certain common computations can take advantage of the massively parallel and optimized hardware constructs of the GPU to allow for throughput that was previously reserved for compute clusters. These common computations have high degrees of data parallelism, that is, they are the same computation applied to a large set of data where the result does not depend on other data elements. Signal (image) processing is one area were GPUs are being used to greatly increase the performance of certain algorithms and analysis techniques. Technical Methodology/Approach: Apply massively parallel algorithms and data structures to the specific analysis requirements presented when working with thermographic data sets.

Ragasa, Scott↗

Constrained Multipoint Aerodynamic Shape Optimization Using an Adjoint Formulation and Parallel Computers

An aerodynamic shape optimization method that treats the design of complex aircraft configurations subject to high fidelity computational fluid dynamics (CFD), geometric constraints and multiple design points is described. The design process will be greatly accelerated through the use of both control theory and distributed memory computer architectures. Control theory is employed to derive the adjoint differential equations whose solution allows for the evaluation of design gradient information at a fraction of the computational cost required by previous design methods. The resulting problem is implemented on parallel distributed memory architectures using a domain decomposition approach, an optimized communication schedule, and the MPI (Message Passing Interface) standard for portability and efficiency. The final result achieves very rapid aerodynamic design based on a higher order CFD method. In order to facilitate the integration of these high fidelity CFD approaches into future multi-disciplinary optimization (NW) applications, new methods must be developed which are capable of simultaneously addressing complex geometries, multiple objective functions, and geometric design constraints. In our earlier studies, we coupled the adjoint based design formulations with unconstrained optimization algorithms and showed that the approach was effective for the aerodynamic design of airfoils, wings, wing-bodies, and complex aircraft configurations. In many of the results presented in these earlier works, geometric constraints were satisfied either by a projection into feasible space or by posing the design space parameterization such that it automatically satisfied constraints. Furthermore, with the exception of reference 9 where the second author initially explored the use of multipoint design in conjunction with adjoint formulations, our earlier works have focused on single point design efforts. Here we demonstrate that the same methodology may be extended to treat complete configuration designs subject to multiple design points and geometric constraints. Examples are presented for both transonic and supersonic configurations ranging from wing alone designs to complex configuration designs involving wing, fuselage, nacelles and pylons.

Reuther, James↗

The role of analysis error in the convergence of reanalysis production streams in MERRA-2

Due to production time constraints, most reanalyses are produced in multiple parallel streams instead of a single continuous one. These streams cover separate segments of the reanalysis time period with short overlaps to allow reconstruction of the official record. A fundamental assumption justifying this approach is that the streams will be assimilating the same observations during the periods where they overlap, and so will eventually converge to a similar atmospheric state, making discontinuities at stream junctions negligible. This assumption is revisited in this work by examining the impact of analysis error on the differences between MERRA-2 overlapping streams in three historical periods. Comparison results are shown in terms of standard deviations of stream differences as well as the spectral decomposition of the variance of their differences. Residual differences were found at the end of each year of overlap, with larger values observed in the earlier segments of the presatellite era. By drawing parallels with analysis error statistics estimated from the GMAO OSSE system, these differences are shown to reflect the varying constraint of data with the varying observing network, and to further carry the imprint of errors that the data assimilation process is not able to mitigate. As such, they are unlikely to be reduced by longer spinup periods. The ability of data assimilation to ensure continuity in the parallel streams is put into question when the observing system coverage is inadequate or simply when the data assimilation system as a whole is suboptimal.

Amal El Akkraoui↗

Experiences with serial and parallel algorithms for channel routing using simulated annealing

Two algorithms for channel routing using simulated annealing are presented. Simulated annealing is an optimization methodology which allows the solution process to back up out of local minima that may be encountered by inappropriate selections. By properly controlling the annealing process, it is very likely that the optimal solution to an NP-complete problem such as channel routing may be found. The algorithm presented proposes very relaxed restrictions on the types of allowable transformations, including overlapping nets. By freeing that restriction and controlling overlap situations with an appropriate cost function, the algorithm becomes very flexible and can be applied to many extensions of channel routing. The selection of the transformation utilizes a number of heuristics, still retaining the pseudorandom nature of simulated annealing. The algorithm was implemented as a serial program for a workstation, and a parallel program designed for a hypercube computer. The details of the serial implementation are presented, including many of the heuristics used and some of the resulting solutions.

Brouwer, Randall Jay↗

Automating CPM-GOMS

CPM-GOMS is a modeling method that combines the task decomposition of a GOMS analysis with a model of human resource usage at the level of cognitive, perceptual, and motor operations. CPM-GOMS models have made accurate predictions about skilled user behavior in routine tasks, but developing such models is tedious and error-prone. We describe a process for automatically generating CPM-GOMS models from a hierarchical task decomposition expressed in a cognitive modeling tool called Apex. Resource scheduling in Apex automates the difficult task of interleaving the cognitive, perceptual, and motor resources underlying common task operators (e.g. mouse move-and-click). Apex's UI automatically generates PERT charts, which allow modelers to visualize a model's complex parallel behavior. Because interleaving and visualization is now automated, it is feasible to construct arbitrarily long sequences of behavior. To demonstrate the process, we present a model of automated teller interactions in Apex and discuss implications for user modeling. available to model human users, the Goals, Operators, Methods, and Selection (GOMS) method [6, 21] has been the most widely used, providing accurate, often zero-parameter, predictions of the routine performance of skilled users in a wide range of procedural tasks [6, 13, 15, 27, 28]. GOMS is meant to model routine behavior. The user is assumed to have methods that apply sequences of operators and to achieve a goal. Selection rules are applied when there is more than one method to achieve a goal. Many routine tasks lend themselves well to such decomposition. Decomposition produces a representation of the task as a set of nested goal states that include an initial state and a final state. The iterative decomposition into goals and nested subgoals can terminate in primitives of any desired granularity, the choice of level of detail dependent on the predictions required. Although GOMS has proven useful in HCI, tools to support the construction of GOMS models have not yet come into general use.

GOMS↗

Distributed Multi-GPU Community Detection on Exascale Computing Platforms

Community detection is a fundamental operation in graph mining, and by uncovering hidden structures and patterns within complex systems it helps solve fundamental problems pertaining to social networks, such as information diffusion, epidemics, and recommender systems. Scaling graph algorithms for massive networks becomes challenging on modern distributed-memory multi-GPU (Graphics Processing Unit) systems due to limitations such as irregular memory access patterns, load imbalances, higher communication-computation ratios, and cross-platform support. We present a novel algorithm HiPDPL-GPU (distributed parallel Louvain) to address these challenges. We conduct experiments involving different partitioning techniques to achieve optimized performance of HiPDPL-GPU on the two largest supercomputers: Frontier and Summit. Remarkably, HiPDPL-GPU processes a graph with 4.2 billion edges in less than 3 minutes using 1024 GPUs. Qualitatively performance of HiPDPL-GPU is similar or better compared to other state-of-the-art CPU- and GPU-based implementations. While prior GPU implementations have predominantly employed CUDA, our first-of-its-kind implementation for community detection is cross-platform, accommodating both AMD and NVIDIA GPUs.

graph algorithms, high performance comptuing↗

An Adaptive Kalman Filter using a Simple Residual Tuning Method

One difficulty in using Kalman filters in real world situations is the selection of the correct process noise, measurement noise, and initial state estimate and covariance. These parameters are commonly referred to as tuning parameters. Multiple methods have been developed to estimate these parameters. Most of those methods such as maximum likelihood, subspace, and observer Kalman Identification require extensive offline processing and are not suitable for real time processing. One technique, which is suitable for real time processing, is the residual tuning method. Any mismodeling of the filter tuning parameters will result in a non-white sequence for the filter measurement residuals. The residual tuning technique uses this information to estimate corrections to those tuning parameters. The actual implementation results in a set of sequential equations that run in parallel with the Kalman filter. Equations for the estimation of the measurement noise have also been developed. These algorithms are used to estimate the process noise and measurement noise for the Wide Field Infrared Explorer star tracker and gyro.

Harman, Richard R.↗

Fabrication and Characterization of A Lunar Simulant-Based Sintered Construction Material

In-situ resource utilization (ISRU) is critical to enable future efforts to have a long-term human presence on the Moon as well as Mars. ISRU technologies are being developed for radiation protection, dust mitigation, thermal insulation, and other applications. One such ISRU technology for creating construction materials out of lunar and Martian regolith is sintering, which is a thermal-based construction process that bonds finely grained material together at temperatures below the melting point. However, the conditions employed during the sintering, such as temperature, atmospheric composition, duration of the process, and pressure, can have a significant impact on the quality and strength of the resulting materials. In parallel, the development of methods for characterizing the quality, porosity, density, and other properties of these materials is critical. X-ray computed tomography (X-ray CT) can image large changes in density within a material, such as the presence of pores throughout an otherwise uniform medium, with relatively high spatial resolution. Similarly, Terahertz time-domain spectroscopic (THz-TDS) imaging is sensitive to density variations within samples, but is restricted to non-conducting materials. Specifically, previous work has shown that the refractive index (n eff ) values obtained through the analysis of THz-TDS images increases with increasing density within plastic samples. Even further, this work showed that it is possible to create a calibration curve for a given material from samples of different, but known density, which can enable one to directly convert n eff to density for samples having the same composition, but unknown density. Here, we report on the fabrication of a lunar simulant-based sintered construction material using vacuum hot pressed (VHP) sintering, then show X-ray CT and THz-TDS imaging results of the sample, which show spatial variations in the material. This has important implications for efforts to improve these types of lunar construction material processes and verify the quality of these materials in terms of consolidation. To the best of our knowledge, there is no previous work utilizing VHP sintering to make lunar simulant-based construction materials or exploring the feasibility of THz imaging to spatially map the density variation through a lunar simulant-based construction material.

Terahertz time-domain spectroscopic imaging↗

Graphics Processing Unit Assisted Thermographic Compositing

Objective Develop a software application utilizing high performance computing techniques, including general purpose graphics processing units (GPGPUs), for the analysis and visualization of large thermographic data sets. Over the past several years, an increasing effort among scientists and engineers to utilize graphics processing units (GPUs) in a more general purpose fashion is allowing for previously unobtainable levels of computation by individual workstations. As data sets grow, the methods to work them grow at an equal, and often greater, pace. Certain common computations can take advantage of the massively parallel and optimized hardware constructs of the GPU which yield significant increases in performance. These common computations have high degrees of data parallelism, that is, they are the same computation applied to a large set of data where the result does not depend on other data elements. Image processing is one area were GPUs are being used to greatly increase the performance of certain analysis and visualization techniques.

Ragasa, Scott↗

Parallel Computational Environment for Substructure Optimization

Design optimization of large structural systems can be attempted through a substructure strategy when convergence difficulties are encountered. When this strategy is used, the large structure is divided into several smaller substructures and a subproblem is defined for each substructure. The solution of the large optimization problem can be obtained iteratively through repeated solutions of the modest subproblems. Substructure strategies, in sequential as well as in parallel computational modes on a Cray YMP multiprocessor computer, have been incorporated in the optimization test bed CometBoards. CometBoards is an acronym for Comparative Evaluation Test Bed of Optimization and Analysis Routines for Design of Structures. Three issues, intensive computation, convergence of the iterative process, and analytically superior optimum, were addressed in the implementation of substructure optimization into CometBoards. Coupling between subproblems as well as local and global constraint grouping are essential for convergence of the iterative process. The substructure strategy can produce an analytically superior optimum different from what can be obtained by regular optimization. For the problems solved, substructure optimization in a parallel computational mode made effective use of all assigned processors.

Gendy, Atef S.↗

Distributed Multi-GPU Community Detection on Exascale Computing Platforms

Community detection is a fundamental operation in graph mining, and by uncovering hidden structures and patterns within complex systems it helps solve fundamental problems pertaining to social networks, such as information diffusion, epidemics, and recommender systems. Scaling graph algorithms for massive networks becomes challenging on modern distributed-memory multi-GPU (Graphics Processing Unit) systems due to limitations such as irregular memory access patterns, load imbalances, higher communication-computation ratios, and cross-platform support. We present a novel algorithm HiPDPL-GPU (Distributed Parallel Louvain) to address these challenges. We conduct experiments involving different partitioning techniques to achieve an optimized performance of HiPDPL-GPU on the two largest supercomputers: Frontier and Summit. Remarkably, HiPDPL-GPU processes a graph with 4.2 billion edges in less than 3 minutes using 1024 GPUs. Qualitatively, the performance of HiPDPL-GPU is similar or better compared to other state-of-the-art CPU- and GPU-based implementations. While prior GPU implementations have predominantly employed CUDA, our first-of-its-kind implementation for community detection is cross-platform, accommodating both AMD and NVIDIA GPUs.

Sattar, Naw Safrin↗

Computer architecture for efficient algorithmic executions in real-time systems: New technology for avionics systems and advanced space vehicles

Improvements and advances in the development of computer architecture now provide innovative technology for the recasting of traditional sequential solutions into high-performance, low-cost, parallel system to increase system performance. Research conducted in development of specialized computer architecture for the algorithmic execution of an avionics system, guidance and control problem in real time is described. A comprehensive treatment of both the hardware and software structures of a customized computer which performs real-time computation of guidance commands with updated estimates of target motion and time-to-go is presented. An optimal, real-time allocation algorithm was developed which maps the algorithmic tasks onto the processing elements. This allocation is based on the critical path analysis. The final stage is the design and development of the hardware structures suitable for the efficient execution of the allocated task graph. The processing element is designed for rapid execution of the allocated tasks. Fault tolerance is a key feature of the overall architecture. Parallel numerical integration techniques, tasks definitions, and allocation algorithms are discussed. The parallel implementation is analytically verified and the experimental results are presented. The design of the data-driven computer architecture, customized for the execution of the particular algorithm, is discussed.

Carroll, Chester C.↗

Real-Time Cognitive Computing Architecture for Data Fusion in a Dynamic Environment

A novel cognitive computing architecture is conceptualized for processing multiple channels of multi-modal sensory data streams simultaneously, and fusing the information in real time to generate intelligent reaction sequences. This unique architecture is capable of assimilating parallel data streams that could be analog, digital, synchronous/asynchronous, and could be programmed to act as a knowledge synthesizer and/or an "intelligent perception" processor. In this architecture, the bio-inspired models of visual pathway and olfactory receptor processing are combined as processing components, to achieve the composite function of "searching for a source of food while avoiding the predator." The architecture is particularly suited for scene analysis from visual data and odorant.

Duong, Tuan A.↗

Parallel VLSI Architecture

Fermat number transformation convolutes two digital data sequences. Very-large-scale integration (VLSI) applications, such as image and radar signal processing, X-ray reconstruction, and spectrum shaping, linear convolution of two digital data sequences of arbitrary lenghts accomplished using Fermat number transform (ENT).

Truong, T. K.↗