Search NASA⌕ Search

SEARCH · Search NASA

Results for “High performance computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Computational Performance of Progressive Damage Analysis of Composite Laminates using Abaqus/Explicit with 16 to 512 CPU Cores

The computational scaling performance of progressive damage analysis using Abaqus/ Explicit is evaluated and quantified using from 16 to 512 CPU cores. Several analyses were conducted on varying numbers of cores to determine the scalability of the code on five NASA high performance computing systems. Two finite element models representative of typical models used for progressive damage analysis of composite laminates were used. The results indicate a 10 to 15 times speed up scaling from 24 to 512 cores. The run times were modestly reduced with newer generations of CPU hardware. If the number of degrees of freedom is held constant with respect to the number of cores, the model size can be increased by a factor of 20, scaling from 16 to 512 cores, with the same run time. An empirical expression was derived relating run time, the number of cores, and the number of degrees of freedom. Analysis cost was examined in terms of software tokens and hardware utilization. Using additional cores reduces token usage since the computational performance increases more rapidly than the token requirement with increasing number of cores. The in- crease in hardware cost with increasing cores was found to be modest. Overall the results show relatively good scalability of the Abaqus/Explicit code on up to 512 cores.

Bergan, A. C.↗

Parallel aeroelastic computations for wing and wing-body configurations

The objective of this research is to develop computationally efficient methods for solving fluid-structural interaction problems by directly coupling finite difference Euler/Navier-Stokes equations for fluids and finite element dynamics equations for structures on parallel computers. This capability will significantly impact many aerospace projects of national importance such as Advanced Subsonic Civil Transport (ASCT), where the structural stability margin becomes very critical at the transonic region. This research effort will have direct impact on the High Performance Computing and Communication (HPCC) Program of NASA in the area of parallel computing.

Byun, Chansup↗

Applications of massively parallel computers in telemetry processing

Telemetry processing refers to the reconstruction of full resolution raw instrumentation data with artifacts, of space and ground recording and transmission, removed. Being the first processing phase of satellite data, this process is also referred to as level-zero processing. This study is aimed at investigating the use of massively parallel computing technology in providing level-zero processing to spaceflights that adhere to the recommendations of the Consultative Committee on Space Data Systems (CCSDS). The workload characteristics, of level-zero processing, are used to identify processing requirements in high-performance computing systems. An example of level-zero functions on a SIMD MPP, such as the MasPar, is discussed. The requirements in this paper are based in part on the Earth Observing System (EOS) Data and Operation System (EDOS).

El-Ghazawi, Tarek A.↗

Load Balancing Strategies for Multi-Block Overset Grid Applications

The multi-block overset grid method is a powerful technique for high-fidelity computational fluid dynamics (CFD) simulations about complex aerospace configurations. The solution process uses a grid system that discretizes the problem domain by using separately generated but overlapping structured grids that periodically update and exchange boundary information through interpolation. For efficient high performance computations of large-scale realistic applications using this methodology, the individual grids must be properly partitioned among the parallel processors. Overall performance, therefore, largely depends on the quality of load balancing. In this paper, we present three different load balancing strategies far overset grids and analyze their effects on the parallel efficiency of a Navier-Stokes CFD application running on an SGI Origin2000 machine.

Djomehri, M. Jahed↗

Parallel Domain Decomposition Formulation and Software for Large-Scale Sparse Symmetrical/Unsymmetrical Aeroacoustic Applications

The overall objectives of this research work are to formulate and validate efficient parallel algorithms, and to efficiently design/implement computer software for solving large-scale acoustic problems, arised from the unified frameworks of the finite element procedures. The adopted parallel Finite Element (FE) Domain Decomposition (DD) procedures should fully take advantages of multiple processing capabilities offered by most modern high performance computing platforms for efficient parallel computation. To achieve this objective. the formulation needs to integrate efficient sparse (and dense) assembly techniques, hybrid (or mixed) direct and iterative equation solvers, proper pre-conditioned strategies, unrolling strategies, and effective processors' communicating schemes. Finally, the numerical performance of the developed parallel finite element procedures will be evaluated by solving series of structural, and acoustic (symmetrical and un-symmetrical) problems (in different computing platforms). Comparisons with existing "commercialized" and/or "public domain" software are also included, whenever possible.

Nguyen, D. T.↗

Advances in PSP Testing in LaRC High Reynolds Number Facilities

The use of luminescent coatings for the global measurement of surface aerodynamic properties at full-flight Reynolds numbers has been ongoing at the NASA Langley Research Center since the late 1990s, beginning with Temperature Sensitive Paint (TSP) for boundary layer analysis at the 0.3-m Transonic Cryogenics Tunnel. Since then, significant work has been made to extend these measurements to pressure using Pressure Sensitive Paint (PSP) as well as develop systems for application in larger scale, full-flight Reynolds number facilities, such as the National Transonic Facility (NTF). Recently, a system for the measurement of time-resolved pressure using recently developed unsteady PSP (uPSP) formulations has been designed and implemented in the Transonic Dynamics Tunnel (TDT). The use of PSP (or uPSP) in either of these large scale, full-fight Reynolds number wind tunnels is complicated by the fact that both facilities operate best in oxygen deficient environments; however, the PSP technique relies on the presence of oxygen to work. The NTF is typically operated in cryogenic conditions, which is achieved using liquid nitrogen. In this facility, temperatures can reach as low as 116 K (-250 °F) with nominal oxygen concentrations of less than 50 ppm. This has resulted in the development of specialized PSP formulations that can operate in the cryogenic environment with the introduction of low amounts of oxygen (typically less than 2000 ppm) for PSP response. Likewise, to achieve full-flight Reynolds number conditions in the TDT, the atmosphere in the tunnel is replaced with a “heavy gas” of R-134A (1,1,1,2-tetrafluoroethane), which also has minimal oxygen native in the flow. The use of uPSP in this case will also depend on the introduction of trace amounts of oxygen, with the precise concentrations to be determined in an upcoming test. This presentation will describe some recent advancements that have been made for PSP measurements in both facilities. For the NTF system, there has been significant development of enhanced lighting for use at cryogenic conditions, as well as improvements in the application efficiency of the PSP. Furthermore, there are efforts underway to incorporate advanced data analysis techniques to acquire additional surface aerodynamic properties using the PSP technique that can yield not only improved experimental efficiency, but also provide data needed for next generation vehicle design and development. For the uPSP system in TDT, significant efforts to improve the data transfer rates to a high-performance computing environment for analysis are underway and performance and initial results from the system in the heavy-gas environment will be presented.

Pressure Sensitive Paint↗

NASA GRC ICME Schema for Materials Data Management: An Executive Summary

Integrated Computational Materials Engineering (ICME) has received a growing emphasis in attention due its potential impact on rapid material design, reduction in cost and time to market for new applications, and the promise of ‘fit-for-purpose’ materials coupled with recent advances in high performance computing and material characterization tools. However, for an organization to implement ICME practices for material discovery and design, a series of both technical and cultural challenges must be overcome to foster an environment that enables efficient, traceable, and predictive multiscale simulations of material behavior to enable virtual design of materials. In 2016, NASA sponsored a 2040 Vision study to define the potential 25-year future state required for integrated multiscale modeling of materials and systems to improve both the associated time and cost for aerospace and aeronautical innovation. The study envisions a cyber-physical-social ecosystem of experimentally validated computational models, tools, and techniques, along with the associated digital tapestry, that can enable rapid, optimized, ‘fit-for-purpose’ design of materials, components, and systems. A key requirement for such an ecosystem is the development of a robust information management system for materials across their full lifecycle, including material pedigree, experimental (real) and virtual (simulation) data, developed material models, and the implementation of models in engineering applications, such that process-structure-property-performance relationships can be established, thereby enabling the virtual design and optimization of materials. Such an information management system must be able to effectively capture: i) material information at each length scale; ii) test data and analysis; iii) associated material models; and iv) material and model deployment in engineering applications. These systems must also provide traceability between experimental and virtual representations of the material to ensure, when appropriate, the material digital twin is maintained. Additionally, this robust material information management system must be able to seamlessly connect with both commercial and an organization’s in-house software tools, be they analysis tools, other material databases, product lifecycle management (PLM) or simulation data management (SDM) tools, etc., such that automation of the design and analysis of a material across multiple length scales is possible. In this paper, an executive summary of the NASA GRC ICME Schema for materials information management is presented. The database best practices and schema design philosophy specifically for ICME materials data management and an overview description of each element in the schema is given, along with its associated role in an ICME workflow. Additionally, auxiliary tools that interact with the database and provide judicious automation with regards to importing, exporting, and analyzing materials data are presented. Such tools are critical to an ICME ecosystem, not only for their role in enabling optimization, but also in relieving users of tedious manual tasks, thus helping to promote adoption and combat the cultural challenges organizations face in enabling ICME.

Materials↗

Space-based crystal growth and thermocapillary flow

The demand for larger crystals is increasing especially in applications associated with the electronic industry, where large and pure electronic crystals (notably silicon) are the essential material to make high-performance computer chips. Crystal growth under weightless conditions has been considered an ideal way to produce bigger and hopefully better crystals. One technique which may benefit from a microgravity environment is the float-zone crystal-growth process, a containerless method for producing high-quality electronic material. In this method, a rod of material to be refined is moved slowly through a heating device which melts a portion of it. Ideally, as the melt resolidifies it does so as a single crystal which is then used as substrate for building microelectronic devices. The possibility of contamination by contact with other material is reduced because of the 'float' configuration. However, since the weight of the material contained in the zone is supported by the surface-tension force, the size of the resulting crystal is limited in Earth-based productions; in fact, some materials have properties which prevent this process from being used to manufacture crystals of reasonable size. Consequently, there has been a great deal of interest in exploiting the microgravity environment of space to grow larger size crystals of electronic material using the float-zone method. In addition to allowing larger crystals to be grown, a microgravity environment would also significantly reduce the magnitude of convection induced by buoyancy forces during the melting state. This type of convection was once thought to be at least partially responsible for the presence of undesirable nonuniformities--called striations--in material properties observed in float-zone material. However, past experiments on crystal growth under weightless conditions found that even with the absence of gravity, the float-zone method sometimes still results striations. It is believed that another mechanism is playing a dominant role in the microgravity environment.

Shen, Yong-Hong↗

An Assessment of the State-of-the-art in Multidisciplinary Aeromechanical Analyses

This paper presents a survey of the current state-of-the-art in multidisciplinary aeromechanical analyses which integrate advanced Computational Structural Dynamics (CSD) and Computational Fluid Dynamics (CFD) methods. The application areas to be surveyed include fixed wing aircraft, turbomachinery, and rotary wing aircraft. The objective of the authors in the present paper, together with a companion paper on requirements, is to lay out a path for a High Performance Computing (HPC) based next generation comprehensive rotorcraft analysis. From this survey of the key technologies in other application areas it is possible to identify the critical technology gaps that stem from unique rotorcraft requirements.

Datta, Anubhav↗

NASA Blazes a Different Path to Energy-Efficient Supercomputing

For years, NASA had a very straightforward process for replacing high-performance computing hardware: over a three-year period, when it became more expensive to operate an older suite of hardware than it did to replace it with new products that could accomplish the same work, we simply replaced the old hardware. For NASA’s High-End Computing Capability (HECC) Project, that process changed when we reached the limits of our facility’s power, cooling, and floor-loading capacity, becoming a strategy of decommissioning the least productive hardware and replacing it with more capable counterparts. The impact was that we provided our users with less supercomputing capability than we would have without the limitations. Additionally, with 25% of our total power consumption going to cool our systems and 50,000 gallons of water per day being evaporated, we wanted a solution that would expand our compute facility while being sensitive to the impact on our environment.

Thigpen, William↗

Merlin - Massively parallel heterogeneous computing

Hardware and software for Merlin, a new kind of massively parallel computing system, are described. Eight computers are linked as a 300-MIPS prototype to develop system software for a larger Merlin network with 16 to 64 nodes, totaling 600 to 3000 MIPS. These working prototypes help refine a mapped reflective memory technique that offers a new, very general way of linking many types of computer to form supercomputers. Processors share data selectively and rapidly on a word-by-word basis. Fast firmware virtual circuits are reconfigured to match topological needs of individual application programs. Merlin's low-latency memory-sharing interfaces solve many problems in the design of high-performance computing systems. The Merlin prototypes are intended to run parallel programs for scientific applications and to determine hardware and software needs for a future Teraflops Merlin network.

Wittie, Larry↗

An Analysis of Failure Handling in Chameleon, A Framework for Supporting Cost-Effective Fault Tolerant Services

The desire for low-cost reliable computing is increasing. Most current fault tolerant computing solutions are not very flexible, i.e., they cannot adapt to reliability requirements of newly emerging applications in business, commerce, and manufacturing. It is important that users have a flexible, reliable platform to support both critical and noncritical applications. Chameleon, under development at the Center for Reliable and High-Performance Computing at the University of Illinois, is a software framework. for supporting cost-effective adaptable networked fault tolerant service. This thesis details a simulation of fault injection, detection, and recovery in Chameleon. The simulation was written in C++ using the DEPEND simulation library. The results obtained from the simulation included the amount of overhead incurred by the fault detection and recovery mechanisms supported by Chameleon. In addition, information about fault scenarios from which Chameleon cannot recover was gained. The results of the simulation showed that both critical and noncritical applications can be executed in the Chameleon environment with a fairly small amount of overhead. No single point of failure from which Chameleon could not recover was found. Chameleon was also found to be capable of recovering from several multiple failure scenarios.

Haakensen, Erik Edward↗

The Automated Instrumentation and Monitoring System (AIMS): Design and Architecture

Whether a researcher is designing the 'next parallel programming paradigm', another 'scalable multiprocessor' or investigating resource allocation algorithms for multiprocessors, a facility that enables parallel program execution to be captured and displayed is invaluable. Careful analysis of such information can help computer and software architects to capture, and therefore, exploit behavioral variations among/within various parallel programs to take advantage of specific hardware characteristics. A software tool-set that facilitates performance evaluation of parallel applications on multiprocessors has been put together at NASA Ames Research Center under the sponsorship of NASA's High Performance Computing and Communications Program over the past five years. The Automated Instrumentation and Monitoring Systematic has three major software components: a source code instrumentor which automatically inserts active event recorders into program source code before compilation; a run-time performance monitoring library which collects performance data; and a visualization tool-set which reconstructs program execution based on the data collected. Besides being used as a prototype for developing new techniques for instrumenting, monitoring and presenting parallel program execution, AIMS is also being incorporated into the run-time environments of various hardware testbeds to evaluate their impact on user productivity. Currently, the execution of FORTRAN and C programs on the Intel Paragon and PALM workstations can be automatically instrumented and monitored. Performance data thus collected can be displayed graphically on various workstations. The process of performance tuning with AIMS will be illustrated using various NAB Parallel Benchmarks. This report includes a description of the internal architecture of AIMS and a listing of the source code.

Yan, Jerry C.↗

Somatic Mutation in Mice on the International Space Station (ISS): Guanine Substitution Suggests Link to Cancer Risk

We conducted comprehensive analysis of single nucleotide somatic mutations in mice exposed to microgravity and other factors aboard the International Space Station (ISS), using data archived in GeneLab. Animals in the experimental cohort consisted of mice that spent 37 days on the ISS within the Rodent Habitat. Ground control animals consisted of mice of identical age, sex, strain, in a terrestrial Rodent Habitat controlled for temperature, humidity and carbon dioxide levels, to match ISS conditions as closely as possible. RNA extracted from eye, liver, skeletal muscle, and kidney tissue specimens was subjected to next-generation sequencing to acquire primary data. Our analysis employed cutting-edge software developed at NASA Ames Research Center, executed on the NASA Ames Supercomputer and on another high-performance computer, for accurate variant calling of single point mutations. ISS-flown mice exhibited a notably heightened level of somatic mutation compared to control mice. The degree of somatic mutation correlated with the degree of gene expression across the four tissue types, i.e., the greatest rate of mutation accumulation was seen in highly expressed genes. We discovered that guanine substitutions were the most common type of somatic mutation. This observation is consistent with the hypothesis that DNA mutation events stem from reactive oxygen/nitrogen/chlorine species-mediated guanine oxidation induced by the spaceflight environment. Since guanine oxidation is a prominent feature of the DNA mutation landscape that accompanies malignant transformation, our findings suggest a possible link between the spaceflight environment and cancer risk that is independent of radiation carcinogenesis.

ISS↗

Electra: A Modular-Based Expansion of NASA's Supercomputing Capability

NASA has increasingly relied on high-performance computing (HPC) re- sources for computational modeling, simulation, and data analysis to meet the science and engineering goals of its missions in space exploration, aeronautics, and Earth and space science. The NASA Advanced Supercomputing (NAS) Division at Ames Research Center in Silicon Valley, Calif., hosts NASA’s premier supercomputing resources, integral to achieving and enhancing the success of the agency’s missions. NAS provides a balanced environment, funded under the High-End Computing Capability (HECC) project, comprised of world-class supercomputers, including its flagship distributed-memory cluster, Pleiades; high-speed networking; and massive data storage facilities, along with multi-disciplinary support teams for user support, code porting and optimization, and large-scale data analysis and scientific visualization. However, as scientists have increased the fidelity of their simulations and engineers are conducting larger parameter-space studies, the requirements for supercomputing resources have been growing by leaps and bounds. With the facility housing the HECC systems reaching its power and cooling capacity, NAS undertook a prototype project to investigate an alternative approach for housing supercomputers. Modular supercomputing, or container-based computing, is an innovative concept for expanding NASA’s HPC capabilities. With modular supercomputing, additional containers—similar to portable storage pods—can be connected together as needed to accommodate the agency’s ever-increasing demand for computing resources. In addition, taking advantage of the local weather permits the use of cooling technologies that would additionally save energy and reduce annual water usage. The first stage of NASA’s Modular Supercomputing Facility (MSF) prototype, which resulted in a 1,000 square-foot module on a concrete pad with room for 16 compute racks, was completed in Fall 2016 and an SGI (now HPE) computer system, named Electra, was deployed there in early 2017. Cooling is performed via an evaporative system built into the module, and preliminary experience shows a Power Usage Effectiveness (PUE) measurement of 1.03. Electra achieved over a petaflop on the LINPACK benchmark, sufficient to rank number 96 on the November 2016 TOP500 list [14]. The system consists of 1,152 InfiniBand-connected Intel Xeon Broadwell-based nodes. Its users access their files on a facility-wide file system shared by all HECC compute assets via Mellanox MetroX InfiniBand extenders, which connect the Electra fabric to Lustre routers in the primary facility over fiber-optic links about 900 feet long. The MSF prototype has exceeded expectations and is serving as a blueprint for future expansions. In the remainder of this chapter, we detail how modular data center technology can be used to expand an existing compute resource. We begin by describing NASA’s requirements for supercomputing and how resources were provided prior to the integration of the Electra module-based system.

Biswas, Rupak↗

By Hand or Not By-Hand: A Case Study of Alternative Approaches to Parallelize CFD Applications

While parallel processing promises to speed up applications by several orders of magnitude, the performance achieved still depends upon several factors, including the multiprocessor architecture, system software, data distribution and alignment, as well as the methods used for partitioning the application and mapping its components onto the architecture. The existence of the Gorden Bell Prize given out at Supercomputing every year suggests that while good performance can be attained for real applications on general purpose multiprocessors, the large investment in man-power and time still has to be repeated for each application-machine combination. As applications and machine architectures become more complex, the cost and time-delays for obtaining performance by hand will become prohibitive. Computer users today can turn to three possible avenues for help: parallel libraries, parallel languages and compilers, interactive parallelization tools. The success of these methodologies, in turn, depends on proper application of data dependency analysis, program structure recognition and transformation, performance prediction as well as exploitation of user supplied knowledge. NASA has been developing multidisciplinary applications on highly parallel architectures under the High Performance Computing and Communications Program. Over the past six years, the transition of underlying hardware and system software have forced the scientists to spend a large effort to migrate and recede their applications. Various attempts to exploit software tools to automate the parallelization process have not produced favorable results. In this paper, we report our most recent experience with CAPTOOL, a package developed at Greenwich University. We have chosen CAPTOOL for three reasons: 1. CAPTOOL accepts a FORTRAN 77 program as input. This suggests its potential applicability to a large collection of legacy codes currently in use. 2. CAPTOOL employs domain decomposition to obtain parallelism. Although the fact that not all kinds of parallelism are handled may seem unappealing, many NASA applications in computational aerosciences as well as earth and space sciences are amenable to domain decomposition. 3. CAPTOOL generates code for a large variety of environments employed across NASA centers: MPI/PVM on network of workstations to the IBS/SP2 and CRAY/T3D.

Yan, Jerry C.↗

Neutron Characterization for Additive Manufacturing

Oak Ridge National Laboratory (ORNL) is leveraging decades of experience in neutron characterization of advanced materials together with resources such as the Spallation Neutron Source (SNS) and the High Flux Isotope Reactor (HFIR) shown in Fig. 1 to solve challenging problems in additive manufacturing (AM). Additive manufacturing, or three-dimensional (3-D) printing, is a rapidly maturing technology wherein components are built by selectively adding feedstock material at locations specified by a computer model. The majority of these technologies use thermally driven phase change mechanisms to convert the feedstock into functioning material. As the molten material cools and solidifies, the component is subjected to significant thermal gradients, generating significant internal stresses throughout the part (Fig. 2). As layers are added, inherent residual stresses cause warping and distortions that lead to geometrical differences between the final part and the original computer generated design. This effect also limits geometries that can be fabricated using AM, such as thin-walled, high-aspect- ratio, and overhanging structures. Distortion may be minimized by intelligent toolpath planning or strategic placement of support structures, but these approaches are not well understood and often "Edisonian" in nature. Residual stresses can also impact component performance during operation. For example, in a thermally cycled environment such as a high-pressure turbine engine, residual stresses can cause components to distort unpredictably. Different thermal treatments on as-fabricated AM components have been used to minimize residual stress, but components still retain a nonhomogeneous stress state and/or demonstrate a relaxation-derived geometric distortion. Industry, federal laboratory, and university collaboration is needed to address these challenges and enable the U.S. to compete in the global market. Work is currently being conducted on AM technologies at the ORNL Manufacturing Demonstration Facility (MDF) sponsored by the DOE's Advanced Manufacturing Office. The MDF is focusing on R&D of both metal and polymer AM pertaining to in-situ process monitoring and closed-loop controls; implementation of advanced materials in AM technologies; and demonstration, characterization, and optimization of next-generation technologies. ORNL is working directly with industry partners to leverage world-leading facilities in fields such as high performance computing, advanced materials characterization, and neutron sciences to solve fundamental challenges in advanced manufacturing. Specifically, MDF is leveraging two of the world's most advanced neutron facilities, the HFIR and SNS, to characterize additive manufactured components.

Watkins, Thomas↗

Parallel Visualization Co-Processing of Overnight CFD Propulsion Applications

An interactive visualization system pV3 is being developed for the investigation of advanced computational methodologies employing visualization and parallel processing for the extraction of information contained in large-scale transient engineering simulations. Visual techniques for extracting information from the data in terms of cutting planes, iso-surfaces, particle tracing and vector fields are included in this system. This paper discusses improvements to the pV3 system developed under NASA's Affordable High Performance Computing project.

Edwards, David E.↗