Search NASA⌕ Search

SEARCH · Search NASA

Results for “High Performance Computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32

NASA GRC ICME Schema for Materials Data Management: An Executive Summary

Integrated Computational Materials Engineering (ICME) has received a growing emphasis in attention due its potential impact on rapid material design, reduction in cost and time to market for new applications, and the promise of ‘fit-for-purpose’ materials coupled with recent advances in high performance computing and material characterization tools. However, for an organization to implement ICME practices for material discovery and design, a series of both technical and cultural challenges must be overcome to foster an environment that enables efficient, traceable, and predictive multiscale simulations of material behavior to enable virtual design of materials. In 2016, NASA sponsored a 2040 Vision study to define the potential 25-year future state required for integrated multiscale modeling of materials and systems to improve both the associated time and cost for aerospace and aeronautical innovation. The study envisions a cyber-physical-social ecosystem of experimentally validated computational models, tools, and techniques, along with the associated digital tapestry, that can enable rapid, optimized, ‘fit-for-purpose’ design of materials, components, and systems. A key requirement for such an ecosystem is the development of a robust information management system for materials across their full lifecycle, including material pedigree, experimental (real) and virtual (simulation) data, developed material models, and the implementation of models in engineering applications, such that process-structure-property-performance relationships can be established, thereby enabling the virtual design and optimization of materials. Such an information management system must be able to effectively capture: i) material information at each length scale; ii) test data and analysis; iii) associated material models; and iv) material and model deployment in engineering applications. These systems must also provide traceability between experimental and virtual representations of the material to ensure, when appropriate, the material digital twin is maintained. Additionally, this robust material information management system must be able to seamlessly connect with both commercial and an organization’s in-house software tools, be they analysis tools, other material databases, product lifecycle management (PLM) or simulation data management (SDM) tools, etc., such that automation of the design and analysis of a material across multiple length scales is possible. In this paper, an executive summary of the NASA GRC ICME Schema for materials information management is presented. The database best practices and schema design philosophy specifically for ICME materials data management and an overview description of each element in the schema is given, along with its associated role in an ICME workflow. Additionally, auxiliary tools that interact with the database and provide judicious automation with regards to importing, exporting, and analyzing materials data are presented. Such tools are critical to an ICME ecosystem, not only for their role in enabling optimization, but also in relieving users of tedious manual tasks, thus helping to promote adoption and combat the cultural challenges organizations face in enabling ICME.

Materials↗

ChatHPC: Building the Foundations for a Productive and Trustworthy AI-Assisted HPC Ecosystem

ChatHPC democratizes large language models for the high-performance computing (HPC) community by providing the infrastructure, ecosystem, and knowledge needed to apply modern generative AI technologies to rapidly create specific capabilities for critical HPC components while using relatively modest computational resources. Our divide-and-conquer approach focuses on creating a collection of reliable, highly specialized, and optimized AI assistants for HPC based on the cost-effective and fast Code Llama fine-tuning processes and expert supervision. We target major components of the HPC software stack, including programming models, runtimes, I/O, tooling, and math libraries. Thanks to AI, ChatHPC provides a more productive HPC ecosystem by boosting important tasks related to portability, parallelization, optimization, scalability, and instrumentation, among others. With relatively small datasets (on the order of KB), the AI assistants, which are created in a few minutes by using one node with two NVIDIA H100 GPUs and the ChatHPC library, can create new capabilities with Meta’s 7-billion parameter Code Llama base model to produce high-quality software with a level of trustworthiness of up to 90% higher than the 1.8-trillion parameter OpenAI ChatGPT-4o model for critical programming tasks in the HPC software stack.

Young, Aaron [ORNL] (ORCID:0000000254484667)↗

Space-based crystal growth and thermocapillary flow

The demand for larger crystals is increasing especially in applications associated with the electronic industry, where large and pure electronic crystals (notably silicon) are the essential material to make high-performance computer chips. Crystal growth under weightless conditions has been considered an ideal way to produce bigger and hopefully better crystals. One technique which may benefit from a microgravity environment is the float-zone crystal-growth process, a containerless method for producing high-quality electronic material. In this method, a rod of material to be refined is moved slowly through a heating device which melts a portion of it. Ideally, as the melt resolidifies it does so as a single crystal which is then used as substrate for building microelectronic devices. The possibility of contamination by contact with other material is reduced because of the 'float' configuration. However, since the weight of the material contained in the zone is supported by the surface-tension force, the size of the resulting crystal is limited in Earth-based productions; in fact, some materials have properties which prevent this process from being used to manufacture crystals of reasonable size. Consequently, there has been a great deal of interest in exploiting the microgravity environment of space to grow larger size crystals of electronic material using the float-zone method. In addition to allowing larger crystals to be grown, a microgravity environment would also significantly reduce the magnitude of convection induced by buoyancy forces during the melting state. This type of convection was once thought to be at least partially responsible for the presence of undesirable nonuniformities--called striations--in material properties observed in float-zone material. However, past experiments on crystal growth under weightless conditions found that even with the absence of gravity, the float-zone method sometimes still results striations. It is believed that another mechanism is playing a dominant role in the microgravity environment.

Shen, Yong-Hong↗

An Assessment of the State-of-the-art in Multidisciplinary Aeromechanical Analyses

This paper presents a survey of the current state-of-the-art in multidisciplinary aeromechanical analyses which integrate advanced Computational Structural Dynamics (CSD) and Computational Fluid Dynamics (CFD) methods. The application areas to be surveyed include fixed wing aircraft, turbomachinery, and rotary wing aircraft. The objective of the authors in the present paper, together with a companion paper on requirements, is to lay out a path for a High Performance Computing (HPC) based next generation comprehensive rotorcraft analysis. From this survey of the key technologies in other application areas it is possible to identify the critical technology gaps that stem from unique rotorcraft requirements.

Datta, Anubhav↗

NASA Blazes a Different Path to Energy-Efficient Supercomputing

For years, NASA had a very straightforward process for replacing high-performance computing hardware: over a three-year period, when it became more expensive to operate an older suite of hardware than it did to replace it with new products that could accomplish the same work, we simply replaced the old hardware. For NASA’s High-End Computing Capability (HECC) Project, that process changed when we reached the limits of our facility’s power, cooling, and floor-loading capacity, becoming a strategy of decommissioning the least productive hardware and replacing it with more capable counterparts. The impact was that we provided our users with less supercomputing capability than we would have without the limitations. Additionally, with 25% of our total power consumption going to cool our systems and 50,000 gallons of water per day being evaporated, we wanted a solution that would expand our compute facility while being sensitive to the impact on our environment.

Thigpen, William↗

Accelerating laser ray tracing in high fidelity physics simulations of laser melting using squeeze U-net

Laser melting is a core component of the ongoing industrial revolution, dubbed Industry 4.0, as lasers facilitate fast and precise melting and fusion in advanced manufacturing. There is a strong need to optimize the laser process using simulations. However, this has proven challenging as high fidelity simulations are needed for predictive modeling and this is currently prohibitively expensive even when run on hundreds of processors on high performance computers. The challenge is capturing complex physics of laser material interaction, fluid dynamics, thermal physics and material phase transformations at various length and time scales. To close this technological gap, we modified a squeeze U-net to accelerate the laser ray tracing component of such high fidelity models by ~4x–40x while preserving the core physics principle of conservation of energy with 97% accuracy. This approach enables the accurate modeling of global laser energy absorption as a function of local surface temperatures and complex surface topologies, which govern the reflection directions and energy losses of laser rays upon interacting with the material surface.

Computer science↗

Automation for Grid Interconnected Laboratory Emulation

As computational capabilities improve, digital twins are becoming vital for evaluating equipment realistically in laboratories. This paper outlines a digital twin architecture for the power grid, employing electromagnetic transient (EMT) simulation alongside real-time simulation of power hardware and hierarchical control systems. EMT simulation occurs on a high-performance computing server for scalability. Additionally, the paper describes a workflow and real-time data streaming software facilitating connectivity among EMT simulation, hierarchical control systems, and power hardware. This software enables automated equipment connectivity in the laboratory for realistic evaluations, aiding in identifying necessary upgrades for both equipment control systems and the power grid.

Marthi, Phani Ratna Vanamali [ORNL] (ORCID:0000000↗

Elastic Stochastic Full Waveform Inversion (eSFWI)

This collaboration between Lawrence Livermore National Security, LLC (LLNS) as manager and operator of Lawrence Livermore National Laboratory (LLNL) and Chevron USA Inc., acting through its Chevron Technical Center division, aimed at developing next-generation computational methods for the Elastic Stochastic Full Waveform Inversion (eSFWI). Seismic imaging is heavily used in the oil and gas industry for identifying and operating subsurface reservoirs. Improved seismic imaging methods can improve productivity, lower costs, and improve operational and environmental safety. This CRADA demonstrated that new high-performance computing (HPC) architectures being rolled out over the next five years can enable unprecedented seismic imaging resolution when using eSFWI techniques to process active seismic data. An open-source computational mini-application was developed, capable of demonstrating near-peak performance for eSFWI algorithms on CPU and GPU enabled HPC platforms. Performance was demonstrated on LLNL HPC systems such as Lassen, as well as on Chevron systems. This project benefited Chevron USA Inc. by demonstrating the potential computational efficiency of their full waveform inversion capabilities used to characterize oil/gas reservoirs, which in turn benefits the public through potential increases in capabilities to perform analysis of leasing sites.

04 OIL SHALES AND TAR SANDS↗

Report of the 2025 Workshop on Next-Generation Ecosystems for Scientific Computing: Harnessing Community, Software, and AI for Cross-Disciplinary Team Science

This report summarizes insights from the 2025 Workshop on Next-Generation Ecosystems for Scientific Computing: Harnessing Community, Software, and AI for Cross-Disciplinary Team Science, which convened more than 40 experts from national laboratories, academia, industry, and community organizations to chart a path toward more powerful, sustainable, and collaborative scientific software ecosystems. To address urgent challenges at the intersection of high-performance computing (HPC), AI, and scientific software, participants envisioned agile, robust ecosystems built through socio-technical co-design—the intentional integration of social and technical components as interdependent parts of a unified strategy. This approach combines advances in AI, HPC, and software with new models for cross-disciplinary collaboration, training, and workforce development. Key recommendations include building modular, trustworthy AI-enabled scientific software systems; enabling scientific teams to integrate AI systems into their workflows while preserving human creativity, trust, and scientific rigor; and creating innovative training pipelines that keep pace with rapid technological change. Pilot projects were identified as near-term catalysts, with initial priorities focused on hybrid AI/HPC infrastructure, cross-disciplinary collaboration and pedagogy, responsible AI guidelines, and prototyping of public-private partnerships. This report presents a vision of next-generation ecosystems for scientific computing where AI, software, hardware, and human expertise are interwoven to drive discovery, expand access, strengthen the workforce, and accelerate scientific progress.

97 MATHEMATICS AND COMPUTING↗

Merlin - Massively parallel heterogeneous computing

Hardware and software for Merlin, a new kind of massively parallel computing system, are described. Eight computers are linked as a 300-MIPS prototype to develop system software for a larger Merlin network with 16 to 64 nodes, totaling 600 to 3000 MIPS. These working prototypes help refine a mapped reflective memory technique that offers a new, very general way of linking many types of computer to form supercomputers. Processors share data selectively and rapidly on a word-by-word basis. Fast firmware virtual circuits are reconfigured to match topological needs of individual application programs. Merlin's low-latency memory-sharing interfaces solve many problems in the design of high-performance computing systems. The Merlin prototypes are intended to run parallel programs for scientific applications and to determine hardware and software needs for a future Teraflops Merlin network.

Wittie, Larry↗

Accelerating Bilevel Optimization With Hierarchical Many-Threaded Parallel Differential Evolution

Bilevel optimization is encountered in many relevant real-world applications. The main feature of this type of problem is that an upper-level optimization problem is constrained by a nested lower-level optimization problem. Because of this nested structure, bilevel problems (BLPs) are usually computationally expensive to solve. Differential evolution (DE) has demonstrated promising results in solving BLPs of relatively small scales. As the problem scale increases, the decision space becomes intrinsically larger, requiring a growing number of function evaluations for the method to work properly. In this context, heavy parallelization and high-performance computing techniques are indispensable to enable the resolution of more complex and challenging optimization problems. Hence, we propose a hierarchical many-threaded parallel DE approach for BLPs, where both levels are parallelized. The computational experiments demonstrate that the parallel implementation achieved runtime speeds ranging from 44 to 2559 times faster than the sequential version on a well-known scalable SMD benchmark test problem when executed on an NVIDIA A100 GPU. The findings indicate that the algorithm’s convergence is strongly influenced by the number of both upper- and lower-level generations. Moreover, the success of experiments with large-scale problems is closely linked to the choice of small population sizes.

Dufek, Amanda S↗

An Analysis of Failure Handling in Chameleon, A Framework for Supporting Cost-Effective Fault Tolerant Services

The desire for low-cost reliable computing is increasing. Most current fault tolerant computing solutions are not very flexible, i.e., they cannot adapt to reliability requirements of newly emerging applications in business, commerce, and manufacturing. It is important that users have a flexible, reliable platform to support both critical and noncritical applications. Chameleon, under development at the Center for Reliable and High-Performance Computing at the University of Illinois, is a software framework. for supporting cost-effective adaptable networked fault tolerant service. This thesis details a simulation of fault injection, detection, and recovery in Chameleon. The simulation was written in C++ using the DEPEND simulation library. The results obtained from the simulation included the amount of overhead incurred by the fault detection and recovery mechanisms supported by Chameleon. In addition, information about fault scenarios from which Chameleon cannot recover was gained. The results of the simulation showed that both critical and noncritical applications can be executed in the Chameleon environment with a fairly small amount of overhead. No single point of failure from which Chameleon could not recover was found. Chameleon was also found to be capable of recovering from several multiple failure scenarios.

Haakensen, Erik Edward↗

A Benchmark Suite for Evaluating Scientific AI Workloads on GPUs

AI applications have been steadily increasing in the allocation portfolio among leadership computing facilities. These applications depend on deep learning frameworks with hardware acceleration and underlying software systems. With the rapid development of applications, software stacks, and hardware devices, it is essential to evaluate the performance of core operations in AI workloads for direction of optimizations and procurement of next-generation high-performance computing (HPC) infrastructures. Currently, most benchmarks lack scientific AI workloads. So, we present DeepKernelBench and the experimental results of evaluating the benchmark suite for early observations and performance comparisons on datacenter GPUs using representative workloads for scientific AI, including Attentions, General matrix multiplications, Geometrics and Fourier neural operations.

Jin, Zheming [Advanced Micro Devices (AMD)]↗

The Automated Instrumentation and Monitoring System (AIMS): Design and Architecture

Whether a researcher is designing the 'next parallel programming paradigm', another 'scalable multiprocessor' or investigating resource allocation algorithms for multiprocessors, a facility that enables parallel program execution to be captured and displayed is invaluable. Careful analysis of such information can help computer and software architects to capture, and therefore, exploit behavioral variations among/within various parallel programs to take advantage of specific hardware characteristics. A software tool-set that facilitates performance evaluation of parallel applications on multiprocessors has been put together at NASA Ames Research Center under the sponsorship of NASA's High Performance Computing and Communications Program over the past five years. The Automated Instrumentation and Monitoring Systematic has three major software components: a source code instrumentor which automatically inserts active event recorders into program source code before compilation; a run-time performance monitoring library which collects performance data; and a visualization tool-set which reconstructs program execution based on the data collected. Besides being used as a prototype for developing new techniques for instrumenting, monitoring and presenting parallel program execution, AIMS is also being incorporated into the run-time environments of various hardware testbeds to evaluate their impact on user productivity. Currently, the execution of FORTRAN and C programs on the Intel Paragon and PALM workstations can be automatically instrumented and monitored. Performance data thus collected can be displayed graphically on various workstations. The process of performance tuning with AIMS will be illustrated using various NAB Parallel Benchmarks. This report includes a description of the internal architecture of AIMS and a listing of the source code.

Yan, Jerry C.↗

Creating Apptainer Workflows with Docker-Compose-like Utilities

Creating Apptainer Workflows with Docker-Compose-like Utilities In this presentation, I will explore the utilization of a tool called process-compose, inspired by docker-compose, to create Apptainer-based services. This approach allows for easy deployment and management of fully containerized applications on High Performance Computing (HPC) systems without requiring elevated privileges. Benefits to the Ecosystem: By incorporating process-compose and Apptainer, I aim to address several key challenges in the HPC ecosystem: Simplified Workflow Management: Process-compose provides a user-friendly interface for defining and managing complex containerized application services, reducing the setup time and lowering the barrier to entry for new users. Enhanced Portability: Apptainer ensures that containerized applications can run consistently across different HPC environments, promoting greater portability and reducing compatibility issues. Process-compose is also a single binary that does not need to be installed by admin level users. Community Driven Solutions: This approach aligns with the goals of the High Performance Software Foundation (HPSF) to advance community-driven solutions. By sharing our experiences and insights, I hope to foster collaboration and innovation within the HPC community. Increased Productivity: The combination of process-compose and Apptainer streamlines the serve deployment process, allowing researchers and developers to focus more on their scientific work rather than the intricacies of system or service administration. Through this presentation, attendees will gain valuable insights into the practical implementation of containerized workflows on HPC systems, learn about the benefits of using process-compose and Apptainer, and understand how these tools can contribute to a more efficient HPC ecosystem.

97 - MATHEMATICS AND COMPUTING↗

Somatic Mutation in Mice on the International Space Station (ISS): Guanine Substitution Suggests Link to Cancer Risk

We conducted comprehensive analysis of single nucleotide somatic mutations in mice exposed to microgravity and other factors aboard the International Space Station (ISS), using data archived in GeneLab. Animals in the experimental cohort consisted of mice that spent 37 days on the ISS within the Rodent Habitat. Ground control animals consisted of mice of identical age, sex, strain, in a terrestrial Rodent Habitat controlled for temperature, humidity and carbon dioxide levels, to match ISS conditions as closely as possible. RNA extracted from eye, liver, skeletal muscle, and kidney tissue specimens was subjected to next-generation sequencing to acquire primary data. Our analysis employed cutting-edge software developed at NASA Ames Research Center, executed on the NASA Ames Supercomputer and on another high-performance computer, for accurate variant calling of single point mutations. ISS-flown mice exhibited a notably heightened level of somatic mutation compared to control mice. The degree of somatic mutation correlated with the degree of gene expression across the four tissue types, i.e., the greatest rate of mutation accumulation was seen in highly expressed genes. We discovered that guanine substitutions were the most common type of somatic mutation. This observation is consistent with the hypothesis that DNA mutation events stem from reactive oxygen/nitrogen/chlorine species-mediated guanine oxidation induced by the spaceflight environment. Since guanine oxidation is a prominent feature of the DNA mutation landscape that accompanies malignant transformation, our findings suggest a possible link between the spaceflight environment and cancer risk that is independent of radiation carcinogenesis.

ISS↗

Electra: A Modular-Based Expansion of NASA's Supercomputing Capability

NASA has increasingly relied on high-performance computing (HPC) re- sources for computational modeling, simulation, and data analysis to meet the science and engineering goals of its missions in space exploration, aeronautics, and Earth and space science. The NASA Advanced Supercomputing (NAS) Division at Ames Research Center in Silicon Valley, Calif., hosts NASA’s premier supercomputing resources, integral to achieving and enhancing the success of the agency’s missions. NAS provides a balanced environment, funded under the High-End Computing Capability (HECC) project, comprised of world-class supercomputers, including its flagship distributed-memory cluster, Pleiades; high-speed networking; and massive data storage facilities, along with multi-disciplinary support teams for user support, code porting and optimization, and large-scale data analysis and scientific visualization. However, as scientists have increased the fidelity of their simulations and engineers are conducting larger parameter-space studies, the requirements for supercomputing resources have been growing by leaps and bounds. With the facility housing the HECC systems reaching its power and cooling capacity, NAS undertook a prototype project to investigate an alternative approach for housing supercomputers. Modular supercomputing, or container-based computing, is an innovative concept for expanding NASA’s HPC capabilities. With modular supercomputing, additional containers—similar to portable storage pods—can be connected together as needed to accommodate the agency’s ever-increasing demand for computing resources. In addition, taking advantage of the local weather permits the use of cooling technologies that would additionally save energy and reduce annual water usage. The first stage of NASA’s Modular Supercomputing Facility (MSF) prototype, which resulted in a 1,000 square-foot module on a concrete pad with room for 16 compute racks, was completed in Fall 2016 and an SGI (now HPE) computer system, named Electra, was deployed there in early 2017. Cooling is performed via an evaporative system built into the module, and preliminary experience shows a Power Usage Effectiveness (PUE) measurement of 1.03. Electra achieved over a petaflop on the LINPACK benchmark, sufficient to rank number 96 on the November 2016 TOP500 list [14]. The system consists of 1,152 InfiniBand-connected Intel Xeon Broadwell-based nodes. Its users access their files on a facility-wide file system shared by all HECC compute assets via Mellanox MetroX InfiniBand extenders, which connect the Electra fabric to Lustre routers in the primary facility over fiber-optic links about 900 feet long. The MSF prototype has exceeded expectations and is serving as a blueprint for future expansions. In the remainder of this chapter, we detail how modular data center technology can be used to expand an existing compute resource. We begin by describing NASA’s requirements for supercomputing and how resources were provided prior to the integration of the Electra module-based system.

Biswas, Rupak↗

Large-scale Multiphysics Simulations of Small Modular Reactors Operating in Natural Circulation

Thanks to the advancements in high-performance computing, advanced modeling and simulation have become crucial in driving the development and deployment of next-generation nuclear reactors, such as small modular reactors (SMRs). SMRs offer the promise of cost-effective baseload electricity production and improved safety, while addressing some of the challenges associated with large reactor designs, such as high capital costs and extended construction timelines. As part of the Exascale Computing Project, the large-scale multiphysics simulation of an entire SMR primary system has been achieved by combining computational fluid dynamics and neutronics. In addition to the successful demonstration of full-core SMR simulations, the current study integrated the impact of natural circulation into the system. Natural circulation is the primary mechanism driving coolant circulation in SMRs. The mass flow rate in the core depends on the core power, and a numerical model has been developed to predict it. The pressure drop caused by the helical coil steam generator was also accounted for by developing a pressure drop correlation based on high-fidelity large eddy simulation results, further improving prediction accuracy. In conclusion, the results of the study demonstrate that the implemented natural circulation model is effective in predicting the responses of SMR full-core multiphysics simulations.

ECP↗