Search NASA⌕ Search

SEARCH · Search NASA

Results for “High Performance Computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Studying CPU and memory utilization of applications on Fujitsu A64FX and Nvidia Grace Superchip

ARM-based manycore CPU architectures are well-positioned to provide the rising memory throughput requirements of modern data intensive scientific applications in High Performance Computing (HPC). The Fujitsu A64FX CPU platform is based on the ARM v8.2A architecture, and is the processor of the flagship Japanese supercomputer - "Fugaku", which was previously ranked as the #1 supercomputer in the world according to the Top500 list. The Nvidia Grace superchip features 144 Neoverse V2 cores based on the ARMv9 architecture with 4x128b SVE2, providing exceptional computational power. The chip supports up to 480GB of memory, making it ideal for AI, machine learning, and scientific computing workloads. In this paper, we conduct a thorough performance exploration of a variety of parallel bandwidth-sensitive benchmarks and applications compiled with the native Fujitsu compiler on a Fugaku A64FX compute node and ARM (LLVM) Compiler on an NVIDIA Grace superchip compute node, engaging all the computational cores per cluster using OpenMP multithreading (assuming the cores can drive the available bandwidth). Our ultimate goals are to study the resource utilization of scientific applications and benchmarks on A64FX and Grace superchip, considering graph application scenarios ( GAP Benchmark suite) and eleven appli- cation proxies from the Rodinia heterogeneous benchmark suite (considering domains such as Data Mining, Bioinformatics, Fluid Dynamics, Pattern Recognition, etc.). Through exhaustive performance monitoring, we quantify the resource utilization of diverse OpenMP-based HPC applications on both the Fujitsu A64FX and the Nvidia Grace Superchip platforms.

benchmarking, Performance Analysis, High performan↗

State of the art, gaps, and prospects in fusion materials theory and modelling

Advancing the theory and simulation of materials for fusion applications remains a key component of global roadmaps aimed at delivering much-needed fusion power. Especially as the drive for commercial application increases, prototypes must be designed against radiation damage before the relevant experimental data can be collected and cost reductions that are possible by testing materials in silico become even more important. Here, we summarise the state of the art as it emerged during the 7 th Fusion Materials Theory & Modelling Workshop that took place in 2024, with the aim to highlight present gaps and future directions for the fusion materials modelling community. Of particular interest were the effects of transmutations, chemical complexity with the development of novel alloys and interatomic potentials, advancements in modelling high-dose microstructures, comparison with experimental data and multiscale models for structural assessment relying on high-performance computing and virtual reality.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

ERF: Energy Research and Forecasting Model

High performance computing (HPC) architectures have undergone rapid development in recent years. As a result, established software suites face an ever increasing challenge to remain performant on and portable across modern systems. Many of the widely adopted atmospheric modeling codes cannot fully (or in some cases, at all) leverage the acceleration provided by General-Purpose Graphics Processing Units, leaving users of those codes constrained to increasingly limited HPC resources. Energy Research and Forecasting (ERF) is a regional atmospheric modeling code that leverages the latest HPC architectures, whether composed of only Central Processing Units (CPUs) or incorporating GPUs. ERF contains many of the standard discretizations and basic features needed to model general atmospheric dynamics. The modular design of ERF provides a flexible platform for exploring different physics parameterizations and numerical strategies. ERF is built on a state-of-the-art, well-supported, software framework (AMReX) that provides a performance portable interface and ensures ERF's long-term sustainability on next generation computing systems. This paper details the numerical methodology of ERF, presents results for a series of verification/validation cases, and documents ERF's performance on current HPC systems. The roughly 5× speed up of ERF (using GPUs) over Weather Research and Forecasting (CPUs only) for a 3D squall line test case highlights the significance of leveraging GPU acceleration.

17 WIND ENERGY↗

Computational technology for high-temperature aerospace structures

The status and some recent developments of computational technology for high-temperature aerospace structures are summarized. Discussion focuses on a number of aspects including: goals of computational technology for high-temperature structures; computational material modeling; life prediction methodology; computational modeling of high-temperature composites; error estimation and adaptive improvement strategies; strategies for solution of fluid flow/thermal/structural problems; and probabilistic methods and stochastic modeling approaches, integrated analysis and design. Recent trends in high-performance computing environment are described and the research areas which have high potential for meeting future technological needs are identified.

Noor, A. K.↗

Mineralogical, Elemental, and Tomographic Reconnaissance Investigation for CLPS (METRIC)

METRIC is a robotic science laboratory that can determine the mineralogy, elemental chemistry, micromorphology, and thermophysical properties of planetary regolith. The METRIC suite comprises METRIC XRD/F, an X-ray diffraction/X-ray fluorescence instrument that can determine the mineralogy and elemental chemistry of regolith samples; METRIC XCT, a micro X-ray computed tomography instrument that can be used to evaluate grain/crystallite sizes and textures; METRIC IRS, an imaging spectrometer mounted on a rover that can determine mineralogy and thermophysical properties at the landing site; and a pneumatic sample collection, processing, distribution system developed by Honeybee Robotics. The payload elements could be deployed on a static lander or a rover. Data returned from the METRIC payload would inform origin, formation, and evolution of rocky planetary bodies. METRIC XRD/F draws on heritage from the CheMin instrument on the Mars Science Laboratory (MSL) Curiosity rover [1], with a few important improvements. Like CheMin, METRIC XRD/F operates in transmission geometry and uses piezoelectric actuators on sample cells in a tuning fork geometry to induce convective grain motion of the regolith to create a randomly oriented powder. MSL CheMin uses an energy-sensitive CCD to collect XRD patterns and XRF spectra simultaneously from the same sample cell, resulting in qualitative XRF data. METRIC XRD/F uses two different sample cells, one optimized for XRD and one optimized for XRF, and a silicon drift detector to detect fluoresced X-rays. This improvement to the XRF capabilities provides quantitative geochemical data of major elements down to Z = 11 and allows for the detection of minor and trace elements that are critical for evaluating geologic evolution of the Moon (e.g., P and Th). Modest improvements to the XRD geometry and hardware allow for better angular resolution and the ability to distinguish between members of the pyroxene group. METRIC XCT uses the same geometry and much of the same hardware as METRIC XRD/F, where a CCD would capture images of a regolith sample in a 3 mm diameter sample tube that is rotated 360° in steps <1°. Image brightness can be used to infer compositional data, where brighter materials indicate a higher Z, much like scanning electron microscopy. Data from METRIC XCT complement those from METRIC XRD/F. Particle size, shape, and texture can provide petrologic and provenance information, whereas vesicle size and morphology in volcanic or impact melt lithologies can inform cooling rates. METRIC IRS is a hyperspectral thermal imager that can be mounted to a lander or rover to provide mineralogical data from the broader landing site and help determine whether the samples analyzed by METRIC XRD/F and XCT are representative. The METRIC IRS spectral range (8–14 μm) and resolution (10.8 cm-1) allow for quantitative mineralogy from modelling Reststrahlen bands of major rock-forming minerals (e.g., silicates, phosphates). Radiance cubes can be processed and modelled with an onboard high-performance computer to determine mineral abundances of plagioclase, high-Ca pyroxene, pigeonite, orthopyroxene, olivine, and glass. Regolith samples can be acquired, processed, and delivered to the X-ray instruments via multiple sample handling systems, but the pneumatic sampling systems developed by Honeybee Robotics [e.g., 2] are best suited for relatively low-cost missions that are being competed for the Moon (e.g., NASA’s Payloads and Research Investigations for the Surface of the Moon program). There are pneumatic sampling systems that collect surface material and other systems that pneumatically drill up to ~1 m below the surface, providing material that has not been space weathered and has not been affected by the lander’s exhaust. [1] Blake, D. F., Vaniman, D., Achilles, C., Anderson, R., Bish, D., et al. (2012). Space Sci. Rev. 170, 341-478. https://doi.org/10.1007/s11214-012-9905-1. [2] Zacny, K., Betts, B., Hedlund, M., Long, P., Gramlich, M., Tura, K., Chu, P., Jacob, A., Garcia, A. (2014). IEEE Aerospace Conference, 3-7 March 2014, Big Sky, MT, U.S.A.

X-ray diffraction↗

S&TR September 2025: Computing Grand Challenge Turns 20

Livermore’s Computing Grand Challenge Program enters its 20th year with more unclassified high-performance computing (HPC) power than ever before. This unique, peer-reviewed competition awards HPC allocations on top supercomputers to multidisciplinary teams with high-impact projects. The Grand Challenge encourages researchers to innovate, pushes scientific discovery to new heights, improves the Laboratory’s HPC capabilities, and extends HPC accessibility to collaborators. Awardees must adapt to successive generations of HPC hardware and learn to run simulations at scale. The feature article spotlights three Grand Challenge teams whose research broke new ground in key scientific pursuits—the essence of dark matter, explosion-generated seismic waves, and protein interactions linked to cancer—while underscoring the importance of academic partnerships and considering the program’s future.

07 ISOTOPE AND RADIATION SOURCES↗

Variable-Complexity Multidisciplinary Optimization on Parallel Computers

This report covers work conducted under grant NAG1-1562 for the NASA High Performance Computing and Communications Program (HPCCP) from December 7, 1993, to December 31, 1997. The objective of the research was to develop new multidisciplinary design optimization (MDO) techniques which exploit parallel computing to reduce the computational burden of aircraft MDO. The design of the High-Speed Civil Transport (HSCT) air-craft was selected as a test case to demonstrate the utility of our MDO methods. The three major tasks of this research grant included: development of parallel multipoint approximation methods for the aerodynamic design of the HSCT, use of parallel multipoint approximation methods for structural optimization of the HSCT, mathematical and algorithmic development including support in the integration of parallel computation for items (1) and (2). These tasks have been accomplished with the development of a response surface methodology that incorporates multi-fidelity models. For the aerodynamic design we were able to optimize with up to 20 design variables using hundreds of expensive Euler analyses together with thousands of inexpensive linear theory simulations. We have thereby demonstrated the application of CFD to a large aerodynamic design problem. For the predicting structural weight we were able to combine hundreds of structural optimizations of refined finite element models with thousands of optimizations based on coarse models. Computations have been carried out on the Intel Paragon with up to 128 nodes. The parallel computation allowed us to perform combined aerodynamic-structural optimization using state of the art models of a complex aircraft configurations.

Grossman, Bernard↗

Energy-efficient scientific computing using chemical reservoirs

The rapid growth of computing demands driven by scientific computing, data analytics, and artificial intelligence (AI) advancements has exposed the limitations of traditional digital processing systems. These systems are nearing physical energy barriers, making significant gains in energy efficiency increasingly unattainable. As we advance toward post-exascale computing, disruptive approaches are critical to overcoming these limitations. Among emerging analog solutions, biochemical computing offers a transformative path for achieving orders-of-magnitude improvements in energy efficiency. By leveraging the natural optimization capabilities of chemical reaction networks (CRNs), biochemical systems have the potential to meet high-performance computing needs through natural scalability. However, numerous challenges remain, including theoretical limitations in mapping computational problems to CRNs and practical barriers in implementing biochemical computing devices. In this paper, we present a framework for chemical computation using biochemical systems and introduce key components of our approach for energy-efficient scientific computing. We showcase the feasibility of this framework by solving a system of ordinary differential equations by emulating a chemical reservoir device, demonstrating its potential for addressing modern computing challenges. This work lays a foundational step toward harnessing the computational power of chemistry to design energy-efficient, scalable, high-performance next-generation computing systems.

Johnson, Connah G. M. [Pacific Northwest National ↗

Modeling Materials: Design for Planetary Entry, Electric Aircraft, and Beyond

NASA missions push the limits of what is possible. The development of high-performance materials must keep pace with the agency's demanding, cutting-edge applications. Researchers at NASA's Ames Research Center are performing multiscale computational modeling to accelerate development times and further the design of next-generation aerospace materials. Multiscale modeling combines several computationally intensive techniques ranging from the atomic level to the macroscale, passing output from one level as input to the next level. These methods are applicable to a wide variety of materials systems. For example: (a) Ultra-high-temperature ceramics for hypersonic aircraft-we utilized the full range of multiscale modeling to characterize thermal protection materials for faster, safer air- and spacecraft, (b) Planetary entry heat shields for space vehicles-we computed thermal and mechanical properties of ablative composites by combining several methods, from atomistic simulations to macroscale computations, (c) Advanced batteries for electric aircraft-we performed large-scale molecular dynamics simulations of advanced electrolytes for ultra-high-energy capacity batteries to enable long-distance electric aircraft service; and (d) Shape-memory alloys for high-efficiency aircraft-we used high-fidelity electronic structure calculations to determine phase diagrams in shape-memory transformations. Advances in high-performance computing have been critical to the development of multiscale materials modeling. We used nearly one million processor hours on NASA's Pleiades supercomputer to characterize electrolytes with a fidelity that would be otherwise impossible. For this and other projects, Pleiades enables us to push the physics and accuracy of our calculations to new levels.

Supercomputing↗

Extreme-scale workflows: A perspective from the JLESC international community

The Joint Laboratory for Extreme-Scale Computing (JLESC) focuses on software challenges in high-performance computing systems to meet the needs of today’s science campaigns, which often require large resources, consist of multiple tasks, and generate vast amounts of data. In this context, extreme-scale workflows have been the key factor in enabling scientific discoveries by helping scientists automate the dependencies and data exchanges between workflow tasks, instead of managing those manually. Here, in this paper, we present representative extreme-scale workflows and feature workflow systems developed by JLESC participating institutions. We present lessons learned while developing these tools, alongside with the open challenges and future research directions in the field of extreme-scale workflows.

97 MATHEMATICS AND COMPUTING↗

ChatMPI: LLM-Driven MPI Code Generation for HPC Workloads

The Message Passing Interface (MPI) standard plays a crucial role in enabling scientific applications for parallel computing and is an essential component in high-performance computing (HPC). However, implementing MPI code manually—especially applying a proper domain decomposition and communication pattern—is a challenging and error-prone task. We present ChatMPI, an AI assistant for MPI parallelization of sequential C codes. In our analysis, we focus on testing six essential HPC workloads, which are based on Basic Linear Algebra Subprograms levels 1, 2, and 3 as well as sparse, stencil, and iterative operations. We analyze the process of creating ChatMPI by using the ChatHPC library. This lightweight large language model (LLM)–based infrastructure enables HPC experts to efficiently create and supervise trustworthy AI capabilities for critical HPC software tasks. We study the data required for training (fine-tuning) ChatMPI to generate parallel codes that not only use MPI syntax correctly but also apply HPC techniques to reduce memory communication and maximize performance by using proper work decomposition. With a relatively small training dataset composed of a few dozen prompts and fewer than 15 minutes of fine-tuning on one node equipped with two NVIDIA H100 GPUs, ChatMPI elevates trustworthiness for MPI code generation of current LLMs (e.g., Code Llama, ChatGPT-4o and ChatGPT 5). Additionally, we evaluate the performance of the MPI codes generated by ChatMPI in comparison with the ones generated by ChatGPT-4o and ChatGPT-5. The codes generated by ChatMPI provide up to a 4 × boost in performance by using better problem decomposition, communication patterns, and HPC techniques (e.g., communication avoiding).

Valero Lara, Pedro [ORNL] (ORCID:0000000214794310)↗

AI-Driven Accelerated Inclusion Analysis for Energy Efficient Steelmaking (Final CRADA Report)

This was a collaborative effort between Lawrence Livermore National Security, LLC (LLNS) as manager and operator of Lawrence Livermore National Laboratory (LLNL) and ArcelorMittal USA Research LLC (“ArcelorMittal” as the Participant), to use scanning electron microscopy (SEM) images, computer vision and machine learning methods, and high-performance computing to accelerate the inclusion analysis process of liquid steel so that new methods can be used for near-real time process control on the shop floor.

36 MATERIALS SCIENCE↗

Properties of Electronic Materials

This final technical report summarizes the research conducted under DOE Grant DE-SC0002623, "Properties of Electronic Materials," led by Principal Investigator Shengbai Zhang at Rensselaer Polytechnic Institute. Over the 16-year period, the project employed first-principles computational methods to investigate the structural, electronic, and dynamic properties of a wide range of electronic materials, with applications in energy technologies, optoelectronics, and data storage. Key areas included topological insulators, phase-change materials, graphene and two-dimensional systems, perovskites for photovoltaics, defect engineering in semiconductors, kagome lattices, and ultrafast carrier dynamics. The research resulted in 115 peer-reviewed publications, advancing fundamental understanding of material behaviors at the atomic scale and contributing to innovations in renewable energy, memory devices, and quantum materials. Findings have implications for improving energy efficiency, developing lead-free solar cells, and enabling high-speed data processing. The work has trained numerous graduate students and postdocs, fostering the next generation of computational materials scientists. The original goals were to develop theoretical models and computational tools to predict and optimize electronic properties of materials for energy applications. All objectives were accomplished, with no major departures from planned methodologies. Challenges in computational scaling were addressed through access to high-performance computing resources.

36 MATERIALS SCIENCE↗

A Geometry Based Infra-Structure for Computational Analysis and Design

The computational steps traditionally taken for most engineering analysis suites (computational fluid dynamics (CFD), structural analysis, heat transfer and etc.) are: (1) Surface Generation -- usually by employing a Computer Assisted Design (CAD) system; (2) Grid Generation -- preparing the volume for the simulation; (3) Flow Solver -- producing the results at the specified operational point; (4) Post-processing Visualization -- interactively attempting to understand the results. For structural analysis, integrated systems can be obtained from a number of commercial vendors. These vendors couple directly to a number of CAD systems and are executed from within the CAD Graphical User Interface (GUI). It should be noted that the structural analysis problem is more tractable than CFD; there are fewer mesh topologies used and the grids are not as fine (this problem space does not have the length scaling issues of fluids). For CFD, these steps have worked well in the past for simple steady-state simulations at the expense of much user interaction. The data was transmitted between phases via files. In most cases, the output from a CAD system could go to Initial Graphics Exchange Specification (IGES) or Standard Exchange Program (STEP) files. The output from Grid Generators and Solvers do not really have standards though there are a couple of file formats that can be used for a subset of the gridding (i.e. PLOT3D data formats). The user would have to patch up the data or translate from one format to another to move to the next step. Sometimes this could take days. Specifically the problems with this procedure are:(1) File based -- Information flows from one step to the next via data files with formats specified for that procedure. File standards, when they exist, are wholly inadequate. For example, geometry from CAD systems (transmitted via IGES files) is defined as disjoint surfaces and curves (as well as masses of other information of no interest for the Grid Generator). This is particularly onerous for modern CAD systems based on solid modeling. The part was a proper solid and in the translation to IGES has lost this important characteristic. STEP is another standard for CAD data that exists and supports the concept of a solid. The problem with STEP is that a solid modeling geometry kernel is required to query and manipulate the data within this type of file. (2) 'Good' Geometry. A bottleneck in getting results from a solver is the construction of proper geometry to be fed to the grid generator. With 'good' geometry a grid can be constructed in tens of minutes (even with a complex configuration) using unstructured techniques. Adroit multi-block methods are not far behind. This means that a million node steady-state solution can be computed on the order of hours (using current high performance computers) starting from this 'good' geometry. Unfortunately, the geometry usually transmitted from the CAD system is not 'good' in the grid generator sense. The grid generator needs smooth closed solid geometry. It can take a week (or more) of interaction with the CAD output (sometimes by hand) before the process can begin. One way Communication. (3) One-way Communication -- All information travels on from one phase to the next. This makes procedures like node adaptation difficult when attempting to add or move nodes that sit on bounding surfaces (when the actual surface data has been lost after the grid generation phase). Until this process can be automated, more complex problems such as multi-disciplinary analysis or using the above procedure for design becomes prohibitive. There is also no way to easily deal with this system in a modular manner. One can only replace the grid generator, for example, if the software reads and writes the same files. Instead of the serial approach to analysis as described above, CAPRI takes a geometry centric approach. This makes the actual geometry (not a discretized version) accessible to all phases of the analysis. The connection to the geometry is made through an Application Programming Interface (API) and NOT a file system. This API isolates the top-level applications (grid generators, solvers and visualization components) from the geometry engine. Also this allows the replacement of one geometry kernel with another, without effecting these top-level applications. For example, if UniGraphics is used as the CAD package then Parasolid (UG's own geometry engine) can be used for all geometric queries so that no solid geometry information is lost in a translation. This is much better than STEP because when the data is queried, the same software is executed as used in the CAD system. Therefore, one analyzes the exact part that is in the CAD system. CAPRI uses the same idea as the commercial structural analysis codes but does not specify control. Software components of the CAD system are used, but the analysis suite, not the CAD operator, specifies the control of the software session. This also means that the license issues (may be) minimized and individuals need not have to know how to operate a CAD system in order to run the suite.

Haimes, Robert↗

INL Poster - Juan Barrera Salazar

Generation IV nuclear reactors introduce several advantages and benefits in terms of safety and efficiency when compared with their predecessors from previous generation. This is due, among many things, to the use of innovative forms of fuel and coolant, different from the conventional ones used in the last decades. Given that these upcoming designs utilize emerging technologies, the related instrumentation is also in the process of being developed; therefore, it is necessary to establish the sensitivity requirements and the effects of uncertainty on different properties of the components and elements of the reactor designs. This report presents the results of simulations that quantify the impacts of the uncertainties of four thermophysical properties of the refrigerant salt (LiF-BeF2) for the Kairos Power benchmark model (g-FHR) for steady state making use of the Sobol’ method through polynomial chaos surrogate modeling. The properties of the salt to which uncertainty was evaluated were density, dynamic viscosity, thermal conductivity and heat capacity. This study was carried out using the Griffin/Pronghorn multiphysics model under the computational resources of the Idaho National Laboratory (INL) High Performance Computing (HPC). The results indicate a weak dependence of the uncertainty of thermal conductivity on the quantities of core pressure drop and core outlet temperature.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN↗

Hot Chips and Hot Interconnects for High End Computing Systems

I will discuss several processors: 1. The Cray proprietary processor used in the Cray X1; 2. The IBM Power 3 and Power 4 used in an IBM SP 3 and IBM SP 4 systems; 3. The Intel Itanium and Xeon, used in the SGI Altix systems and clusters respectively; 4. IBM System-on-a-Chip used in IBM BlueGene/L; 5. HP Alpha EV68 processor used in DOE ASCI Q cluster; 6. SPARC64 V processor, which is used in the Fujitsu PRIMEPOWER HPC2500; 7. An NEC proprietary processor, which is used in NEC SX-6/7; 8. Power 4+ processor, which is used in Hitachi SR11000; 9. NEC proprietary processor, which is used in Earth Simulator. The IBM POWER5 and Red Storm Computing Systems will also be discussed. The architectures of these processors will first be presented, followed by interconnection networks and a description of high-end computer systems based on these processors and networks. The performance of various hardware/programming model combinations will then be compared, based on latest NAS Parallel Benchmark results (MPI, OpenMP/HPF and hybrid (MPI + OpenMP). The tutorial will conclude with a discussion of general trends in the field of high performance computing, (quantum computing, DNA computing, cellular engineering, and neural networks).

Saini, Subhash↗

Applied Information Systems Research Program (AISRP). Workshop 2: Meeting Proceedings

The Earth and space science participants were able to see where the current research can be applied in their disciplines and computer science participants could see potential areas for future application of computer and information systems research. The Earth and Space Science research proposals for the High Performance Computing and Communications (HPCC) program were under evaluation. Therefore, this effort was not discussed at the AISRP Workshop. OSSA's other high priority area in computer science is scientific visualization, with the entire second day of the workshop devoted to it.

Source record↗

Job Management Requirements for NAS Parallel Systems and Clusters

A job management system is a critical component of a production supercomputing environment, permitting oversubscribed resources to be shared fairly and efficiently. Job management systems that were originally designed for traditional vector supercomputers are not appropriate for the distributed-memory parallel supercomputers that are becoming increasingly important in the high performance computing industry. Newer job management systems offer new functionality but do not solve fundamental problems. We address some of the main issues in resource allocation and job scheduling we have encountered on two parallel computers - a 160-node IBM SP2 and a cluster of 20 high performance workstations located at the Numerical Aerodynamic Simulation facility. We describe the requirements for resource allocation and job management that are necessary to provide a production supercomputing environment on these machines, prioritizing according to difficulty and importance, and advocating a return to fundamental issues.

Saphir, William↗