Search NASASearch

SEARCH · Search NASA

Results for “High performance computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

PETSc/TAO Users Manual Revision 3.24

This manual describes the use of the Portable, Extensible Toolkit for Scientific Computation (PETSc) and the Toolkit for Advanced Optimization (TAO) for the numerical solution of partial differential equations (PDEs) and related problems on high-performance computers. PETSc/TAO is a suite of data structures and routines that provide the building blocks for implementing large-scale application codes on parallel (and serial) computers. PETSc uses the MPI standard for all distributed memory communication.

97 MATHEMATICS AND COMPUTING

PETSc/TAO Users Manual Revision 3.25

This manual describes the use of the Portable, Extensible Toolkit for Scientific Computation (PETSc) and the Toolkit for Advanced Optimization (TAO) for the numerical solution of partial differential equations (PDEs) and related problems on high-performance computers. PETSc/TAO is a suite of data structures and routines that provide the building blocks for implementing large-scale application codes on parallel (and serial) computers. PETSc uses the MPI standard for all distributed memory communication.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION

Accelerating discoveries at DIII-D with the Integrated Research Infrastructure

DIII-D research is being accelerated by leveraging high performance computing (HPC) and data resources available through the National Energy Research Scientific Computing Center (NERSC) Superfacility initiative. As part of this initiative, a high-resolution, fully automated, whole discharge kinetic equilibrium reconstruction workflow was developed that runs at the NERSC for most DIII-D shots in under 20 min. This has eliminated a long-standing research barrier and opened the door to more sophisticated analyses, including plasma transport and stability. These capabilities would benefit from being automated and executed within the larger Department of Energy Advanced Scientific Computing Research program’s Integrated Research Infrastructure (IRI) framework. The goal of IRI is to empower researchers to meld DOE’s world-class research tools, infrastructure, and user facilities seamlessly and securely in novel ways to radically accelerate discovery and innovation. For transport, we are looking at producing flux matched profiles and also using particle tracing to predict fast ion heat deposition from neutral beam injection before a shot takes place. Our starting point for evaluating plasma stability focuses on the pedestal limits that must be navigated to achieve better confinement. This information is meant to help operators run more effective experiments, so it needs to be available rapidly inside the DIII-D control room. So far this has been achieved by ensuring the data is available with existing tools, but as more novel results are produced new visualization tools must be developed. In addition, all of the high-quality data we have generated has been collected into databases that can unlock even deeper insights. This has already been leveraged for model and code validation studies as well as for developing AI/ML surrogates. The workflows developed for this project are intended to serve as prototypes that can be replicated on other experiments and can be run to provide timely and essential information for ITER, as well as next stage fusion power plants.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

CFD Simulation of Aerobic Gas Fermentation to Enable Commercial Conversion of CO 2 into Aquaculture and Animal Feed: Cooperative Research and Development (Final Report)

NovoNutrients’ fermentation technology uses energy from hydrogen to transform industrial CO2 emissions into premium animal feed ingredients and other valuable products. A single NovoNutrients’ commercial manufacturing plant will capture and convert over 200,000 tons/yr of CO2 into over 100,000 tons/yr of high-protein feed. Key to the rapid and widespread deployment of the technology is maximization of its productivity and energy efficiency. Robust, physically based computational models of the technology will significantly increase productivity and efficiency, accelerating NovoNutrients’ technology to manufacturing scale. NREL has unique capabilities for creating and running such computational models. NREL's existing aerobic bioreaction computational fluid dynamics (CFD) models will be adapted to NovoNutrients’ gas fermentation (CO2, H2, O2) technology. The multiphysics CFD simulations require thousands of high-performance computing (HPC) node hours to simulate the complex geometries and contents of NovoNutrients’ industrial bioreactors. The experimentally validated CFD models were used to identify optimally efficient and productive bioreactor designs and operating conditions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Extending SEER for Extreme Heterogeneity

Heterogeneous and multi-device nodes are increasingly common in high-performance computing and data centers, yet existing programming models often lack simple, transparent, and portable support for these diverse architectures. The main contribution of this work is the development of novel SEER capabilities to address this challenge by providing a descriptive programming model that allows applications to seamlessly leverage heterogeneous nodes across various device types. SEER uses efficient memory management and can select the proper device[s] depending on the computational cost of the applications. This is completely transparent to the programmer, thereby providing a highly productive programming environment. Integrating extreme heterogeneity into the SEER library as shown with the use of NVIDIA and AMD GPUs simultaneously allows it to expand and exploit the performance possibilities. Our analysis based on the well-known Conjugate Gradient algorithm reports accelerations above 1.5 × on computationally demanding steps of such an algorithm by using both architectures simultaneously.

Teranishi, Keita [ORNL] (ORCID:0000000166472690)

Evaluating the potential of disaggregated memory systems for HPC applications

Summary Disaggregated memory is a promising approach that addresses the limitations of traditional memory architectures by enabling memory to be decoupled from compute nodes and shared across a data center. Cloud platforms have deployed such systems to improve overall system memory utilization, but performance can vary across workloads. High‐performance computing (HPC) is crucial in scientific and engineering applications, where HPC machines also face the issue of underutilized memory. As a result, improving system memory utilization while understanding workload performance is essential for HPC operators. Therefore, learning the potential of a disaggregated memory system before deployment is a critical step. This paper proposes a methodology for exploring the design space of a disaggregated memory system. It incorporates key metrics that affect performance on disaggregated memory systems: memory capacity, local and remote memory access ratio, injection bandwidth, and bisection bandwidth, providing an intuitive approach to guide machine configurations based on technology trends and workload characteristics. We apply our methodology to analyze thirteen diverse workloads, including AI training, data analysis, genomics, protein, fusion, atomic nuclei, and traditional HPC bookends. Our methodology demonstrates the ability to comprehend the potential and pitfalls of a disaggregated memory system and provides motivation for machine configurations. Our results show that eleven of our thirteen applications can leverage injection bandwidth disaggregated memory without affecting performance, while one pays a rack bisection bandwidth penalty and two pay the system‐wide bisection bandwidth penalty. In addition, we also show that intra‐rack memory disaggregation would meet the application's memory requirement and provide enough remote memory bandwidth.

Ding, Nan

Integrating DOE ASCR Computing into HEPCloud through GlideinWMS

Fermilab's HEPCloud facility expands the laboratory's computing capacity by provisioning resources beyond the local grid, using GlideinWMS to deliver pilots to where experiments such as CMS and DUNE run. The High-Performance Computing (HPC) facilities of the DOE Office of Advanced Scientific Computing Research (ASCR) are a growing part of that pool. HEPCloud currently provisions NERSC over SSH, but NERSC is moving away from that path as it adopts multi-factor authentication and directs automated access to its Superfacility API and the DOE Integrated Research Infrastructure (IRI) APIs. Maintaining and extending access across the ASCR ecosystem now requires provisioning through these interfaces. This work adds new pilot submission paths to GlideinWMS for the NERSC Superfacility API, IRI, and Globus Compute. Each uses the provisioning model GlideinWMS already applies to batch resources, so experiments can run on ASCR computing resources without changes to their existing workflows. This work finally presents a comparison of the paths to guide which interfaces are best suited for different workflows.

Majumder, Meghanto [U. Houston (main)]

Energy Systems Integration Facility Stewardship Summary: Fiscal Year 2025

A summary of NLR's stewardship of the nationally unique Energy Systems Integration Facility (ESIF) highlighting performance metrics, capability upgrades, and examples of R&D impact. 2025 brought a national focus on energy and ESIF is meeting the moment. All eyes are on data centers and domestic manufacturing and bringing the benefits of artificial intelligence to power system planning and operations. In step with national priorities, ESIF is building out capabilities that advance secure, reliable, and affordable power. ESIF hosted 190 multidisciplinary research projects, 855 high-performance computer users, and collaborated with 81 partners from industry, academia, research, and federal agencies. These research projects resulted in an AI method for detecting high-impedance faults with 90% accuracy, a grid controls demonstration in Connecticut, power quality validation of CorePower's flagship inductor, and a cybersecurity assessment of potential rogue capabilities in digitally connected energy devices. Facility infrastructure improvements enhanced the thermal research network, the SCADA system, the cyber range, power hardware-in-the-loop testing, and more. With support from the U.S Department of Energy (DOE), the ESIF laboratories continue to deliver leading solutions for secure, reliable, and affordable power.

24 POWER TRANSMISSION AND DISTRIBUTION

PV Performance Modeling and Stakeholder Engagement (Final Technical Report)

This core capability project’s objective is to increase the value of photovoltaic (PV) performance models by improving their functionality, demonstrating, and quantifying their validity, and offering a wide range of stakeholder engagement opportunities. In FY22-24, we developed new and improved modeling algorithms and functions to represent PV performance more accurately in a variety of environments and conditions. The “Model parameter toolkit” was developed and includes functions to translate between different module temperature models, incidence angle modifier models, and single-diode models. A new modeling capability named “PV Atlas” was also developed leveraging Sandia’s High Performance Computing resources. This capability allows us to investigate several questions and provide climate-specific best practices and geographic data files; all these are hosted on an interactive website on Sandia’s GitHub and can be used for training, system optimization, or to provide best practices for uncertainty reduction. For model validation, we published high-quality PV performance, and weather data; these data are well documented, filtered, and processed for quality and include examples on how to run PV simulations. We also developed well documented, standardized methods for validating PV models and ran independent model validation and 2 blind modeling intercomparisons engaging with 49 organizations from 17 countries. We co-led and contributed to a growing, well documented and maintained suite of open-source functions for PV modeling (i.e., the pvlib-python) and we outreached to the PV modeling stakeholders via the PVPMC workshops and web resources. In addition, this project supported US representation and leadership for the International Energy Agency (IEA) PVPS Task 13; specifically, members of our team led and supported 3 subtasks on: 1) Best practices for the optimization of bifacial photovoltaic tracking, 2) Extreme weather events and their multiple impact on PV power plants: Risks, failure mechanisms and mitigation strategies, and 3) Best practice guidelines for the use of economic and technical Key Performance Indicators (KPIs). This project resulted in the publications of 14 peer reviewed journal papers, 37 conference presentations, 6 SAND reports, 5 public datasets and 6 new webpages on the PVPMC website. It supported the release of 13 pvlib-python versions where 28 enhancements were from this PV Performance Modeling project. We co-organized 5 PVPMC workshops in FY22-24 with the participation of 214 unique institutions and around 700 participants. The PVPMC website was redesigned, and its reliability was improved; it receives over 50,000 visitors/year from 202 unique countries.

14 SOLAR ENERGY

Leveraging AI for Productive and Trustworthy HPC Software: Challenges and Research Directions

We discuss the challenges and propose research directions for using AI to revolutionize the development of high-performance computing (HPC) software. AI technologies, in particular large language models, have transformed every aspect of software development. For its part, HPC software is recognized as a highly specialized scientific field of its own. We discuss the challenges associated with leveraging state-of-the-art AI technologies to develop such a unique and niche class of software and outline our research directions in the two US Department of Energy–funded projects for advancing HPC Software via AI: Ellora and Durban.

Teranishi, Keita [ORNL] (ORCID:0000000166472690)

District heating utilizing waste heat of a data center: High-temperature heat pumps

Data centers are energy-intensive facilities with substantial low-grade waste heat. High-temperature heat pumps can be critical in boosting the data center’s waste heat for district heating, improving the system-level energy efficiency of data centers, and reducing CO 2 emissions in district heating. This study built thermodynamic models to assess high-temperature heat pumps with six configurations using low global warming potential refrigerants to supply heat up to 120 °C. The heat pump configurations include single-stage or two-stage cycles with advanced components, such as internal heat exchanger, economizer, flash tank, or parallel compressor. The refrigerants include R1234ze(Z), R1233ed(E), R1224yd(Z), R600, and R600a, and R245fa is used as a reference. A case study was carried out to recover the waste heat from the Frontier high-performance computing data center and provide hot water for district heating at the US Department of Energy’s Oak Ridge National Laboratory campus. The optimized performance of high-temperature heat pumps is characterized with various effectiveness of internal heat exchangers, and the operating parameters of economizer or flash tank, as well as their combination. The results show that the configurations of two-stage cycles with internal heat exchanger + flash tank and internal heat exchanger + economizer/parallel-compressor provide the highest coefficient of performance under scenarios of the maximum allowable value and a fixed value (0.3) of the internal heat exchangers’ effectiveness, respectively. R1234ze(Z) and R600a are the most promising refrigerants, considering trade-offs between the coefficient of performance and the volumetric heating capacity. The single-stage cycle with internal heat exchanger + economizer/parallel-compressor using R1234ze(Z) is recommended for utilizing Fronter’s waste heat in district heating. A one mega-watt high-temperature heat pump will reduce 33,100–33,200 metric tons of CO2 emission annually, corresponding to 85.4 %–85.6 % of equivalent CO2 emissions from natural gas boilers. Here, this study provides good guidelines for designing and deploying high-temperature heat pumps to support sustainable data centers and decarbonize district heating in the US.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Packaging HEP Heterogeneous Mini-apps for Portable Benchmarking and Facility Evaluation on Modern HPCs

High Energy Physics (HEP) experiments are making increasing use of GPUs and GPU dominated High Performance Computer facilities. Both the software and hardware of these systems are rapidly evolving, creating challenges for experiments to make informed decisions as to where they wish to devote resources. In its first phase, the High Energy Physics Center for Computational Excellence (HEP-CCE) produced portable versions of a number of heterogeneous HEP mini-apps, such as p2r, FastCaloSim, Patatrack and the WireCell Toolkit, that exercise a broad range of GPU characteristics, enabling cross platform and facility benchmarking and evaluation. However, these miniapps still require a significant amount of manual intervention to deploy on a new facility. We present our work in developing turn-key deployments of these mini-apps, where by means of containerization and automated configuration and build techniques such as Spack, we are able to quickly test new hardware, software, environments and entire facilities with minimal user intervention, and then track performance metrics over time.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Packaging HEP Heterogeneous Mini-apps for Portable Benchmarking and Facility Evaluation on Modern HPCs

High Energy Physics (HEP) experiments are making increasing use of GPUs and GPU dominated High Performance Computer facilities. Both the software and hardware of these systems are rapidly evolving, creating challenges for experiments to make informed decisions as to where they wish to devote resources. In its first phase, the High Energy Physics Center for Computational Excellence (HEP-CCE) produced portable versions of a number of heterogeneous HEP mini-apps, such as \ptor, FastCaloSim, Patatrack and the WireCell Toolkit, that exercise a broad range of GPU characteristics, enabling cross platform and facility benchmarking and evaluation. However, these mini-apps still require a significant amount of manual intervention to deploy on a new facility. We present our work in developing turn-key deployments of these mini-apps, where by means of containerization and automated configuration and build techniques such as Spack, we are able to quickly test new hardware, software, environments and entire facilities with minimal user intervention, and then track performance metrics over time.

Atif, Mohammad [Brookhaven] (ORCID:000000026889770

Optimization of Scrap Melting Using an Electric Arc in Steel Manufacturing

Steel industry is crucial to the national economy and security. Around 67% of crude steel in the U.S is produced in electric arc furnaces (EAF), which is energy intensive. Around 140 EAFs operate in the U.S., consuming about 8.6x10 7 MMBtu/year of electricity. One of major challenges for EAFs includes maximizing the efficiency of the electrical energy provided in the form of electric arcs to melt various scrap mixes. To address this issue, a computational fluid dynamics (CFD) methodology is chosen to analyze scrap melting using the electric arc. Due to complex furnace phenomena and the wide variety of potential scenarios, high performance computing (HPC) is essential to yield comprehensive and detailed CFD analyses and systematic parametric studies for optimized EAF operation. The objectives are to 1) simulate scrap melting using electric arc, 2) evaluate electrode/arc position for optimum scrap melting and 3) establish reduced order model for CFD data-base for fast model calculation.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Viskores: Integrating Parallel Scientific Visualization Research into Applications

Viskores is a scientific visualization library that is the primary deployment of such algorithms to the parallel accelerated processors of modern DOE supercomputers. In this paper, we review the capabilities provided by Viskores and how these capabilities are leveraged by other software in the high-performance computing ecosystem. We discuss the Viskores data representation and pay particular attention to array management. Through this array management we describe how data is adapted between Viskores and other software along with strategies for converting dynamic, polymorphic objects to static representations better suited to GPU processing. We conclude with several examples of Viskores integrating with high-performance software that is used in production today.

Moreland, Ken [ORNL] (ORCID:0000000270513288)

Custom Accessors: Enabling Scalable Data Ingestion, (Re-)Organization, and Analysis on Distributed Systems

The emerging class of high velocity and high volume data analytic workflows comprise interwoven data ingestion, organization, and processing stages, with ingestion and organization steps often contributing comparable or even higher computational costs than actual processing steps. Since complex workflows consist of a variety of phases that view and use data differently, being able to construct efficient, scalable, distributed data structures (arrays, vectors, sets, maps, and multi-maps) is essential and requires custom methods to extend and shrink containers, analyze and position data, and, maintain globallyconsistent meta-data. In this paper, we propose a novel datastructure access paradigm based on the concept of Accessors. At a high level, accessors are customizable callable objects that can modify the behavior of insert, read, update, and delete operations for distributed containers while preserving atomicity guarantees. Accessors provide a very clean and natural way to implement a variety of programming patterns, e.g., conditional insertion/deletion and cascading computations, which would be otherwise hard (or even impossible) to express in parallel and distributed settings without using locks. We demonstrate the practicality and usefulness of our approach with two representative use cases and study the performance of these applications on a distributed High-Performance Computing system. Our analysis highlights that our proposed abstraction allows for an effective overlapping and concurrent execution of different workflow steps (e.g., data ingestion and analysis), which in a conventional analytics pipeline would execute sequentially, contributing cumulatively to the overall latency.

Castellana, Vito G. [BATTELLE (PACIFIC NW LAB)] (O

Robust wind farm layout optimization

Wake interactions in wind farms cause losses in annual energy production (AEP) on the order of 10%. Wind farm designers optimize the layout of the farm to mitigate wake losses, especially in the dominant site-specific wind directions. As wind turbines and wind farms grow in scale, optimization becomes more complex. Offshore wind farms regularly comprise more than 100 wind turbines and are characterized by complex boundaries due to shipping lanes, neighboring wind farms, and other constraints. Layout optimization methods are broadly split between gradient-based and gradient-free approaches. Gradient-based approaches can converge quickly and perform well for smaller, academic problems but are often sensitive to initial conditions and tuning parameters and require expert knowledge to use. On the other hand, gradient-free approaches can be more robust to problem complexities. We present a robust layout optimization approach based on a random search algorithm. The algorithm is intended for those who are not optimization experts and has few tuning parameters that need specification to achieve satisfactory results. Unlike off-the-shelf methods, which use generally available, non-domain-specific optimization routines that accept as inputs an optimization function and constraint definitions, this approach takes advantage of the relative computational costs of the different evaluations by evaluating cheaper computations first (boundary and minimum distance constraints) and running expensive AEP evaluations only if all other checks pass. Moreover, an outer genetic algorithm allows multiple solutions to evolve in parallel, enabling rapid solution development on high-performance computers. We discuss the relative ease of selecting necessary tuning parameters and demonstrate the efficacy of the genetic random search on a complex layout problem consisting of placing 70 turbines in a nonconvex and unconnected boundary region.

17 WIND ENERGY

Beyond microbial abundance: metadata integration enhances disease prediction in human microbiome studies

Multiple studies have highlighted the interaction of the human microbiome with physiological systems such as the gut, immune, liver, and skin, via key axes. Advances in sequencing technologies and high-performance computing have enabled the analysis of large-scale metagenomic data, facilitating the use of machine learning to predict disease likelihood from microbiome profiles. However, challenges such as compositionality, high dimensionality, sparsity, and limited sample sizes have hindered the development of actionable models. One strategy to improve these models is by incorporating key metadata from both the human host and sample collection/processing protocols. This remains challenging due to sparsity and inconsistency in metadata annotation and availability. In this paper, we introduce a machine learning-based pipeline for predicting human disease states by integrating host and protocol metadata with microbiome abundance profiles from 68 different studies, processed through a consistent pipeline. Our findings indicate that metadata can enhance machine learning predictions, particularly at higher taxonomic ranks like Kingdom and Phylum, though this effect diminishes at lower ranks. Our study leverages a large collection of microbiome datasets comprising 11,208 samples, therefore enhancing the robustness and statistical confidence of our findings. This work is a critical step toward utilizing microbiome and metadata for predicting diseases such as gastrointestinal infections, diabetes, cancer, and neurological disorders.

Mathematics and Computing