Search NASA⌕ Search

SEARCH · Search NASA

Results for “scalable performance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Programming approaches for scalability, performance, and portability of combustion physics codes

Here, this paper presents the process, strategy, and results associated with porting a typical combustion physics flow solver to current state-of-the-art and future massively-parallel computer architectures. Major focus is placed on the distinct algorithmic structure of these types of codes and how it can be integrated with modern programming paradigms for heterogeneous platforms (i.e., distributed many-core systems with accelerators). An end-to-end case study is presented that exemplifies the process in a generic manner, which then serves as a clear guide with respect to the strategy and best practices leading to a robust and adaptable framework that performs well, is durable over time, is portable, and requires minimal human-effort. This end is accomplished beginning with the use of a mature, validated, structured, multiblock code framework optimized for application of both Large Eddy Simulation (LES) and Direct Numerical Simulation (DNS). This code has been ported to a variety of platforms over the past decade, including most recently the Oak Ridge Leadership Computing Facility’s “Summit” Platform. The experience gained on these multiple platforms provides general insights and thus the results presented are not specific to any one code or platform other than the overarching trend toward distributed many-core systems with accelerators in order to move toward exascale performance. The resultant performance and scalability of the ported code is demonstrated on a real-world application; a state-of-the-art rotating detonation rocket engine simulation that matches the complex geometry and boundary conditions imposed as part of a companion experimental campaign.

97 MATHEMATICS AND COMPUTING↗

A Performant, Scalable Processing Pipeline for High‐Quality and FAIR Environmental Sensor Data

High-resolution environmental monitoring is necessary to record, understand, and predict biogeochemical and ecological changes particularly in coastal systems but brings significant challenges in processing and making rapidly available the resulting data. The COMPASS-FME project established a network of coastal observational sites across the Chesapeake Bay and western Lake Erie regions extensively instrumented with soil, vegetation, and weather sensors logging data every 15 min. Our data processing framework, written in R and completely open source, prioritizes rapid model-experiment iteration and makes biogeochemical data rapidly available for quality assurance/quality control, analysis, and model ingestion. This pipeline is distinguished by a standardized and modular approach to data curation, extensive metadata and documentation, and its high performance. These attributes combine to make biogeochemical data rapidly accessible across COMPASS-FME and the broader community. Flexible, powerful, and reproducible approaches to handling high-volume environmental data are crucial for accelerating biogeosciences research.

Pennington, Stephanie C. [Pacific Northwest Nation↗

ExaCA v2.0: A versatile, scalable, and performance portable cellular automata application for additive manufacturing solidification

The previously established ExaCA software for performance portable alloy grain structure simulation has been updated to better represent the solidification behavior during complex alloy processing conditions, such as those encountered during metal additive manufacturing (AM), and for improved performance and scalability. Here, an extension to the time–temperature history input data format and the core ExaCA algorithm to include an arbitrary number of melting and solidification events yielded improved prediction of texture for various melt pool geometries, expanding the range of AM-relevant conditions that can be accurately simulated. Improved heat transport process simulation coupling, including the creation of large raster datasets from single track time–temperature history data and in-memory coupling with the new, performance portable finite difference code Finch, were also demonstrated in example studies on the effect of multilayer AM microstructure predictions on hatch spacing and cell size, respectively. Additional new features are detailed and demonstrated, including the ability to perform simulations using various interfacial response function forms, execute simulations on state-of-the-art hardware, improved usability through post-processing versatility, and improved strong and weak scaling performance. The performance, physics, and versatility improvements demonstrated here will further enable large-scale studies on AM process–microstructure relationships that were not previously possible. Furthermore, the usability improvements and ability to run coupled AM process–microstructure simulations using the Finch-ExaCA workflow will facilitate broader use of this open-source software by the computational materials community.

36 MATERIALS SCIENCE↗

Toucan: A performance portable, scalable implementation of the DECA algorithm

In the field of additive manufacturing (AM), cellular automata (CA) is extensively used to simulate microstructural evolution during solidification. However, while traditional CA approaches are relatively fast, they still require a substantial number of time steps, are limited to moderate volumes, and are relatively difficult to improve through parallelism due to the highly localized nature of the solidification front. Here, to address these issues of time to solution and load balancing, we introduce Toucan, a parallel, performance-portable, and scalable code written in C++ with the Kokkos library that leverages the discrete event inspired cellular automata (DECA) algorithm to perform parallel-in-time (PinT) grain growth simulations. Toucan effectively mitigates load balancing issues by distributing the computational workload more evenly across processors, enhancing scalability and efficiency. We conduct both strong and weak scaling studies on up to 64 GPUs on the Frontier supercomputer, demonstrating that Toucan significantly outperforms the current state-of-the-art, time-stepped CA code, ExaCA, on both single and multi-GPU simulations. Even in AM-specific weak scaling scenarios, Toucan maintains near-ideal scaling, in contrast to the linear increase observed with ExaCA due to the moving laser raster pattern. This study highlights Toucan’s potential to transform microstructural simulations in AM by radically improving both efficiency and scalability over existing methods.

36 MATERIALS SCIENCE↗

Semantic Property Graph for Scalable Knowledge Graph Analytics

Graphs are a natural and fundamental representation to describe entities, relationships, activities, and evolution of complex systems. Many domains such as communication, citation, procurement, biology, social media, and transportation can be modeled as a set of entities and their relationships. Resource Description Framework (RDF) and Labeled Property Graph (LPG) are two of the most used data models to encode information in a graph. Both models are similar in terms of using basic graph elements such as nodes and edges but differ in terms of the modeling approach, expressibility, serialization, and target applications. RDF is a flexible data exchange model for expressing information about entities but it tends to a have high memory footprint and inefficient storage, which does not make it a natural choice to perform scalable graph analytics. In contrast, LPG has gained traction as a reliable model to perform scalable graph analytic tasks such as sub-graph matching, network alignment, and real-time knowledge graph query. It provides efficient storage, fast traversal, and flexibility to model various real-world domains. At the same time, the LPG lacks the support of a formal knowledge representation such as an ontology to provide automated knowledge inference. We propose Semantic Property Graph (SPG) as a logical projection of reified RDF into the LPG model. SPG continues to use RDF ontology to define the type hierarchy of the projected graph and validate it against a given ontology. We present a framework to convert reified RDF graphs into SPG using two different computing environments. We also present cloud-based graph migration capabilities using Amazon Web Services.

Purohit, Sumit↗

Enabling kilometer-scale E3SM land model simulation over North America: A new integrated framework solution

This study introduces a novel framework designed to enhance the performance, scalability, and portability of the kilometer-scale E3SM Land Model (km-ELM) within the E3SM modeling infrastructure. By seamlessly integrating cutting-edge data tools, we address existing challenges such as slow performance, limited scalability, and difficulties in software integration in current data-driven ELM simulation over large geographic areas. Our innovative approach leverages the KiloCraft data toolkit to generate unified inputs for simulations ranging from a single-cite case, to a 72,083-cell regional case to a continental configuration encompassing 21.6 million land grid cells at a 1 km × 1 km resolution. We conduct extensive strong- and weak-scaling experiments on three state-of-the-art supercomputers, utilizing up to 100,800 CPU cores across 2400 compute nodes to evaluate end-to-end metrics including wall-clock time, simulation-years-per-day (SYPD), initialization costs, and I/O throughput. Our results reveal the land (LND) component’s efficient scaling, demonstrating near-ideal weak scaling and strong-scaling parallel efficiencies reaching up to 87% at 50,400 cores. We confirm portability and reproducibility through bitwise-equivalent outputs across different machines using identical inputs over supported machines. Notably, at extreme scales, we identify I/O as a critical bottleneck and that leads to effective solution with the SCORPIO/ADIOS stack. Collectively, these findings validate the deployment of km-ELM at a continental scale with high parallel efficiency and provide essential guidance on configuration, decomposition, and I/O settings for optimized kilometer-scale land simulations in E3SM. This work emphasizes the innovative design and practical solutions that enhance the operational capabilities of km-ELM, focusing on software performance and scalability while leaving detailed scientific evaluations of simulated land processes for future investigations.

E3SM land model (ELM), km-ELM, scalability, perfor↗

Large‐Area, High‐Numerical‐Aperture, Freeform Metasurfaces

Abstract Nanophotonic devices are optical platforms capable of unprecedented wavefront control. To push the limits of experimental device performance, scalable design methodologies that combine the simplicity and fabricability of conventional design paradigms with the extended capabilities of freeform optimization are required. A novel gradient‐based design framework for large‐area freeform metasurfaces is introduced in which nonlocal interactions between simply shaped nanostructures, placed on an irregular lattice, are tailored to produce high‐order hybridized modes that support customizable large‐angle scattering profiles. Utilizing this approach, multifunctional super‐dispersive metalenses are designed and experimentally demonstrated. The approach to high‐numerical‐aperture radial metalenses capable of diffraction limited focusing and the generation of donut‐shaped point spread functions is also extended. It is anticipated that these concepts will have utility in super‐resolution microscopy, particle trapping, additive manufacturing, and metrology applications that require ultra‐high numerical apertures.

Optics↗

Eco–Friendly Solvent Engineered CsPbI 2.77 Br 0.23 Ink for Large–Area and Scalable High Performance Perovskite Solar Cells

The performance of large-area perovskite solar cells (PSCs) has been assessed for typical compositions, such as methylammonium lead iodide (MAPbI 3 ), using a blade coater, slot-die coater, solution shearing, ink-jet printing, and thermal evaporation. However, the fabrication of large-area all-inorganic perovskite films is not well developed. This study develops, for the first time, an eco-friendly solvent engineered all-inorganic perovskite ink of dimethyl sulfoxide (DMSO) as a main solvent with the addition of acetonitrile (ACN), 2-methoxyethanol (2-ME), or a mixture of ACN and 2-ME to fabricate large-area CsPbI 2.77 Br 0.23 films with slot-die coater at low temperatures (40–50 °C). The perovskite phase, morphology, defect density, and optoelectrical properties of prepared with different solvent ratios are thoroughly examined and they are correlated with their respective colloidal size distribution and solar cell performance. Here, the optimized slot-die-coated CsPbI 2.77 Br 0.23 perovskite film, which is prepared from the eco-friendly binary solvents dimethyl sulfoxide:acetonitrile (0.8:0.2 v/v), demonstrates an impressive power conversion efficiency (PCE) of 19.05%. Moreover, the device maintains ≈91% of its original PCE after 1 month at 20% relative humidity in the dark. It is believed that this study will accelerate the reliable manufacturing of perovskite devices.

binary and ternary solvent↗

Light, High Performance and Scalable Coal-Derived Composites for Construction: Precast and Cast-in-Place Applications

The overall objective of this project was to produce a coal-based construction material that has up to ~95 weight percent (wt. %) coal with physical, chemical, and thermal properties exceeding those of ordinary Portland cement (OPC)-based construction materials. Additionally, the project aimed to minimize external binders by implementing novel mixing techniques, while exceeding the performance/cost ratio of OPC. Finally, the project was to demonstrate production of precast products via the design and fabrication of products via a bench scale process. Consistent with some of these objectives, the project successfully fabricated samples of coal-based composite materials with >80 wt% coal with physical, chemical and thermal properties on par with cement-based concrete. Select samples demonstrated compressive strengths with >7,000 psi and flexural strength of >420 psi. The composite materials minimized external binders and also demonstrated durability, as evidenced by resistance to acidic and basic solutions. Finally, larger slab and beam type samples were produced using a process developed by the Recipient, although, the process was not semicontinuous in nature. Taken together, the results of this project suggest that domestic coal has potential to serve as a replacement for cementitious materials utilized in incumbent construction technologies, which could significantly reduce the energy and emissions of the construction industry

01 COAL, LIGNITE, AND PEAT↗

TChem-atm (v2.0.0): scalable performance-portable multiphase atmospheric chemistry

We present TChem-atm, a performance-portable approach that enables efficient simulation of chemically detailed and multiphase atmospheric chemistry on modern heterogeneous computing architectures. Unlike previous efforts that rely on architecture-specific code or focus exclusively on gas-phase chemistry, TChem-atm supports fully coupled gas–aerosol systems with execution across CPUs, NVIDIA GPUs, and AMD GPUs through the Kokkos programming model. It integrates the flexible multiphase capabilities of the Community Atmospheric Model Chemistry Package (CAMP) with the high-performance kinetic routines of TChem, and includes automatic Jacobian construction with support for a range of stiff ODE solvers. In a proof-of-concept integration with the particle-resolved model PartMC, TChem-atm reproduces the existing PartMC–CAMP implementation within solver tolerances and delivers substantial GPU speedups, especially for large particle populations. Performance benchmarks reveal substantial speedups on GPU platforms, particularly for large particle populations, with consistent results across hardware backends. TChem-atm enables performance-portable execution across CPUs and GPUs, though optimal efficiency may require modest architecture-specific tuning (e.g., team and vector sizes), with up to a twofold improvement on the NVIDIA H100. It directly supports sectional and particle-resolved host models, while modal aerosol schemes require minor adaptation to provide particle-scale quantities such as representative diameters. By enabling chemically detailed, multiphase simulations with performance portability and host-model flexibility, TChem-atm facilitates the incorporation of advanced chemistry into atmospheric models.

Díaz-Ibarra, Oscar Homero [Sandia National Laborat↗

Scaling Ensembles of Data-Intensive Quantum Chemical Calculations for Millions of Molecules

Deep learning models are efficient computational tools that can accelerate the inverse design of molecules with desired functional properties by generating predictions at a fraction of the time required by traditional quantum chemical approaches. To ensure that a model maintains accuracy and transferability across broad regions of the chemical space explored during the inverse design, it must be trained on massively large volumes of simulation data. This requires running large-scale ensemble quantum chemical calculations on high-performance computing (HPC) systems for data collection. However, the efficient execution of such large ensemble calculations and the management of large volumes of output data require tools that can judiciously utilize computational resources and manage metadata overhead on the file system. Therefore, we present a high-performance, scalable, ensemble management framework for performing data-intensive quantum chemical electronic structure calculations for organic molecules. This framework provides abstractions to plug different ab initio, first principles, and first principles-based semi-empirical methods and executes them efficiently at large scale on HPC systems. It dynamically distributes tasks to resources and uses tiered storage for managing large collections of files. We employed this framework to process over ten million organic molecules and generate open-source datasets that provide UV-vis absorption spectra by running time-dependent density-functional tight-binding calculations. It is the largest database containing molecular optical spectra that were simulated with quantum chemical methods in a consistent manner.

Mehta, Kshitij↗

Fast HARDI Uncertainty Quantification and Visualization with Spherical Sampling

In this paper, we study uncertainty quantification and visualization of orientation distribution functions (ODF), which corresponds to the diffusion profile of high angular resolution diffusion imaging (HARDI) data. The shape inclusion probability (SIP) function is the state‐of‐the‐art method for capturing the uncertainty of ODF ensembles. The current method of computing the SIP function with a volumetric basis exhibits high computational and memory costs, which can be a bottleneck to integrating uncertainty into HARDI visualization techniques and tools. We propose a novel spherical sampling framework for faster computation of the SIP function with lower memory usage and increased accuracy. In particular, we propose direct extraction of SIP isosurfaces, which represent confidence intervals indicating spatial uncertainty of HARDI glyphs, by performing spherical sampling of ODFs. Our spherical sampling approach requires much less sampling than the state‐of‐the‐art volume sampling method, thus providing significantly enhanced performance, scalability, and the ability to perform implicit ray tracing. Our experiments demonstrate that the SIP isosurfaces extracted with our spherical sampling approach can achieve up to 8164× speedup, 37282× memory reduction, and 50.2% less SIP isosurface error compared to the classical volume sampling approach. We demonstrate the efficacy of our methods through experiments on synthetic and human‐brain HARDI datasets.

97 MATHEMATICS AND COMPUTING↗

Scalable membrane-less microbial electrolysis cell with multiple compact electrode assemblies for high performance hydrogen production

Bioelectrochemical hydrogen production via microbial electrolysis cells (MECs) is a promising method for sustainable energy production and decarbonization of energy systems. However, the application of MECs is limited by the electrochemical performance, scalability, and the cost associated with expensive materials. Here, in this study, a scalable MEC (500 mL) with novel compact electrode assemblies and high electrode surface area to volume ratio (160 m 2 /m 3 ) was designed and constructed. The use of membranes, precious metal catalyst, and current collectors with high costs was avoided. A high current density at the steady state of 49.5 ± 5.3 A/m 2 was achieved using acetate as the substrate with phosphate buffer under the applied voltage of 1.01 V. The corresponding volumetric current density was 3948 ± 422 A/m 3 . The compact electrode assembly design limited methane production rate to 3.9 ± 0.2 L/L/D, while achieving a hydrogen production rate of 33.7 ± 1.7 L/L/D. With the suppression of microbial hydrogen consumption, the hydrogen production rate was 39.8 ± 1.9 L/L/D, higher by almost one order of magnitude than those of MECs with scaling up attempts. The compact electrode configuration reduced internal resistance to 88.5 ± 4.4 Ω cm 2 . The energy efficiency based on input electricity was 146 ± 7 % to 189 ± 9 % within the applied voltage range of 0.71 to 1.05 V. The results in this study demonstrated successful scaling up of high performance small MECs and offered a new possible approach of scaling up MECs by stacking high-performance subunits, with no trade-offs on electrochemical performance.

08 HYDROGEN↗

Scalable augmented enumeration and metadata operations for large filesystems

Systems, apparatus, and methods are disclosed for performing scalable operations in a file system. Metadata entries in a namespace or directory tree are sharded across multiple file metadata servers. An augmented enumeration operation, such as listing a directory, is parallelized across the multiple file metadata servers, transparently to clients. Exemplary augmentation features can include filtering and sorting. Augmentation features can be executed concurrently with enumeration, prior to enumeration, after enumeration, or as a combination of these, and can utilize pre-built index structures or holding structures for intermediate results. Augmented enumeration operations can also include no-output operations such as changing file attributes or deleting a file, and cumulative operations such as counting total disk space usage. The parallelization is compatible with tree-level parallelization and storage-level parallelization. Disclosed technologies can be applied to other fields requiring scalable enumeration, such as database and network applications.

Grider, Gary A.↗

Scalable and Regenerable Fibrous Amine-functionalized Matrix (FAM) sorbent for Efficient Enrichment of Critical Minerals from Coal Wastewaters

The poster presents the latest progress on utilizing a commercially scalable flat sheet sorbent for the effective enrichment of critical minerals from coal wastewater. It highlights the performance, scalability, and potential for industrial applications, addressing key challenges in critical recovery from complex wastewater streams.

critical metals↗

Non-dimensional performance and safety parameters for heat pipes

The use of heat pipes in safety-critical systems such as nuclear microreactors dictates the development of generalized, practical, scalable performance and safety parameters. Traditional dimensional metrics, while informative, lack the universality required for comparative analysis across varying designs and operating regimes. Here, this work introduces a comprehensive set of non-dimensional parameters to characterize heat pipe performance and safety, including capillary performance, effective thermal conductivity, response time, exergetic efficiency, allowable temperature gradients, allowable rate of temperature change, priming coefficients, and factor of safety. A reference heat pipe design representative of microreactor applications was analyzed via the developed parameters using both traditional analytical models and Sockeye simulations under transient and steady-state conditions. Sodium, potassium, and water were evaluated as working fluids to demonstrate the applicability of the framework across a broad temperature range. The proposed non-dimensional parameters effectively captured key thermal-hydraulic behaviors and safety concerns, as was demonstrated via Sockeye simulations. This framework supports the development of design optimization strategies, operational protocols, and safety assurance practices for advanced reactor systems and other high-reliability applications.

42 - ENGINEERING↗

Nek5000/RS performance on advanced GPU architectures

The authors explore performance scalability of the open-source thermal-fluids code, NekRS, on the U.S. Department of Energy's leadership computers, Crusher, Frontier, Summit, Perlmutter, and Polaris. Particular attention is given to analyzing performance and time-to-solution at the strong-scale limit for a target efficiency of 80%, which is typical for production runs on the DOE's high-performance computing systems. Several examples of anomalous behavior are also discussed and analyzed.

97 MATHEMATICS AND COMPUTING↗

hPIC2: A hardware-accelerated, hybrid particle-in-cell code for dynamic plasma-material interactions

The exascale era of high performance computing promises to bring the field of computational plasma physics ever closer to the goal of accurate multiscale modeling. Such computers will rely on hardware acceleration to offload work to dedicated components, notably general-purpose graphics processing units (GPUs). However, devices from different manufacturers require software to be written with different parallel programming models, greatly increasing the code maintenance burden of applications designed to perform on more than one such device. hPIC2 is a hybrid plasma simulation code developed with the Kokkos performance portability framework to target the architectures that will drive exascale computing for the foreseeable future. As a hybrid simulation code, hPIC2 investigates the simultaneous use of various plasma models on the same domain, at the same time. hPIC2 also optionally couples to RustBCA, which accurately models ion-material interactions using the binary collision approximation (BCA) method. In conclusion, hPIC2 therefore achieves scalable performance on a variety of computing architectures when simulating complex and diverse plasmas, particularly near plasma-material interfaces.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗