Search NASA⌕ Search

SEARCH · Search NASA

Results for “Portability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

HPDR: High-Performance Portable Scientific Data Reduction Framework

The rapid growth in scientific data generation is outpacing advancements in computing systems necessary for efficient storage, transfer, and analysis, particularly in the context of exascale computing. With the deployment of first-generation exascale computing systems and next-generation experimental facilities, this gap is widening and necessitates effective data reduction techniques to manage enormous data volumes. Over the past decade, various data reduction methods, including lossless compression, error-controlled lossy compression, and data refactoring, have been developed to accelerate I/O in scientific workflows. Despite significant reductions in data volume, these methods introduce considerable computational overhead, which can become the new bottleneck in data processing. To mitigate this, GPU-accelerated data reduction algorithms have been introduced. However, challenges remain in their integration into exascale workflows, including limited portability across different GPU architectures, substantial memory transfer overhead, and reduced scalability on dense multi-GPU systems. To address these challenges, we propose HPDR, a high-performance and portable data reduction framework. HPDR is designed to enable the execution of state-of-the-art reduction algorithms across diverse processor architectures while reducing memory transfer overhead to 2.3 % of the original, resulting in up to 3.5× faster throughput compared to existing solutions. It also achieves up to 96% of the theoretical speedup in multi-GPU settings. In addition, evaluations on accelerating I/O operations at scale up to 1,024 nodes of the Frontier supercomputer demonstrate that HPDR can achieve up to 103 TB/s reduction throughput, providing up to 4× acceleration in parallel I/O performance compared to existing data reduction routines. This work highlights the potential of HPDR to significantly enhance data reduction efficiency in exascale computing environments.

Chen, Jieyang [University of Oregon]↗

ChatBLAS: The First AI-Generated and Portable BLAS Library

We present ChatBLAS, the first AI-generated and portable Basic Linear Algebra Subprograms (BLAS) library on different CPU/GPU configurations. The purpose of this study is (i) to evaluate the capabilities of current large language models (LLMs) to generate a portable and HPC library for BLAS operations and (ii) to define the fundamental practices and criteria to interact with LLMs for HPC targets to elevate the trustworthiness and performance levels of the AI-generated HPC codes. The generated C/C++ codes must be highly optimized using device-specific solutions to reach high levels of performance. Additionally, these codes are very algorithm-dependent, thereby adding an extra dimension of complexity to this study. We used OpenAI’s LLM ChatGPT and focused on vector-vector BLAS level-1 operations. ChatBLAS can generate functional and correct codes, achieving high-trustworthiness levels, and can compete or even provide better performance against vendor libraries.

Valero Lara, Pedro↗

Mojo: MLIR-based Performance-Portable HPC Science Kernels on GPUs for the Python Ecosystem

We explore the performance and portability of the novel Mojo language for scientific computing workloads on GPUs. As the first language based on the LLVM’s Multi-Level Intermediate Representation (MLIR) compiler infrastructure, Mojo aims to close performance and productivity gaps by combining Python’s interoperability and CUDA-like syntax for compile-time portable GPU programming. We target four scientific workloads: a seven-point stencil (memory-bound), BabelStream (memory-bound), miniBUDE (compute-bound), and Hartree–Fock (compute-bound with atomic operations); and compare their performance against vendor baselines on NVIDIA H100 and AMD MI300A GPUs. We show that Mojo’s performance is competitive with CUDA and HIP for memory-bound kernels, whereas gaps exist on AMD GPUs for atomic operations and for fast-math compute-bound kernels on both AMD and NVIDIA GPUs. Although the learning curve and programming requirements are still fairly low-level, Mojo can close significant gaps in the fragmented Python ecosystem in the convergence of scientific computing and AI.

Godoy, William [ORNL] (ORCID:0000000225905178)↗

Portable HCAL reconstruction in the CMS detector using the Alpaka library

CMS has deployed a number of different GPU algorithms at the High-Level Trigger (HLT) in Run 3. As the code base for GPU algorithms continues to grow, the burden for developing and maintaining separate implementations for GPU and CPU becomes increasingly challenging. To mitigate this, CMS has adopted the Alpaka (Abstraction Library for Parallel Kernel Acceleration) library as the performance portability solution to provide a single-code base for parallel execution on both GPUs and CPUs in CMS software (CMSSW). A direct CUDA version of HCAL energy reconstruction, called Minimization At Hcal, Iteratively (MAHI), has been deployed at the HLT in the 2022-2023 data taking period. This contribution will describe how the CUDA version is converted into a portable implementation using the Alpaka library. We will discuss the porting experience from CUDA to Alpaka, the validation process and the performance of the Alpaka version in CPU and GPU.

Kwok, Martin↗

Image Reconstruction from Sparse-view Data Acquired with Portable X-ray Devices

• Portable X-ray systems enable on-site 3D imaging for non-invasive inspection of suspicious packages and explosives. • Existing reconstruction algorithms (e.g., FDK or Feldkamp, Davis and Kress) require hundreds of projections over 360 degrees. • Sparse-view scan reduces scanning time and setup effort, making it ideal for field use in timecritical scenarios. • Existing reconstruction algorithms introduce severe artifacts when applied to sparse-view data. • We developed a total variation (TV)-based optimization algorithm for yielding 3D images from sparse-view data collected with our portable X-ray imaging system.

Xia, Dan [University of Chicago, Chicago, IL]↗

A Novel, Low-Cost, Portable Device for Counterfeit and Noncompliant Refrigerant Detection

The increasing prevalence of counterfeit refrigerants presents significant risks to Heating, Ventilation, Air Conditioning, and Refrigeration (HVACR) systems, including compromised equipment performance, safety hazards, and environmental non-compliance. This report details the development of a novel, cost-effective, and portable detection device designed to accurately identify counterfeit refrigerants. The device utilizes a controlled gas sampling and analysis system within a sealed chamber, ensuring precise measurements while maintaining safety through a purging mechanism. The system features a high-sensitivity sensor integrated with an onboard control module that analyzes gas composition in real-time, providing users with clear visual indicators for refrigerant authenticity. Laboratory validation demonstrated the device’s high accuracy (exceeding 95%) in detecting unauthorized refrigerant blends. Key advantages include affordability, ease of use, rapid response time, and compatibility with a wide range of refrigerants. This solution supports compliance with regulatory frameworks such as the AIM Act, enhances safety in HVACR operations, and mitigates the risks associated with counterfeit refrigerants. Future developments will focus on expanding refrigerant detection capabilities, integrating machine learning for enhanced accuracy, implementing cost-reduction strategies to improve accessibility and market adoption, and optimizing system packaging for enhanced field portability.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Rescue of Unique Liquid IP-2 Portable Tanks at Savannah River Site – 25078

Westinghouse Savannah River Company, the SRS Management and Operations Contractor at the time, procured six large IP-2 portable tanks capable of transporting liquids in 2003 to support the layup of the F Canyon at SRS. F Canyon was one of two full-scale hardened radiochemical separations facilities built at SRS in the 1950s as part of the initial construction of the Site. F Canyon and H Canyon provided actinide separations capabilities for SRS through the end of the Cold War. Following the end of the Cold War era, the decision was eventually made that only one of these radiochemical separations facilities at SRS would still be needed moving forward. As H Canyon was considered better suited for ongoing SRS missions, F Canyon was placed into a layup state in the first decade of the 21st Century, with F Area deactivation completing in 2006. This process resulted in the creation of excess, radioactively contaminated solutions, which required disposition. These solutions included contaminated organic solutions as well as DU-bearing aqueous solutions. The six portable tanks were procured to support the transportation of these solutions.

Fee, Nathaniel C. [Savannah River Nuclear Solution↗

Performance Portability Evaluation of Fluid-Structure Interaction Simulations on Heterogeneous Platforms

The rapid proliferation of heterogeneous programming languages and multi-vendor hardware has underscored the critical need to evaluate the performance portability of scientific applications. In this work, we present the systematic porting and optimization of a massively parallel fluid-structure interaction code across multiple heterogeneous programming frameworks for deployment on leadership-class supercomputers from major vendors. Our analysis focuses on at-scale performance for simulations involving hundreds of millions of deformable cells, executed on a combination of CPUs and GPUs spanning thousands of nodes on exascale machines. We benchmark the performance of each implementation, highlighting the trade-offs inherent in adopting diverse programming models. Key insights regarding the portability of CUDA on multi-vendor platforms, the superior multi-core CPU performance from SYCL, and architectural considerations on performance optimization are distilled from our experience, offering guidance to other users of high performance computing based on our findings.

Martin, Aristotle [Duke University]↗

Portable and Cost-Effective Device for Reliable Detection of Counterfeit and Non-compliant Refrigerants in Diverse Applications

Counterfeit refrigerants pose significant challenges to safety, system reliability, and operational effectiveness due to their harmful contaminants or incompatible chemical compositions. Utilizing these noncompliant products can lead to reduced efficiency, equipment failures, and expensive repairs. Additionally, heightened demand for alternative refrigerants during the industry's transition has created supply gaps, enabling counterfeit products to proliferate. Accurate detection and analysis tools are therefore essential to verify refrigerant authenticity and ensure system integrity in diverse applications. This paper presents the development of a portable device designed for reliable identification and detailed analysis of refrigerant composition. By integrating precision gas sampling, controlled pressure regulation, and automated sensor technology, the device not only detects deviations from standard refrigerant properties but also provides a comprehensive composition breakdown. Pre-calibrated sensors measure the refrigerant gas to identify specific concentrations and contaminants, with an intuitive LED-based indicator system ensuring quick interpretation of results. The user-friendly interface enables operators to select refrigerant types for targeted testing, further enhancing accuracy and usability for field technicians. Comprehensive testing was conducted on mildly flammable A2L refrigerants, showcasing the device’s robustness and adaptability in analyzing composition and detecting discrepancies. The device demonstrated consistent accuracy across a range of refrigerant samples, affirming its reliability in diverse operational environments. Its design minimizes contamination risks during sampling and provides detailed composition results within 90 seconds, ensuring efficient and precise analysis. With a projected price point under $150, the proposed solution delivers affordability alongside its lightweight portability and straightforward operation. Unlike complex and costly alternatives, such as gas chromatography systems, this device provides an accessible option for technicians, customs personnel, and industry operators in need of quick and effective refrigerant verification. Compatible with both current formulations and emerging refrigerant technologies, the device addresses critical counterfeit detection needs across a range of applications. By delivering accurate composition analysis and counterfeit identification, this innovation enhances system performance, safety, and operational reliability in crucial industries.

Cheekatamarla, Praveen [ORNL] (ORCID:0000000248827↗

AMR-Wind: A Performance-Portable, High-Fidelity Flow Solver for Wind Farm Simulations

We present AMR-Wind, a verified and validated high-fidelity computational-fluid-dynamics code for wind farm flows. AMR-Wind is a block-structured, adaptive-mesh, incompressible-flow solver that enables predictive simulations of the atmospheric boundary layer and wind plants. It is a highly scalable code designed for parallel high-performance computing with a specific focus on performance portability for current and future computing architectures, including graphical processing units (GPUs). In this paper, we detail the governing equations, the numerical methods, and the turbine models. Establishing a foundation for the correctness of the code, we present the results of formal verification and validation. The verification studies, which include a novel actuator line test case, indicate that AMR-Wind is spatially and temporally second-order accurate. The validation studies demonstrate that the key physics capabilities implemented in the code, including actuator disk models, actuator line models, turbulence models, and large eddy simulation (LES) models for atmospheric boundary layers, perform well in comparison to reference data from established computational tools and theory. We conclude with a demonstration simulation of a 12-turbine wind farm operating in a turbulent atmospheric boundary layer, detailing computational performance and realistic wake interactions.

17 WIND ENERGY↗

A portable application framework for energy management and information systems (EMIS) solutions using Brick semantic schema

This paper introduces a portable framework for developing, scaling and maintaining energy management and information systems (EMIS) applications using an ontology-based approach. Key contributions include an interoperable layer based on Brick schema, the formalization of application constraints pertaining metadata and data requirements, and a field demonstration. The framework allows for querying metadata models, fetching data, preprocessing, and analyzing data, thereby offering a modular and flexible workflow for application development. Its effectiveness is demonstrated through a case study involving the development and implementation of a data-driven anomaly detection tool for the photovoltaic systems installed at the Politecnico di Torino, Italy. During eight months of testing, the framework was used to tackle practical challenges including: (i) developing a machine learning-based anomaly detection pipeline, (ii) replacing data-driven models during operation, (iii) optimizing model deployment and retraining, (iv) handling critical changes in variable naming conventions and sensor availability (v) extending the pipeline from one system to additional ones.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Innovative approach to counterfeit and noncompliant refrigerant detection: A cost-effective, portable solution

The increasing prevalence of counterfeit and incompatible refrigerants presents significant risks to Heating, Ventilation, Air Conditioning, and Refrigeration (HVAC&R) systems, including compromised equipment performance, safety hazards, and non-compliance. This article details the development of a novel, cost-effective, and portable detection device designed to accurately verify refrigerants. The device utilizes a controlled gas sampling and analysis system within a sealed chamber, ensuring precise measurements while maintaining safety through a purging mechanism. The system features a high-sensitivity sensor integrated with an onboard control module that analyzes gas composition in real-time, providing feedback within a 2-minute duration. Laboratory validation demonstrated the device’s high accuracy (>95 % based on correct identification of compliant vs. non-compliant blends) in detecting unauthorized refrigerant blends. The projected cost of the product stands at ∼ $150, based on the retail pricing of individual components. Laboratory validation demonstrated the device’s high accuracy (>95 % for composition identification, 100 % rejection of tested counterfeit/incorrect blends) in detecting unauthorized refrigerant blends with a response time <2 min. The device correctly identified authentic R-454A/B/C blends and reliably rejected R-407F and closely related counterfeit mixtures. Key advantages include affordability, ease of use, rapid response time, and compatibility with a wide range of refrigerants. This solution supports compliance with regulatory frameworks, enhances safety in HVAC&R operations, and mitigates the risks associated with counterfeit refrigerants.

Counterfeit refrigerants↗

Portable Software Environment for Ultrahigh-Resolution ELM Development on GPUs

This paper presents our endeavors in developing the large-scale, ultra-high-resolution E3SM Land Model (uELM), specifically designed for exascale computers furnished with accelerators such as Nvidia GPUs. The uELM is a sophisticated code that substantially relies on High-Performance Computing (HPC) environments, necessitating particular machine and software configurations. To facilitate community-based uELM developments employing GPUs, we have created a portable, standalone software environment preconfigured with uELM input datasets, simulation cases, and source code. This environment, utilizing Docker, encompasses all essential code, libraries, and system software for uELM development on GPUs. It also features a functional unit test framework and an offline model testbed for comprehensive numerical experiments. From a technical perspective, the paper discusses GPU-ready container generations, uELM code management, and input data distribution across computational platforms. Lastly, the paper demonstrates the use of environment for functional unit testing, end-to-end simulation on CPUs and GPUs, and collaborative code development.

E3SM Land Model↗

Sample screening of uranium ore concentrates using portable spectrophotometers: investigating the correlation between visible colors and chemical signatures

A research collaboration between the Japan Atomic Energy Agency and the Department of Energy’s National Nuclear Security Administration examined nuclear forensic signatures and analytical methods for tracing the origins of uranium ore concentrates (UOCs). Here, this study focuses on utilizing portable spectrophotometers capable of reflectance measurements in the visible light spectrum as a potential rapid screening tool for nuclear forensics analysis. Unlike laboratory-based near-infrared spectroscopy or digital image analysis, this research investigated the potential to correlate visible color measurements with key nuclear forensics signatures using seven types of UOC samples with known origins and three UOC certified reference materials. Results demonstrated that distinct color groups, quantified using CIELAB values, correlated with major uranium compounds. Furthermore, the findings indicated that trace elements can influence the UOC colors, providing additional insights into material characteristics. Although this approach requires further validation across a broader range of UOC species, this study demonstrated that simple colorimetric analysis using visible spectrophotometry, which does not require complex sample preparation or data processing, can serve as a practical and novel rapid tool for preliminary screening and attribution in nuclear forensics investigations.

and nuclear chemistry↗

Portable Acceleration of CMS Computing Workflows with Coprocessors as a Service

Computing demands for large scientific experiments, such as the CMS experiment at the CERN LHC, will increase dramatically in the next decades. To complement the future performance increases of software running on central processing units (CPUs), explorations of coprocessor usage in data processing hold great potential and interest. Coprocessors are a class of computer processors that supplement CPUs, often improving the execution of certain functions due to architectural design choices. We explore the approach of Services for Optimized Network Inference on Coprocessors (SONIC) and study the deployment of this as-a-service approach in large-scale data processing. In the studies, we take a data processing workflow of the CMS experiment and run the main workflow on CPUs, while offloading several machine learning (ML) inference tasks onto either remote or local coprocessors, specifically graphics processing units (GPUs). With experiments performed at Google Cloud, the Purdue Tier-2 computing center, and combinations of the two, we demonstrate the acceleration of these ML algorithms individually on coprocessors and the corresponding throughput improvement for the entire workflow. This approach can be easily generalized to different types of coprocessors and deployed on local CPUs without decreasing the throughput performance. We emphasize that the SONIC approach enables high coprocessor usage and enables the portability to run workflows on different types of coprocessors.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Toucan: A performance portable, scalable implementation of the DECA algorithm

In the field of additive manufacturing (AM), cellular automata (CA) is extensively used to simulate microstructural evolution during solidification. However, while traditional CA approaches are relatively fast, they still require a substantial number of time steps, are limited to moderate volumes, and are relatively difficult to improve through parallelism due to the highly localized nature of the solidification front. Here, to address these issues of time to solution and load balancing, we introduce Toucan, a parallel, performance-portable, and scalable code written in C++ with the Kokkos library that leverages the discrete event inspired cellular automata (DECA) algorithm to perform parallel-in-time (PinT) grain growth simulations. Toucan effectively mitigates load balancing issues by distributing the computational workload more evenly across processors, enhancing scalability and efficiency. We conduct both strong and weak scaling studies on up to 64 GPUs on the Frontier supercomputer, demonstrating that Toucan significantly outperforms the current state-of-the-art, time-stepped CA code, ExaCA, on both single and multi-GPU simulations. Even in AM-specific weak scaling scenarios, Toucan maintains near-ideal scaling, in contrast to the linear increase observed with ExaCA due to the moving laser raster pattern. This study highlights Toucan’s potential to transform microstructural simulations in AM by radically improving both efficiency and scalability over existing methods.

36 MATERIALS SCIENCE↗