Search NASA⌕ Search

SEARCH · Search NASA

Results for “Portable application”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

PETSc/TAO Users Manual Revision 3.23

This manual describes the use of the Portable, Extensible Toolkit for Scientific Computation (PETSc) and the Toolkit for Advanced Optimization (TAO) for the numerical solution of partial differential equations (PDEs) and related problems on high-performance computers. PETSc/TAO is a suite of data structures and routines that provide the building blocks for implementing large-scale application codes on parallel (and serial) computers. PETSc uses the MPI standard for all distributed memory communication.

97 MATHEMATICS AND COMPUTING↗

PETSc/TAO Users Manual Revision 3.24

This manual describes the use of the Portable, Extensible Toolkit for Scientific Computation (PETSc) and the Toolkit for Advanced Optimization (TAO) for the numerical solution of partial differential equations (PDEs) and related problems on high-performance computers. PETSc/TAO is a suite of data structures and routines that provide the building blocks for implementing large-scale application codes on parallel (and serial) computers. PETSc uses the MPI standard for all distributed memory communication.

97 MATHEMATICS AND COMPUTING↗

PETSc/TAO Users Manual Revision 3.25

This manual describes the use of the Portable, Extensible Toolkit for Scientific Computation (PETSc) and the Toolkit for Advanced Optimization (TAO) for the numerical solution of partial differential equations (PDEs) and related problems on high-performance computers. PETSc/TAO is a suite of data structures and routines that provide the building blocks for implementing large-scale application codes on parallel (and serial) computers. PETSc uses the MPI standard for all distributed memory communication.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

A Synthetic Pathway for Producing Carbon Dots for Detecting Iron Ions Using a Fiber Optic Spectrometer

Iron detection is of growing importance in the critical minerals sector, where unwanted iron ions are typically removed during the processing of target critical metals. The ideal sensor should utilize inexpensive, scalable materials along with a low-cost, robust, and easy-to-use analysis platform. Here, we demonstrate a simple acid–base synthesis of luminescent iron-responsive carbon dots by reacting ethanolamine, phosphoric acid, and m-phenylenediamine. The carbon dots exhibit selective, iron-specific emission quenching, with the ability to detect part-per-billion levels of iron ions even in 0.1 M HCl. After benchmarking the purified materials using a commercial spectrometer, a “low-cost” process is demonstrated in which carbon dots with minimal purification are coupled with a portable fiber-optic spectrometer for analyzing iron content. Carbon dot-coated paper strips are also evaluated as another convenient platform for iron analysis. Taken together, the sensing material and platforms demonstrated here are well-suited for detecting trace quantities of iron in environmentally relevant conditions, with potential applications in tracking iron removal processes during critical mineral production as one exciting area of interest.

acid mine draining↗

Fabrication of metal-organic framework thin films for luminescent sensing applications

Metal−organic frameworks (MOFs) have emerged as an exciting material class due to their nearly infinite design space: an inexhaustible combination of metal centers, organic linkers, and reaction conditions may be leveraged to obtained desired optical, physical, and chemical properties. This rich synthetic diversity enables properties such as pore dimensions and functionality to be rationally tailored for highly selective and sensitive detection of target analytes. Luminescent MOF-based sensors in particular offer advantages including low cost, portability, and ease-of-use. Often, performance is maximized by utilizing the luminescent MOF in thin film form, for example by integrating the film onto an optical fiber, yet the synthesis of high-quality MOF thin films often requires tedious, expensive, and/or slow approaches. Here, we demonstrate a rapid, simple strategy for growing copper MOFs using metal oxide templates; the concept is first demonstrated with a well-studied system, copper-1,3,5-benznetricarboxylate (Cu-BTC). Variables such as the choice of solvent, pH, and the choice of the copper salt anion all govern thin film quality. We then extend this method to synthesize a solvent and metal-responsive copper MOF that is particularly effective at detecting aluminum, an economically critical metal, at trace concentrations using a fully portable luminescence spectrometer. Taken together, this work demonstrates a convenient, versatile, and sustainable method for developing high quality MOF thin film-based optical sensors.

Crawford, Scott↗

NOVEL METHODS FOR IN SITU HIGH DENSITY SURFACE CLEANING SCRUBBING OF ULTRAHIGH VAC LONG NARROW TUBES TO REDUCE SECONDARY ELECTRON YIELD AND OUT GASSING

This project developed new methods for cleaning the inside surfaces of very long, narrow vacuum tubes used in particle accelerators. Traditional cleaning approaches are expensive, slow, or difficult to implement in accelerator tunnels. We designed and tested a portable plasma discharge cleaning system that uses lower-cost microwave and magnetron technologies to reduce outgassing and secondary electron emission from stainless steel and copper surfaces. The system, called the Plasma Discharge Test System (PDTS), allows accelerator components to be scrubbed more efficiently, which can improve performance and reduce maintenance costs for research and industrial applications.

POOLE, JOE HENRY [PRESIDENT]↗

Experimental Report: Multi-Instrument Comparison of AAF Size Distribution Instruments

Aerosols are particles suspended in the atmosphere, ranging in size from nanometers to micrometers. Their size distribution affects key atmospheric processes, including nucleation, coagulation, scavenging, activation, and radiative properties (Seinfeld and Pandis 2016). Aerosol size distribution is a critical parameter in atmospheric science, influencing processes such as cloud formation and radiative forcing. Accurate representation of aerosol size distributions is essential for understanding their impact on climate, air quality, and human health. However, aerosol size and composition vary significantly across time and space due to meteorological conditions and natural or anthropogenic sources. (Wu and Boor 2021). Various instruments are used to measure aerosol size distributions, each with distinct principles, advantages, and limitations. This report begins with an in-depth overview of aerosol size-distribution comparison studies, focusing on the passive cavity aerosol spectrometer probe (PCASP), portable optical particle spectrometer (POPS), ultra-high-sensitivity aerosol spectrometer (UHSAS), aerodynamic particle sizer (APS), and scanning mobility particle sizer (SMPS). All of these instruments are used by the U.S. Department of Energy Atmospheric Radiation Measurement (ARM) Aerial Facility (AAF), which commissioned this comparison and report. The report evaluates the strengths and weaknesses of these instruments, highlights their applications, and discusses efforts to merge data from multiple instruments for comprehensive analysis.

54 ENVIRONMENTAL SCIENCES↗

An Integrated Framework for Memory-Centric Analysis: From Trace Collection to Co-Design

The memory wall phenomenon—where advances in processor performance significantly outpace those in memory subsystems—poses a fundamental challenge for contemporary computing systems. In memory-bound applications, memory subsystem behavior dominates performance, yet existing analysis approaches present significant limitations: detailed microarchitectural simulators require days to weeks to simulate modest workloads; hardware performance counters provide only aggregate statistics that obscure temporal and spatial access patterns; and scaled simulation approaches face challenges in capturing certain behaviors that emerge at larger scales. These limitations reflect a processor-centric design philosophy increasingly misaligned with memory-bound workloads where detailed understanding of memory access patterns, cache hierarchy interactions, and contention is critical for effective optimization. This paper presents an integrated framework for memory-centric analysis that enables effective hardware-software co-design. We describe practical trace collection techniques, including hardware-assisted processor tracing with minimal overhead and portable software-based instrumentation with statistical sampling. We present multi-perspective analysis methods that examine memory behavior from temporal, sequential, spatial, and relational viewpoints, revealing distinct optimization opportunities invisible in aggregate metrics. We detail an architectural modeling framework that uses sampled traces with temporal interpolation and confidence-based filtering to evaluate cache and memory configurations. Evaluation on representative benchmarks demonstrates that this framework achieves practical accuracy (L2 cache errors of 2.64\%, confidence-filtered L3 errors of 9.92\%, bandwidth errors of 7.33\%) while providing substantial speedup (26.8×) over cycle-accurate simulation, enabling rapid design space exploration. We demonstrate how this integrated framework enables systematic identification of both hardware optimizations (memory controller tuning, bank partitioning, NUMA configuration) and software optimizations (data layout restructuring, prefetching strategies, memory-aware scheduling). Through this comprehensive treatment of the memory-centric analysis pipeline—from trace collection through architectural modeling to co-design application—we provide researchers and practitioners with practical techniques for addressing memory bottlenecks in contemporary computing systems.

Gajaria, Dhruv Mayur↗

Distributed-Memory Sparse Deep Neural Network Inference Using Global Arrays

Partitioned Global Address Space (PGAS) models exhibit tremendous promise in developing efficient and productive distributed-memory parallel applications. They have been used extensively in scientific computations due to conveniently offering a ``shared-memory''-like model and convenient interfaces that separate communication with synchronization. Traditionally, PGAS communication models have been applied to dense/contiguously distributed data, but most modern applications depict varied levels of sparsity. Existing PGAS models require certain adaptations to support distributed sparse computations, since associated computations often require matrix arithmetic, in addition to data movement. The Global Arrays toolkit from Pacific Northwest National Laboratory (PNNL) is one of the earliest PGAS models to combine one-sided data communication and distributed matrix operations and is still used in the popular NWChem quantum chemistry suite. Recently, we have expanded the Global Arrays toolkit to support common sparse operations, like sparse matrix-dense matrix multiplies (SpMM), sparse matrix-sparse matrix multiplication (SpGEMM) and Sampled Dense-Dense Matrix Multiplication (SDDMM). As it turns out, these operations are the bedrock of sparse Deep Learning (DL); sparse deep neural networks and Graph Neural Networks (GNNs) have gained increasing attention recently in achieving speedups on training and inference with reduced memory footprints. Unlike scientific applications in High Performance Computing (HPC), modern (distributed-memory capable) DL toolkits often rely on non-standardized and closed-source vendor software optimizations, creating challenges in software-hardware co-design at scale. Our goal is to support a variety of distributed-memory sparse matrix operations and helper functions in the newly created Sparse Global Arrays (SGA), such that it is possible to build portable and productive Machine Learning scenarios for algorithm/software and hardware codesign purposes. Contemporary data-parallel schemes for training/inference are undergoing a major overhaul since model replication limits scalability and causes resource inefficiencies. As such, we have adopted tensor parallelism in decomposing the model and inputs, to mitigate memory issues. Current implementation is built on top of MPI and uses CPUs to maximize the portability across the platforms.

Distributed computing, machine learning↗

Trilinos: Enabling Scientific Computing across Diverse Hardware Architectures at Scale

Trilinos is a community-developed, open-source software framework that facilitates building large-scale, complex, multiscale, multiphysics simulation code bases for scientific and engineering problems. Since the Trilinos framework has undergone substantial changes to support new applications and new hardware architectures, this document is an update to “An Overview of the Trilinos project” by Heroux et al. (ACM Transactions on Mathematical Software, 31(3):397–423, 2005). It describes the design of Trilinos, introduces its new organization in product areas, and highlights established and new features available in Trilinos. Particular focus is put on the modernized software stack based on the Kokkos ecosystem to deliver performance portability across heterogeneous hardware architectures. This article also outlines the organization of the Trilinos community and the contribution model to help onboard interested users and contributors.

Heterogeneous Hardware Architectures↗

ECP libraries and tools: An overview

The Exascale Computing Project (ECP) Software Technology and Co-Design teams addressed the growing complexities in high-performance computing (HPC) by developing scalable software libraries and tools that leverage exascale system capabilities. As we enter the exascale era, the need for reusable, optimized software solutions that can handle the unique challenges posed by these systems becomes increasingly important. The primary challenges the ECP teams faced were to create software libraries and tools that are performant on exascale architectures and portable and usable across diverse hardware platforms. Efforts addressed issues related to concurrent execution, memory management, and the integration of heterogeneous computing resources, such as GPUs from multiple vendors. The ECP’s strategy involved a structured development process encompassing the creation, optimization, and deployment of software in collaboration with industry, academia, and national laboratories. The project was organized into several technical areas: co-design of domain-specific suites with target applications, programming models and runtimes, development tools, mathematical libraries, data and visualization tools, and software ecosystem and delivery mechanisms. ECP has successfully developed a large portfolio of software libraries and tools that demonstrate significant improvements in performance and scalability on exascale systems. These products have been integrated into the Department of Energy’s computing facilities, supporting various scientific applications and ensuring robust performance across different hardware setups. ECP advancements in software development for exascale computing highlight the importance of a collaborative and adaptive approach to handling next-generation HPC systems complexities. The lessons learned emphasize the need for continuous engagement with end-users and vendors, and the importance of maintaining a balance between innovation and practical implementation. Future efforts will focus on ensuring scalability, keeping pace with rapid hardware advancements, and further enhancing the interoperability and usability of the software ecosystem. In conclusion, subsequent articles in this special issue provide in-depth discussions and case studies into specific library and tool efforts.

97 MATHEMATICS AND COMPUTING↗

Lessons Learned and Scalability Achieved When Porting Uintah to DOE Exascale Systems

A key challenge faced when preparing codes for Department of Energy (DOE) exascale systems was designing scalable applications for systems featuring hardware and software not yet available at leadership-class scale. With such systems now available, it is important to evaluate scalability of the resulting software solutions on these target systems. One such code designed with the exascale DOE Aurora and DOE Frontier systems in mind is the Uintah Computational Framework, an open-source asynchronous many-task (AMT) runtime system. To prepare for exascale, Uintah adopted a portable MPI+X hybrid parallelism approach using the Kokkos performance portability library (i.e., MPI+Kokkos). This paper complements recent work with additional details and an evaluation of the resulting approach on Aurora and Frontier. Results are shown for a challenging benchmark demonstrating interoperability of 3 portable codes essential to Uintah-related combustion research. These results demonstrate single-source portability across Aurora and Frontier with scaling characteristics shown to 3,072 Aurora nodes and 9,216 Frontier nodes. In addition to showing results run to new scales on new systems, this paper also discusses lessons learned through efforts preparing Uintah for exascale systems.

Holmen, John [ORNL] (ORCID:0000000259342641)↗

Multinonmetal-Doped V 2 O 5 Nanocomposites for Lithium-Ion Battery Cathodes

Lithium-ion batteries (LIBs) are critical for portable electronics and electric vehicles, demanding higher energy density to meet increasing energy storage needs. Current commercial cathode materials, such as LiFePO 4 and LiCoO 2 , are limited by a single electron transfer, restricting their energy density. Vanadium pentoxide (V 2 O 5 ) emerges as a promising high-capacity cathode due to its high theoretical capacity of 443 mA h g –1 with three Li storage capacities, significantly surpassing conventional materials. However, the practical application of V 2 O 5 is hindered by a large structural evolution and rapid capacity fading during full lithium intercalation. Here, this study introduces a multinonmetal doping (MNM) strategy to enhance V 2 O 5 cathodes by incorporating all-nonmetal dopants (B, P, and Si) and graphene (G). MNM-V 2 O 5 -G exhibits increased surface oxygen defects, improving charge transfer kinetics and thus enhancing the rate performance and cycling stability. Our results provide valuable insights into the role of surface oxygen defects in stabilizing V 2 O 5 with element doping. This research highlights the potential of multinonmetal doping to improve LIB cathode materials, offering a promising pathway for design of high-energy-density V 2 O 5 cathodes and advancing the development of next-generation energy storage solutions.

25 ENERGY STORAGE↗

Challenges and Technology-Driven Opportunities for Safeguarding Microreactors

Nuclear microreactors (MRs) represent a new class of reactors characterized by their compactness, portability, and low power output. These features enable MRs to supply electricity and process heat to remote areas like military bases; inaccessible locations; small grids, such as on islands; or disaster impacted areas. Compared to traditional light water reactors, MRs have a unique set of attributes that need to be considered for the implementation of safeguard strategies. Current safeguard methodologies are reactor technology specific and are employed on large, stationary reactors where there is easy access by safeguards inspectors and where safeguard equipment can be easily installed and retrofitted. While there are numerous benefits to MRs, their compact size, portability, scalability, and operational lifetime create challenges to the traditional safeguard approaches, thus needing novel safeguard strategies. Here, this paper addresses the unique challenges posed by MRs to the international nuclear safeguards regime, including limited human resources, and explores how technology advancements can help mitigate these challenges. Specifically, it examines novel technologies that could contribute to establishing a comprehensive safeguards framework for MRs. These safeguards-enabling technologies encompass safeguards by design, remote sensing and monitoring technologies, applications of artificial intelligence and machine learning algorithms, utilization of digital twins, and system of systems assessments. While each of these safeguards-enabling technologies offers partial solutions to the challenges posed by MRs for the international safeguards regime, none of them alone can entirely address these challenges. Consequently, a combination of the safeguards-enabling technologies outlined in this paper is recommended to establish a robust safeguards regime for MRs.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Static Subspace Approximation for Random Phase Approximation Correlation Energies: Implementation and Performance

Developing theoretical understanding of complex reactions and processes at interfaces requires using methods that go beyond semilocal density functional theory to accurately describe the interactions between solvent, reactants and substrates. Methods based on many-body perturbation theory, such as the random phase approximation (RPA), have previously been limited due to their computational complexity. However, this is now a surmountable barrier due to the advances in computational power available, in particular through modern GPU-based supercomputers. In this work, we describe the implementation of RPA calculations within BerkeleyGW and show its favorable computational performance on large complex systems relevant for catalysis and electrochemistry applications. Our implementation builds off of the static subspace approximation which, by employing a compressed representation of the frequency dependent polarizability, enables the evaluation of the RPA correlation energy with significant acceleration and systematically controllable accuracy. We find that the computational cost of calculating the RPA correlation energy scales only linearly with system size for systems containing up to 50 thousand bands, and is expected to scale quadratically thereafter. We also show excellent strong scaling results across several supercomputers, demonstrating the performance and portability of this implementation.

algorithmic development↗

Performance Results on CPU/GPU Exascale Architectures for OMEGA: The Ocean Model for E3SM Global Applications

The US Department of Energy (DOE) conducts climate simulations on some of the world’s largest supercomputers. These exascale machines use heterogeneous architectures with both CPUs and GPUs, and scientific codes must adapt to make full use of this computing power. Los Alamos National Lab is developing Omega: The Ocean Model for E3SM Global Applications, which is specifically designed for modern exascale computers. It uses external libraries that have been optimized for a variety of architectures to run on different supercomputers. Omega is an unstructured-mesh ocean model based on TRiSK numerical methods. It will be the new ocean component of the DOE’s Energy Exascale Earth System Model (E3SM). The algorithms in Omega follow those of the current ocean component, MPAS-Ocean, but it will be written in C++ rather than Fortran to take advantage of the Kokkos performance portability library. Omega spatial operators are written as Kokkos kernels to run efficiently on both CPUs and GPUs. Work on Omega began in 2023 with a new C++ framework for unstructured mesh partitioning, halo exchanges, parallel IO, and Kokkos interfaces. The current version, Omega-0, is being developed to solve the shallow water equations and at present includes all of the tendency terms but not time stepping. Here we share the results of Omega-0 verification and performance testing. Verification includes unit tests implemented with CTest as well as convergence tests in Polaris, an in-house python package with a large suite of test problems. Performance tests compare simulations conducted on CPUs versus GPUs and across different architectures: tests are run on Frontier, which has AMD “Optimized 3rd Gen EPYC” CPUs and AMD MI250X GPUs, as well as Perlmutter, which is composed of AMD EPYC 7763 CPUs and NVIDIA A100 GPUs.

58 GEOSCIENCES↗

Nondestructive Analysis of Commercial Batteries

Electrochemical batteries play a crucial role for powering portable electronics, electric vehicles, large-scale electric grids, and future electric aircraft. However, key performance metrics such as energy density, charging speed, lifespan, and safety raise significant consumer concerns. Enhancing battery performance hinges on a deep understanding of their operational and degradation mechanisms, from material composition and electrode structure to large-scale pack integration, necessitating advanced characterization methods. These methods not only enable improved battery performance but also facilitate early detection of substandard or potentially hazardous batteries before they cause serious incidents. Here, this review comprehensively examines the operational principles, applications, challenges, and prospects of cutting-edge characterization techniques for commercial batteries, with a specific focus on in situ and operando methodologies. Furthermore, it explores how these powerful tools have elucidated the operational and degradation mechanisms of commercial batteries. By bridging the gap between advanced characterization techniques and commercial battery technologies, this review aims to guide the design of more sophisticated experiments and models for studying battery degradation and enhancement.

36 MATERIALS SCIENCE↗

Development of a portable high-T c spherical neutron polarimetry device at the Oak Ridge National Laboratory

Spherical neutron polarimetry is a powerful polarized neutron scattering technique used to determine complex magnetic structures which are only partly accessible by other methods. This technique measures the full neutron polarization change upon scattering from a sample by fully decoupling the incoming and outgoing neutron polarization with a zero-field chamber placed at the sample position. Recent advancements and testing are presented for a new spherical neutron polarimetry device utilizing high-T c superconducting YBCO films, PHiTPAD, at the High Flux Isotope Reactor at Oak Ridge National Laboratory. Furthermore, we introduce a conceptual design that utilizes wavelength-independent adiabatic transitions to adapt spherical neutron polarimetry for use with pulsed neutron sources, thereby expanding its potential applications in neutron scattering research.

47 OTHER INSTRUMENTATION↗