Search NASA⌕ Search

SEARCH · Search NASA

Results for “Computer Systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Advanced Computing is at the Forefront of a New “Moonshot” Revolutionizing the North American Power Grid

In the 50+ years since the first humans landed on the moon, computing has grown at breakneck speed. We are faced with another challenge that is just as daunting, and just as important to overcome-modernizing the North American electric power grid-and high-performance computing (HPC) systems with specialized software will be an important element in rising to this challenge. We describe at a high level how software developed in the ExaSGD project addresses this "moonshot" goal by utilizing exascale computing and a novel high performance solver software stack to support the mission of decarbonizing power grid operations in an environment of uncertain weather and climate. To reach the exascale benchmark the team has made a number of first-of-their-kind innovations, including novel method for stochastic optimization, fine grained parallel methods for modeling power systems, and GPU resident sparse numerical linear solvers.

17 WIND ENERGY↗

System for controller area network payload decoding

A system for decoding an unknown automotive controller area network (“CAN”) message definitions. CAN data vehicle signal mappings are typically held in secret and varied by automotive model and year. Without knowledge of the mappings, the wealth of real-time vehicle data hidden in the automotive CAN packets is uninterpretable—impeding research, after-market tuning, efficiency and performance monitoring, fault diagnosis, and privacy-related technologies. This system can ascertain the CAN signals' boundaries (start bit and length), endianness (byte ordering), signedness (binary-to-integer encoding) from raw CAN data. This allows conversion of CAN data to time series. Interpreting the translated CAN data's physical meaning and finding a linear mapping to standard units (e.g., knowing the signal is speed and scaling values to represent units of miles per hour) can be achieved for many signals by leveraging diagnostic standards to obtain real-time measurements of in-vehicle systems. The system can be integrated into lightweight hardware enabling an OBD-II plugin for real-time in-vehicle CAN decoding or run on standard computers. The system can output a standard DBC file with the signal definition information.

Verma, Kiren E.↗

I/O in Machine Learning Applications on HPC Systems: A 360-degree Survey

Growing interest in Artificial Intelligence (AI) has resulted in a surge in demand for faster methods of Machine Learning (ML) model training and inference. This demand for speed has prompted the use of high performance computing (HPC) systems that excel in managing distributed workloads. Because data is the main fuel for AI applications, the performance of the storage and I/O subsystem of HPC systems is critical. In the past, HPC applications accessed large portions of data written by simulations or experiments or ingested data for visualizations or analysis tasks. ML workloads perform small reads spread across a large number of random files. This shift of I/O access patterns poses several challenges to modern parallel storage systems. In this paper, we survey I/O in ML applications on HPC systems, and target literature within a 6-year time window from 2019 to 2024. We define the scope of the survey, provide an overview of the common phases of ML, review available profilers and benchmarks, examine the I/O patterns encountered during offline data preparation, training, and inference, and explore I/O optimizations utilized in modern ML frameworks and proposed in recent literature. Lastly, we seek to expose research gaps that could spawn further R&D.

97 MATHEMATICS AND COMPUTING↗

Advanced Computing Annual Report 2025 [Slides]

In Fiscal Year (FY) 2025, the National Laboratory of the Rockies (NLR) continued to advance computing as a cornerstone of energy innovation, expanding the Kestrel high-performance computing (HPC) system to 56 peak petaflops. This growth strengthened Kestrel's role as a national asset for applied energy research, enabling larger, more complex simulations and accelerating the integration of artificial intelligence (AI) methods across the laboratory's computing portfolio. In FY 2025, AI was a component of most projects running on Kestrel, underscoring its central role in modern energy science and engineering. Kestrel supported a broad and diverse set of 507 modeling and simulation projects, engaging 855 researchers across the U.S. Department of Energy's (DOE's) Office of Critical Minerals and Energy Innovation (CMEI) portfolio and other offices, as well as partners from industry, academia, and utilities. These efforts span critical materials discovery, energy systems modeling, grid modernization, advanced manufacturing, and other areas essential to strengthening U.S. energy security and competitiveness. Together, these collaborations produced 708 technical outputs, including 293 peer-reviewed publications, reflecting both the depth and impact of the science enabled by NLR's computing capabilities. This year's report highlights the growing importance and benefit of AI throughout NLR's research programs and features work by early career researchers who are helping shape the future of computing-enabled energy innovation. Explore these sections and the many project successes captured in the pages that follow.

97 MATHEMATICS AND COMPUTING↗

A Quantum Computational Determination of the Weak Mixing Angle in the Standard Model

The weak mixing angle $s_W$ is a fundamental constant in the Standard Model (SM) and measured at the $Z$ boson mass to be $\widehat{s}^2_W(m_Z) = 0.23129 \pm 0.00004$ in the $\overline{\rm MS}$ renormalization scheme, where $m_Z=91.2\ \text{GeV}$. On the other hand, non-stabilizerness - the magic - characterizes the computational advantage of a quantum system over classical computers. We consider the production of magic from stabilizer initial states, which carry zero magic, in the 2-to-2 scattering of charged leptons in the SM at the tree level, which is mediated by the photon and the $Z$ boson. Using the second order stabilizer Rényi entropy, and averaging over all 60 initial stabilizer states and the scattering angle, we compute and minimize the magic production as a function of $s^2_W$ in the Møller scattering $e^-e^-\to e^-e^-$, which is free of kinematic thresholds. At the centre-of-mass energy $\sqrt{s}=m_Z$, there is a unique minimum in magic production at $\mathbf{s}^{2}_W(m_Z)=0.2317$, which agrees with the measured $\widehat{s}^2_W(m_Z)$ at the sub-percent level. At higher energies, the magic-minimizing $\mathbf{s}^{2}_W$ continues to agree with the empirical value at the percent level or better, up to 10 TeV. The finding suggests the electroweak sector of the SM tends to generate minimal quantum resources from the computational viewpoint.

Liu, Qiaofeng [Northwestern U.]↗

A Quantum Computational Determination of the Weak Mixing Angle in the Standard Model

The weak mixing angle $s_W$ is a fundamental constant in the Standard Model (SM) and measured at the $Z$ boson mass to be $\widehat{s}^2_W(m_Z) = 0.23129 \pm 0.00004$ in the $\overline{\rm MS}$ renormalization scheme, where $m_Z=91.2\ \text{GeV}$. On the other hand, non-stabilizerness - the magic - characterizes the computational advantage of a quantum system over classical computers. We consider the production of magic from stabilizer initial states, which carry zero magic, in the 2-to-2 scattering of charged leptons in the SM at the tree level, which is mediated by the photon and the $Z$ boson. Using the second order stabilizer Rényi entropy, and averaging over all 60 initial stabilizer states and the scattering angle, we compute and minimize the magic production as a function of $s^2_W$ in the Møller scattering $e^-e^-\to e^-e^-$, which is free of kinematic thresholds. At the centre-of-mass energy $\sqrt{s}=m_Z$, there is a unique minimum in magic production at $\mathbf{s}^{2}_W(m_Z)=0.2317$, which agrees with the measured $\widehat{s}^2_W(m_Z)$ at the sub-percent level. At higher energies, the magic-minimizing $\mathbf{s}^{2}_W$ continues to agree with the empirical value at the percent level or better, up to 10 TeV. The finding suggests the electroweak sector of the SM tends to generate minimal quantum resources from the computational viewpoint.

Liu, Qiaofeng [Northwestern U.]↗

A Quantum Computational Determination of the Weak Mixing Angle in the Standard Model

The weak mixing angle $s_W$ is a fundamental constant in the Standard Model (SM) and measured at the $Z$ boson mass to be $\widehat{s}^2_W(m_Z) = 0.23129 \pm 0.00004$ in the $\overline{\rm MS}$ renormalization scheme, where $m_Z=91.2\ \text{GeV}$. On the other hand, non-stabilizerness - the magic - characterizes the computational advantage of a quantum system over classical computers. We consider the production of magic from stabilizer initial states, which carry zero magic, in the 2-to-2 scattering of charged leptons in the SM at the tree level, which is mediated by the photon and the $Z$ boson. Using the second order stabilizer Rényi entropy, and averaging over all 60 initial stabilizer states and the scattering angle, we compute and minimize the magic production as a function of $s^2_W$ in the Møller scattering $e^-e^-\to e^-e^-$, which is free of kinematic thresholds. At the centre-of-mass energy $\sqrt{s}=m_Z$, there is a unique minimum in magic production at $\mathbf{s}^{2}_W(m_Z)=0.2317$, which agrees with the measured $\widehat{s}^2_W(m_Z)$ at the sub-percent level. At higher energies, the magic-minimizing $\mathbf{s}^{2}_W$ continues to agree with the empirical value at the percent level or better, up to 10 TeV. The finding suggests the electroweak sector of the SM tends to generate minimal quantum resources from the computational viewpoint.

Liu, Qiaofeng [Northwestern U.]↗

Data-driven discovery of dynamics from time-resolved coherent scattering

Coherent X-ray scattering (CXS) techniques are capable of interrogating dynamics of nano- to mesoscale materials systems at time scales spanning several orders of magnitude. However, obtaining accurate theoretical descriptions of complex dynamics is often limited by one or more factors—the ability to visualize dynamics in real space, computational cost of high-fidelity simulations, and effectiveness of approximate or phenomenological models. In this work, we develop a data-driven framework to uncover mechanistic models of dynamics directly from time-resolved CXS measurements without solving the phase reconstruction problem for the entire time series of diffraction patterns. Our approach uses neural differential equations to parameterize unknown real-space dynamics and implements a computational scattering forward model to relate real-space predictions to reciprocal-space observations. This method is shown to recover the dynamics of several computational model systems under various simulated conditions of measurement resolution and noise. Moreover, the trained model enables estimation of long-term dynamics well beyond the maximum observation time, which can be used to inform and refine experimental parameters in practice. Finally, we demonstrate an experimental proof-of-concept by applying our framework to recover the probe trajectory from a ptychographic scan. Our proposed framework bridges the wide existing gap between approximate models and complex data.

36 MATERIALS SCIENCE↗

Synergistic Solvent-Surface Interactions Enable Alkyne Semihydrogenation at Palladium

Enabling higher yield and better selectivity for fine-chemical synthesis through heterogeneous catalysis is intricately linked to the interplay of active sites, reaction conditions, and mass transfer influence provided by the catalyst. Alkyne semihydrogenation is ubiquitous in the production of bulk chemicals in the pharmaceutical, polymer, or fine-chemical industries, but product selectivity remains a major challenge. Here, in this study, we demonstrate that the design of catalysts encompassing nickel (Ni) foams as contiguous monolith supports, decorated with ultralow loading of Pd/PdO x nanoparticles on a carbonized polydopamine interface and tuned with a thin layer of Al 2 O 3 , in conjunction with an optimized reaction environment leads to highly selective alkyne semihydrogenation. The reactions demonstrate good functional group tolerance and applicability to flow reactor systems. Combined computational and experimental studies are presented to describe the synergistic effect between the solvent-surface interaction and the degree of Pd surface reduction that are necessary to promote this selectivity. The system highlights the opportunity for catalyst-solvent codesign as a benign alternative to more complex reactants featuring extrinsic poisons or less-favored dopants.

atomic layer deposition↗

Developing an Energy-Conscious Traffic Signal Control System for Optimized Fuel Consumption in Connected Vehicle Environments

The project titled “Developing an Energy-Conscious Traffic Signal Control System for Optimized Fuel Consumption in Connected Vehicle Environments” addresses energy-related challenges associated with adaptive traffic control systems by integrating connected vehicles (CV) and connected infrastructure (CI). The system developed in this project, a CV-based adaptive traffic control system, aims to improve fuel consumption in mixed traffic environments by capitalizing on emerging CV and CI communication technologies, as well as leveraging recent advances in Artificial Intelligence (AI), optimization, and edge computing. The system was tested at the MLK Smart Corridor, an urban testbed managed by the University of Tennessee at Chattanooga (UTC) and the City of Chattanooga. The system was validated through extensive simulations, both Software-in-the-Loop (SILS) and Hardware-in-the-Loop (HILS), and was further implemented and tested in real-world conditions at several intersections along the corridor. The Fuel Consumption Performance Index (FC-PI) and the Ecological Performance Index (Eco-PI) were developed as the key components for evaluating the system’s impact on fuel consumption and emissions. These metrics provided a comprehensive means of understanding the impact of traffic signal control optimization in mixed traffic environments. The report presents an in-depth analysis of the Eco-PI, FC-PI, adaptive traffic control system integration, and the testing and field implementation of the system. The results demonstrate significant reductions in fuel consumption and emissions, showcasing the system’s capability to contribute to more sustainable urban traffic management. The report also documents the challenges encountered and recommendations for scaling and further improving the system.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Testing convolutional neural network based deep learning systems: a statistical metamorphic approach

Machine learning technology spans many areas and today plays a significant role in addressing a wide range of problems in critical domains,i.e., healthcare, autonomous driving, finance, manufacturing, cybersecurity,etc. Metamorphic testing (MT) is considered a simple but very powerful approach in testing such computationally complex systems for which either an oracle is not available or is available but difficult to apply. Conventional metamorphic testing techniques have certain limitations in verifying deep learning-based models (i.e., convolutional neural networks (CNNs)) that have a stochastic nature (because of randomly initializing the network weights) in their training. In this article, we attempt to address this problem by using a statistical metamorphic testing (SMT) technique that does not require software testers to worry about fixing the random seeds (to get deterministic results) to verify the metamorphic relations (MRs). We propose seven MRs combined with different statistical methods to statistically verify whether the program under test adheres to the relation(s) specified in the MR(s). We further use mutation testing techniques to show the usefulness of the proposed approach in the healthcare space and test two CNN-based deep learning models (used for pneumonia detection among patients). The empirical results show that our proposed approach uncovers 85.71% of the implementation faults in the classifiers under test (CUT). Furthermore, we also propose an MRs minimization algorithm for the CUT, thus saving computational costs and organizational testing resources.

Computer Science↗

Visualizing an Exascale Data Center Digital Twin: Considerations, Challenges and Opportunities

Digital twins are an excellent tool to model, visualize, and simulate complex systems, to understand and optimize their operation. In this work, we present the technical challenges of real-time visualization of a digital twin of the Frontier supercomputer.We show the initial prototype and current state of the twin and highlight technical design challenges of visualizing such a large High Performance Computing (HPC) system. The goal is to understand the use of augmented reality as a primary way to extract information and collaborate on digital twins of complex systems. This leverages the spatio-temporal aspect of a 3D representation of a digital twin, with the ability to view historical and real-time telemetry, triggering simulations of a system state and viewing the results, which can be augmented via dashboards for details. Finally, we discuss considerations and opportunities for augmented reality of digital twins of large-scale, parallel computers.

Maiterth, Matthias↗

Quantum/AI Topology-Aware Latency-Adaptive HPC Workflow Scheduling Optimization

The growing demand for more powerful high-performance computing (HPC) systems has led to a steady rise in energy consumption by supercomputing worldwide. This study is focused on comparing our Application-Topology Mapper (ATMapper) to the popular Simple Linux Utility for Resource Management (SLURM) for the purpose of exploring methods that can further optimize job-scheduling within HPC systems. ATMapper is an Artificial-Intelligence based approach to job-scheduling that is currently being enhanced with quantum annealing (QA) to generate optimal schedules faster. We are applying QA to speedup our ATMapper process to achieve higher computing efficiency, thereby reducing HPC energy consumption. Here, we examine how four job-scheduling approaches perform in processor node assignment when using an example network architecture of 4 interconnected nodes. Using a specialized script, we are assessing the schedule of a computation flow with 11 interdependent tasks. The data movements among nodes were tracked to count for the number of interactions (network hops) between nodes needed to complete the tasks. The total number of hops and the job completion time were then used to quantify the efficiency of the different mapping approaches. In addition to SLURM, we also compare our ATMapper to the QA-enabled LBNL TIGER and the D-Wave Distributed Computing processor assignment approaches. The preliminary results showed that our topology-aware, latency-adaptive ATMapper is significantly more efficient when compared to the other scheduling approaches due to its load-imbalance network allocation. The scheduler displayed a computing efficiency of 53% by performing significantly fewer network hops than its alternatives. By reducing the number of hops, ATMapper was able to perform all 11 tasks by using only 3 nodes out of given 4. This research indicates the potential to use QA/AI for HPC job-scheduling. Later, we will test a SLURM simulator program to draw further comparisons on the effectiveness of ATMapper's scheduling approach. The results of this comparison will serve as a baseline for later improving SLURM's performance using a QA-enhanced ATMapper approach.

Caraveo, Braulio [University of Huston - Clear Lak↗

A Performance Model of In-Situ Techniques

The computational capacity of High-Performance Computing (HPC) systems increases continuously with the rapid development of central processing units (CPUs) and graphic processing units (GPUs), while the in-/output (IO) subsystem develops relatively slowly and storage capacity is also limited. Data-intensive applications, which are designed to leverage the high computational capacity of HPC resources, typically generate a considerable amount of data for post-processing visualizations and data analytics. The limited IO speed and storage space could lead to constraints in the actual performance of these applications and, therefore, scientific discovery. In-situ techniques, where data is visualized/analysed while still in memory rather than through disk, can contribute to alleviating these problems as they can reduce or even fully avoid data writing/reading through the IO subsystem to/from storage. However, the overall efficiency of insitu techniques crucially depends on the characteristics of both the in-situ tasks and the applications, and the resource distribution among them. Therefore, choosing the right in-situ approach (synchronous, asynchronous, or hybrid) and resource allocation is essential to minimize overhead and maximize the benefits of concurrent execution. In this paper, we present a performance model of in-situ techniques to find the most beneficial in-situ approach and the preferred resource configuration. We verify the high accuracy of our approach with over 6800 measurements and provide use cases with different applications.

Ju, Yi [Max Planck Computing and Data Facility, Ga↗

Celeritas Midterm SciDAC Report

Celeritas is a new Monte Carlo (MC) code that helps satisfy the increasing demand for high energy physics (HEP) detector simulation, using Graphics Processing Unit (GPU) hardware on high performance computing (HPC) systems to model Large Hadron Collider (LHC) experiments and beyond. This report details the project’s progress midway through its SciDAC funding period, highlighting the first complete implementation of standard electromagnetic (EM) physics on GPUs, initial results for performance and scalability on Leadership Computing Facilities (LCFs), and preliminary integration into the CMS and ATLAS experiments. By integrating HEP domain knowledge with expertise in MC transport, Celeritas has catalyzed a shift in the HEP community’s perception of GPU platforms as the future for HPC simulations.

97 MATHEMATICS AND COMPUTING↗

R-matrix calculations for opacities: I. Methodology and computations

Abstract An extended version of the R -matrix methodology is presented for calculation of radiative parameters for improved plasma opacities. Contrast and comparisons with existing methods primarily relying on the distorted wave approximation are discussed to verify accuracy and resolve outstanding issues, particularly with reference to the opacity project (OP). Among the improvements incorporated are: (i) large-scale Breit–Pauli R -matrix calculations for complex atomic systems including fine structure, (ii) convergent close coupling wave function expansions for the ( e + ion) system to compute oscillator strengths and photoionization cross sections, (iii) open and closed shell iron ions of interest in astrophysics and experiments, (iv) a treatment for plasma broadening of autoionizing resonances as function of energy-temperature-density dependent cross sections, (v) a ‘top-up’ procedure to compare convergence with R -matrix calculations for highly excited levels, and (vi) spectroscopic identification of resonances and bound ( e + ion) levels. The present R -matrix monochromatic opacity spectra are fundamentally different from OP and lead to enhanced Rosseland and Planck mean opacities. An outline of the work reported in other papers in this series and those in progress is presented. Based on the present re-examination of the OP work, opacities of heavy elements might require revisions in high temperature-density plasma sources.

Pradhan, A. K. (ORCID:0000000187753643)↗

OLCF Test Harness

Acceptance and regression testing of a High Performance Computing (HPC) system requires an automated and reproducible framework and tool for running and logging results. Manually running tests across a system is labor intensive and prone to reproducibility errors. The OLCF Test Harness (OTH) provides a framework in which to document required tests for a HPC system. The OTH then provides tools to execute and log results of these tests in an automated fashion.

Dietz, Dan [Oak Ridge National Laboratory (ORNL), ↗

Creating Apptainer Workflows with Docker-Compose-like Utilities

Creating Apptainer Workflows with Docker-Compose-like Utilities In this presentation, I will explore the utilization of a tool called process-compose, inspired by docker-compose, to create Apptainer-based services. This approach allows for easy deployment and management of fully containerized applications on High Performance Computing (HPC) systems without requiring elevated privileges. Benefits to the Ecosystem: By incorporating process-compose and Apptainer, I aim to address several key challenges in the HPC ecosystem: Simplified Workflow Management: Process-compose provides a user-friendly interface for defining and managing complex containerized application services, reducing the setup time and lowering the barrier to entry for new users. Enhanced Portability: Apptainer ensures that containerized applications can run consistently across different HPC environments, promoting greater portability and reducing compatibility issues. Process-compose is also a single binary that does not need to be installed by admin level users. Community Driven Solutions: This approach aligns with the goals of the High Performance Software Foundation (HPSF) to advance community-driven solutions. By sharing our experiences and insights, I hope to foster collaboration and innovation within the HPC community. Increased Productivity: The combination of process-compose and Apptainer streamlines the serve deployment process, allowing researchers and developers to focus more on their scientific work rather than the intricacies of system or service administration. Through this presentation, attendees will gain valuable insights into the practical implementation of containerized workflows on HPC systems, learn about the benefits of using process-compose and Apptainer, and understand how these tools can contribute to a more efficient HPC ecosystem.

97 - MATHEMATICS AND COMPUTING↗