Search NASA⌕ Search

SEARCH · Search NASA

Results for “Computer systems performance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39

A Flexible Quasi-Static Mooring Design Optimization Method for Floating Structures

This paper presents a flexible and efficient design method for optimizing the mooring systems of floating structures. Mooring system optimization is challenging because of the strong nonlinearity of mooring system behavior and the many technical constraints that must be satisfied. Furthermore, different mooring configurations can have very different design spaces. While some successful examples of mooring design optimization exist in the literature, developing an optimization approach that can work across various mooring design problems is a larger challenge. We present such a method based on a flexible parameterization that allows a wide variety of mooring designs to be described by a list of variables, a quasi-static mooring model that provides efficient evaluation of a mooring design without directly considering mooring system dynamics, and an optimization framework that generates, evaluates, and adjusts the mooring design while considering user-specified constraints such as offset limits, strength safety factors, and seabed contact limits. We demonstrate the design optimization framework on four mooring design problems, each for a different type of mooring system. We compare the use of different design modes to simplify the optimization problem, showing that they can reduce the computation time by up to 75%. We also compare different optimization algorithms and find that the resulting computational speed can vary by up to 51 times. We perform a sensitivity study on one design and find that the local sensitivity of anchoring radius to water depth has a positive correlation of 0.29, but the global sensitivity shows large nonlinearities. Lastly, we perform a coupled dynamic analysis on one of the optimized designs and find that the predicted mean platform motions and mooring line tensions are within 1% of dynamic results and the extreme motions and tensions are within 14%. Lastly, we show that a DEA-Chain-Polyester mooring configuration is cost-optimal for the given design problem of the demonstrations, which aligns with general industry practice.

16 TIDAL AND WAVE POWER↗

Modeling of H2 Dispersion at ARIES

Hydrogen is a versatile and clean energy carrier that can be produced from various renewable sources such as wind, solar, and hydropower and help decarbonize electricity grids, industry, and transportation. Using the Hydrogen Research Facility under Advanced Research on Integrated Energy Systems (ARIES) at the National Renewable Energy Laboratory's (NREL) Flatirons campus as a test bench, the study examines the feasibility, useability, and value of using computational fluid dynamics (CFD) techniques to model hydrogen dispersion. The ARIES facility was chosen because controlled hydrogen releases can be performed at a rate of 27 kg-H2/hr. Site-specific atmospheric and weather condition data such as wind speed and temperature were used as inputs to the model. The results show statistical distributions and ranges of hydrogen concentrations at locations throughout the domain. Wind conditions are found to significantly impact the release behavior, including the hydrogen cloud's direction and concentrations. At low wind speeds (below 1 mph), hydrogen forms a cloud and at higher wind speeds (> 2-4 mph) hydrogen plume stretches in the direction of wind momentum. From >100 simulations for ARIES site-specific conditions, statistical quantities combined with a clustering algorithm were used to propose sensor location at various elevations from ground.

dispersion↗

A Deep Learning-Driven Sampling Technique to Explore the Phase Space of an RNA Stem-Loop

The folding and unfolding of RNA stem-loops are critical biological processes; however, their computational studies are often hampered by the ruggedness of their folding landscape, necessitating long simulation times at the atomistic scale. Here, we adapted DeepDriveMD (DDMD), an advanced deep learning-driven sampling technique originally developed for protein folding, to address the challenges of RNA stem-loop folding. Although tempering- and order parameter-based techniques are commonly used for similar rare-event problems, the computational costs or the need for a priori knowledge about the system often present a challenge in their effective use. DDMD overcomes these challenges by adaptively learning from an ensemble of running MD simulations using generic contact maps as the raw input. DeepDriveMD enables on-the-fly learning of a low-dimensional latent representation and guides the simulation toward the undersampled regions while optimizing the resources to explore the relevant parts of the phase space. We showed that DDMD estimates the free energy landscape of the RNA stem-loop reasonably well at room temperature. Our simulation framework runs at a constant temperature without external biasing potential, hence preserving the information on transition rates, with a computational cost much lower than that of the simulations performed with external biasing potentials. Here, we also introduced a reweighting strategy for obtaining unbiased free energy surfaces and presented a qualitative analysis of the latent space. This analysis showed that the latent space captures the relevant slow degrees of freedom for the RNA folding problem of interest. Finally, throughout the manuscript, we outlined how different parameters are selected and optimized to adapt DDMD for this system. We believe this compendium of decision-making processes will help new users adapt this technique for the rare-event sampling problems of their interest.

Gupta, Ayush↗

Quantum simulations of nuclear resonances with variational methods

Background: The many-body nature of nuclear physics problems poses significant computational challenges. These challenges become even more pronounced when studying the resonance states of nuclear systems, which are governed by the non-Hermitian Hamiltonian. Quantum computing, particularly for quantum many-body systems, offers a promising alternative, especially within the constraints of current noisy intermediate-scale quantum (NISQ) devices. Purpose: This work aims to simulate nuclear resonances using quantum algorithms by developing a variational framework compatible with non-Hermitian Hamiltonians and implementing it fully on a quantum simulator. Methods: We employ the complex scaling technique to extract resonance positions classically and adapt it for quantum simulations using a two-step algorithm. First, we transform the non-Hermitian Hamiltonian into a Hermitian form by using the energy variance as a cost function within a variational framework. Second, we perform 𝜃-trajectory calculations to determine optimal resonance positions in the complex energy plane. To address resource constraints on NISQ devices, we utilize Gray code (GC) encoding to reduce qubit requirements. Results: We first validate our approach using a schematic potential model that mimics a nuclear potential, successfully reproducing known resonance energies with high fidelity. We then extend the method to a more realistic 𝛼−𝛼 nuclear potential and compute the 𝐷- and 𝐺-wave resonance energies with a basis size of 𝑁=16, using only four qubits. The quantum simulation results closely match the classical values, demonstrating the feasibility of our approach. Conclusions: This study demonstrates, for the first time, that the complete 𝜃-trajectory method can be implemented on a quantum computer without relying on any classical input beyond the Hamiltonian. The results establish a scalable and efficient quantum framework for simulating resonance phenomena in nuclear systems. This work represents a significant step toward quantum simulations of open quantum systems and lays the foundation for future investigations into resonance structures in nuclear, atomic, and molecular physics.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Automated ICRF heating surrogate modeling via machine learning

This work introduces automated machine learning workflows that address critical bottlenecks in surrogate model development for Ion Cyclotron Range of Frequencies (ICRF) heating applications. The automated framework includes data analysis tools that transform raw datasets into actionable insights in seconds, replacing weeks of manual exploratory effort and ensuring consistent, reproducible dataset characterization. By integrating advanced hyperparameter optimization (HPO) methods including Bayesian optimization via BoTorch and Tree-structured Parzen Estimators (TPE), the framework significantly reduces model development time from weeks to hours, decreasing computational cost and required expertise, while enabling high-accuracy surrogate models. Compared to traditional hyperparameter scanning (HPS) techniques such as methodical, randomized, and grid searches, HPO methods achieve superior convergence and predictive performance, even when compared to already well-tuned reference models. On NSTX High Harmonic Fast Wave (HHFW) heating datasets, both Random Forest Regressor (RFR) and neural network surrogates demonstrate improved accuracy, achieving R 2 values beyond 0.97 and 0.98, respectively. The results show that while HPO gains are modest for robust architectures like RFR, they become essential for more sensitive models such as neural networks, highlighting the trade-offs across optimization strategies. Through automated workflows that eliminate manual hyperparameter tuning and require minimal ML expertise, this work enables widespread adoption of high-fidelity surrogate models across the fusion community for real-time plasma control, uncertainty quantification, rapid experimental scenario development, and integrated system optimization.

Sanchez-Villar, Alvaro [Princeton Plasma Physics L↗

Analytic Neural Network Gaussian Process Enabled Chance-Constrained Voltage Regulation for Active Distribution Systems with PVs, Batteries and EVs

This paper proposes an analytic neural network Gaussian process (NNGP)-based chance-constrained real-time voltage regulation method for active distribution systems with photovoltaics (PVs), batteries, and electric vehicles (EVs). NNGP can utilize historical measurement data to achieve real-time probabilistic node voltage estimation through Bayesian inference. Then, NNGP is fully analytically embedded into the optimal power flow model to perform voltage regulation and adapt to various topological changes. The uncertainties of voltage estimations are easily considered via the chance constraint, and it has been shown that the adoption of this chance constraint can significantly improve the reliability of voltage regulation under various scenarios. The comparison results with other methods, carried out on a real 759-node distribution system located in western Colorado, U.S., show that the proposed method can achieve accurate voltage estimation across different topologies and reliably perform voltage regulation considering PVs, batteries, and EVs.

active distribution systems↗

Development of Machine-Learned Interatomic Potentials to Predict Structure, Transport, and Reactivity in Platinum-Based Fuel Cells

Machine-learned interatomic potentials (MLIPs) have rapidly progressed in accuracy, speed, and data efficiency in recent years. However, training robust MLIPs in multicomponent systems remains a challenge. In this work, we train an MLIP to describe hydrated Nafion ionomers and platinum catalysts, which are important components of fuel cells, by constructing a diverse training set to describe the bulk polymer and interfacial catalyst–polymer interactions well. We use our trained MLIP to study the properties of the platinum–Nafion system, including polymer structure, proton mobility in a bulk Nafion polymer and near a platinum-Nafion interface, and reactions near and far from the interface, finding excellent results for structure and reactions contained within our training set. Transport seems to be well described, with both vehicular transport and Grotthuss hopping captured, although converged calculations of diffusivities were not computed because they require calculations of tens of nanoseconds that are challenging with current state-of-the-art MLIPs. The combined insights that this model provides can be leveraged to optimize fuel cell performance, and the approach can be applied to other chemical processes and devices where structure, transport, and reactivity all contribute to the overall observed performance.

33 ADVANCED PROPULSION SYSTEMS↗

An Evaluation of the Effect of Network Cost Optimization for Leadership Class Supercomputers

Dragonfly-based networks are an extensively deployed network topology in large-scale high-performance computing due to their cost-effectiveness and efficiency. The US will soon have three Exascale supercomputers for leadership class workloads deployed using dragonfly networks. Compared to indirect networks of similar scale, the dragonfly network has considerably reduced cable lengths, cable counts, and switch counts, resulting in significant network cost savings for a given system size, however, these cost reductions result in reduced global minimal paths and more challenging routing. Additionally, large scale dragonfly networks often require a taper at the global link level, resulting in less bisection bandwidth than is achievable in other traditional non-blocking topologies of equivalent scale. While dragonfly networks have been extensively studied, they have yet to be fully evaluated in an extreme scale (i.e., exascale) system that targets capability workloads. In this paper, we present the results of the first large scale evaluation of a dragonfly network on an exascale system (Frontier) and compare its behavior to a similar scale fat-tree network on a previous generation TOP500 system (Summit). This evaluation aims to determine the effect of network cost optimizations by measuring a tapered topology’s impact on capability workloads. Our evaluation is based on a collection of synthetic microbenchmarks, mini-apps, and full scale applications. It compares the scaling efficiencies of each benchmark between the dragonfly-based Frontier and the fat-tree-based Summit systems. Our results show that a dragonfly network is $\sim \mathbf{3 0 \%}$ more cost efficient than a fat-tree topology, which amortizes to $\sim 3 \%$ of an exascale system cost. Furthermore, while tapered dragonfly networks impose significant tradeoffs, the impacts are not as broad as initially thought and are mostly seen in applications with global communication patterns, particularly all-to-all (e.g., FFT-based algorithms), but also local communication patterns (e.g., nearest-neighbor algorithms) that are sensitive to network performance variability.

Khan, Awais↗

Collisional Excitation of HCN by CO to Refine the Modeling of Cometary Comae

Here, we present the first dataset of collisional (de)-excitation rate coefficients of HCN induced by CO, one of the main perturbing gases in cometary atmospheres. The dataset spans the temperature range of 5–50 K. It includes both state-to-state rate coefficients involving the lowest ten and nine rotational levels of HCN and CO, respectively, and the so-called “thermalized” rate coefficients over the rotational population of CO at each kinetic temperature. The derivation of these coefficients exploited the good performance of the statistical adiabatic channel model (SACM) on top of an accurate interaction potential computed at the CCSD(T)-F12b/CBS level of theory. The reliability of the SACM approach was validated by comparison with full quantum calculations restricted at the lowest total angular momentum of the system. These results provide essential input to accurately model the distribution among the rotational energy levels and the abundance of HCN in cometary atmospheres, accounting for deviations from local thermodynamic equilibrium that typically occurs in such environments.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

When in-memory computing meets spiking neural networks—A perspective on device-circuit-system-and-algorithm co-design

This review explores the intersection of bio-plausible artificial intelligence in the form of spiking neural networks (SNNs) with the analog in-memory computing (IMC) domain, highlighting their collective potential for low-power edge computing environments. Through detailed investigation at the device, circuit, and system levels, we highlight the pivotal synergies between SNNs and IMC architectures. Additionally, we emphasize the critical need for comprehensive system-level analyses, considering the inter-dependencies among algorithms, devices, circuit, and system parameters, crucial for optimal performance. An in-depth analysis leads to the identification of key system-level bottlenecks arising from device limitations, which can be addressed using SNN-specific algorithm–hardware co-design techniques. This review underscores the imperative for holistic device to system design-space co-exploration, highlighting the critical aspects of hardware and algorithm research endeavors for low-power neuromorphic solutions.

Physics↗

Distributed Multi-GPU Community Detection on Exascale Computing Platforms

Community detection is a fundamental operation in graph mining, and by uncovering hidden structures and patterns within complex systems it helps solve fundamental problems pertaining to social networks, such as information diffusion, epidemics, and recommender systems. Scaling graph algorithms for massive networks becomes challenging on modern distributed-memory multi-GPU (Graphics Processing Unit) systems due to limitations such as irregular memory access patterns, load imbalances, higher communication-computation ratios, and cross-platform support. We present a novel algorithm HiPDPL-GPU (distributed parallel Louvain) to address these challenges. We conduct experiments involving different partitioning techniques to achieve optimized performance of HiPDPL-GPU on the two largest supercomputers: Frontier and Summit. Remarkably, HiPDPL-GPU processes a graph with 4.2 billion edges in less than 3 minutes using 1024 GPUs. Qualitatively performance of HiPDPL-GPU is similar or better compared to other state-of-the-art CPU- and GPU-based implementations. While prior GPU implementations have predominantly employed CUDA, our first-of-its-kind implementation for community detection is cross-platform, accommodating both AMD and NVIDIA GPUs.

graph algorithms, high performance comptuing↗

Portable Software Environment for Ultrahigh-Resolution ELM Development on GPUs

This paper presents our endeavors in developing the large-scale, ultra-high-resolution E3SM Land Model (uELM), specifically designed for exascale computers furnished with accelerators such as Nvidia GPUs. The uELM is a sophisticated code that substantially relies on High-Performance Computing (HPC) environments, necessitating particular machine and software configurations. To facilitate community-based uELM developments employing GPUs, we have created a portable, standalone software environment preconfigured with uELM input datasets, simulation cases, and source code. This environment, utilizing Docker, encompasses all essential code, libraries, and system software for uELM development on GPUs. It also features a functional unit test framework and an offline model testbed for comprehensive numerical experiments. From a technical perspective, the paper discusses GPU-ready container generations, uELM code management, and input data distribution across computational platforms. Lastly, the paper demonstrates the use of environment for functional unit testing, end-to-end simulation on CPUs and GPUs, and collaborative code development.

E3SM Land Model↗

Circuit-Based Leakage-to-Erasure Conversion in a Neutral-Atom Quantum Processor

Atom-loss errors are a major limitation of current state-of-the-art neutral-atom quantum computers and pose a significant challenge for scalable systems. In a quantum processor with cesium atoms, we demonstrate proof-of-principle circuit-based conversion of this form of leakage error to erasure errors via leakage-detection units (LDUs), which nondestructively map information about the presence or absence of the qubit onto the state of an ancilla. We benchmark the performance of the LDU using a three-outcome low-loss state-detection method and find that the LDU detects atom-loss errors with approximately 93.4% accuracy, limited by technical imperfections of our apparatus. We further compile and execute a SWAP LDU, wherein the roles of the original data atom and ancilla atom are exchanged under the action of the LDU, providing “free refilling” of atoms in the case of atom loss. This circuit-based leakage-to-erasure error conversion is a critical component of a neutral-atom quantum processor where the quantum information may significantly outlive the lifetime of any individual atom in the quantum register. Finally, we demonstrate that LDUs may also be used to handle other forms of leakage errors where population moves to states outside of the computational subspace.

Chow, Matthew N. H. [Sandia National Laboratories ↗

First-Principles Studies on Sc 2 RuZ (Z = Si, Ge, Sn) Inverse Heusler Alloys: Structural, Electronic, and Transport Properties

The continuous demand for efficient, nontoxic, and thermally stable materials for room-temperature energy conversion motivates the exploration of novel thermoelectric systems beyond the traditional magnetic Heusler alloys. While full and half-Heusler compounds, especially Co-, Ni-, and Mn-based systems, have demonstrated promising thermoelectric properties, their typically high operating temperatures and magnetic complexities limit their applicability in ambient thermal management. In this context, we investigate whether Sc-based inverse Heusler alloys can offer a viable nonmagnetic alternative with competitive thermoelectric performance. In this work, we perform a systematic first-principles study of the inverse Heusler compounds Sc 2 RuZ (Z = Si, Ge, Sn), focusing on their structural, electronic, mechanical, and thermodynamic-thermoelectric properties. Density Functional Theory (DFT) was employed to compute optimized lattice structures and band dispersion, while dynamical stability was assessed via phonon calculations. Thermoelectric transport coefficients, including Seebeck coefficient, electrical conductivity, and thermal conductivity, were estimated using the semiclassical Boltzmann transport theory within the constant relaxation time approximation. Our results show that all Sc 2 RuZ compounds are thermodynamically stable semiconductors with indirect band gaps of 0.12–0.16 eV and exhibit high elastic moduli, especially Sc 2 RuSn, which demonstrates superior stiffness and incompressibility. Importantly, all compounds display promising room-temperature thermoelectric characteristics, including high Seebeck coefficients and power factors. These findings reveal that Sc 2 RuZ alloys represent a rare class of stable, nonmagnetic inverse Heusler semiconductors with intrinsic thermoelectric potential at room temperature, unlike many existing Heusler systems optimized for spintronics or high-temperature operation. This work expands the known design space for Heusler-based thermoelectrics and offers a theoretical basis for experimental realization of efficient, low-temperature, nonmagnetic thermoelectric materials.

alloys↗

Numerical eigen-spectrum slicing, accurate orthogonal eigen-basis, and mixed-precision eigenvalue refinement using OpenMP data-dependent tasks and accelerator offload

Performing a variety of numerical computations efficiently and, at the same time, in a portable fashion requires both an overarching design followed by a number of implementation strategies. All of these are exemplified below as we present transitioning the PLASMA numerical library from relying on dependence-driven large tasks to achieving utilization of fine grain tasking and offload to hardware accelerators while keeping its core dependence sets: OpenMP source code pragmas and runtime for most system-level functionality and basic low-level numerical kernels provided directly by hardware vendors or open source projects with vendor contributions. We also present new algorithmic methods and their efficient parallel implementations including fine grained tasking for eigen-spectrum slicing and offload for mixed-precision eigenvalue refinement. We provide performance, scaling, and numerical results showing sizable gains over the available solutions from either the open source and vendor-provided packages.

Luszczek, Piotr↗

Thermodynamic modeling of CsF with LiF-NaF-KF for molten fluoride-fueled reactors

Gibbs energy models were developed to describe the thermochemical behavior of CsF in molten FLiNaK (46.5LiF-11.5NaF-42KF mol%), a proposed molten salt reactor (MSR) fuel solvent and coolant, as cesium is of concern due to its high radiotoxicity and volatility. Initially, it was necessary to obtain a more accurate Gibbs energy function for CsF which required fitting parameters to reported vapor pressures over condensed phase CsF. The pseudo-binary systems CsF-LiF, CsF-NaF and CsF-KF were then evaluated utilizing phase equilibria and enthalpy of mixing (Δ mix H) values, together with original differential scanning calorimetry (DSC) measurements performed for the CsF-LiF and CsF-KF systems. The CsF-LiF-NaF, CsF-LiF-KF and CsF-NaF-KF pseudo-ternary system representations were obtained by interpolation of the constituent pseudo-binary systems, with DSC measurements performed for the CsF-LiF-NaF system to corroborate the calculated liquidus temperature. Ultimately, the pseudo-ternary systems were interpolated to obtain Gibbs energy models for the pseudo-quaternary CsF-LiF-NaF-KF system, supported by DSC measurements at low CsF compositions (1–10 mol%), yielding computed equilibria and cesium-containing vapor pressures that compare favorably with reported values. In conclusion, the Molten Salt Thermal Properties Database – Thermochemical (MSTDB-TC) was subsequently expanded to include these Gibbs energy models allowing description of the thermochemical behavior of the CsF-LiF-NaF-KF system.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

From molecular to macroscopic: predicting liquid–liquid phase equilibria and small-angle scattering of mixtures of organic liquids from atomistic simulation using Kirkwood–Buff theory

Macroscopic phase equilibria between solutions define the functionality of many biological and industrial processes, yet they are challenging to predict due to the inherent complexity of liquids containing large molecules. This work introduces an approach for the purely predictive calculation of such phase equilibria in temperature-composition space from molecular dynamics (MD) simulations at one temperature in the single-phase region. We use an approach developed previously to obtain the entropic and enthalpic contributions to the free energy of mixing from the atomic-scale information given by MD simulations via Kirkwood–Buff theory. This allows us to accurately estimate the free energy of mixing as a function of temperature, and thus obtain liquid–liquid phase equilibria, including liquid–liquid critical points, associated binodal and spinodal lines, and composition fluctuations across a region of temperature and composition. Results for binary malonamide–alkane systems are validated by comparison to a direct experimental probe of the fluctuations: the small angle X-ray scattering intensity near zero wavenumber. The MDKB → Phase method demonstrated here provides a significant improvement in predicting liquid–liquid equilibria and free energy as a function of temperature for our systems of interest compared to conventional thermodynamic models. The accurate performance of this purely predictive approach lies in its preservation of atomistic details when determining thermodynamic properties. Furthermore, its inherent extensibility to multi-component systems will likely make the MDKB → Phase approach a valuable general tool for connecting molecular interactions to macroscopic phase equilibria and for the computational screening of materials for targeted thermodynamic behavior.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Scalable Multi-Modal Framework for High-Fidelity Distributed Human Mobility Simulations

The development of data-driven models for human mobility in urban settings requires access to substantial and diverse real-world data. However, existing historical data often presents challenges such as limited volume, variety, and veracity, as well as missing data and privacy preservation concerns. Also, urban mobility modeling is inherently time-variant, complex, and multi-modal, encompassing everything from individual walking and running to private road travel and large-scale public transportation. These challenges call for innovative solutions to overcome data limitations and compute needs to model mobility behaviors accurately. To address these challenges, we propose a distributed, co-simulation-based architecture DURMOSim that integrates real-world data with scalable, high-fidelity simulations, demonstrating distributed co-simulation feasibility with existing mobility models. DURMOSim underpins a modular integration that would enable using any available mobility simulators for greater extensibility and scalability in performing various urban scenarios. In this paper, we present the design, implementation, and performance evaluation of DURMOSim, highlighting its capability to model population-scale mobility patterns. Our initial results show its ability to dynamically synchronize multiple simulation models at runtime with negligible computational overhead. We believe DURMOSim could be a robust tool for advancing urban mobility research and intelligent transportation systems.

Yoginath, Srikanth [ORNL] (ORCID:0000000184236050)↗