Search NASA⌕ Search

SEARCH · Search NASA

Results for “hybrid computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Taylor-Expansion-Based Robust Power Flow in Unbalanced Distribution Systems: A Hybrid Data-Aided Method

Traditional power flow methods often adopt certain assumptions designed for passive balanced distribution systems, thus lacking practicality for unbalanced operation. moreover, their computation accuracy and efficiency are heavily subject to unknown errors and bad data in measurements or prediction data of distributed energy resources (ders). to address these issues, this paper proposes a hybrid data-aided robust power flow algorithm in unbalanced distribution systems, which combines taylor series expansion knowledge with a data-driven regression technique. the proposed method initiates a linearization power flow model to derive an explicitly analytical solution by modified taylor expansion. to mitigate the approximation loss that surges due to the der integration and bad data, we further develop a data-aided robust support vector regression approach to estimate the errors efficiently. comparative analysis in the 13-bus and 123-bus ieee unbalanced feeders shows that the proposed hybrid algorithm achieves superior computational efficiency, with guaranteed accuracy and robustness against outliers.

data-driven↗

Validation of a Hybrid Domain Overlapping Coupling Between SAM and CFD Against the TALL-3D Transients

The System Thermal Hydraulics (STH) code SAM has been coupled to the Computational Fluid Dynamics (CFD) code Simcenter STAR-CCM+ utilizing a hybrid domain overlapping method with an explicit coupling in time. The coupling aims to extend the STH code’s applicability to scenarios where local momentum and energy transfers are important yet difficult for STH codes to capture, such as three-dimensional (3D) mixing. The coupling method’s numerical stability was verified in the past against two closed-loop configurations, and it was validated against a double T-junction experiment with 3D scalar mixing. In the present work, the coupling method is validated against the TALL-3D STH/CFD coupling benchmark facility. TALL-3D is a three-legged, liquid-metal facility with a large, pool-type enclosure (test section) that exhibits 3D flow effects to be modeled by a CFD code. The rest of the system exhibits approximately 1D behavior well-predicted by an STH code. First, the present STAR-CCM+ CFD model of the 3D test section is validated against experimental data. Then, the SAM-STARCCM+ coupled model is validated against six different TALL-3D steady states, including SAM standalone model results for comparison. Lastly, the SAM-STARCCM+ coupled model is validated against two TALL-3D transients, one exhibiting flow reversal in the test section and one exhibiting nonlinear, Limit Cycle Oscillations (LCO). For the first transient, the SAM-STARCCM+ coupled model properly predicts an increase in the test section’s inlet temperature during flow reversal, and this is not predicted by the SAM standalone model. Following flow reversal, the SAM-STARCCM+ coupled model better-predicts the initial flow recovery and following oscillations as the system approaches a final natural circulation state. For the second transient, no true final steady state is observed due to LCO. Neither the SAM-STARCCM+ coupled model nor the SAM standalone model can perfectly capture the experiment’s changing oscillation frequency during the transient. However, the SAM-STARCCM+ coupled model does reproduce the oscillatory feedback observed in the system. This is a significant achievement as the SAM-STARCCM+ coupled model only uses an explicit coupling in time, as opposed to a semi-implicit coupling. In comparison, previous STH/CFD coupling efforts of the TALL-3D facility required semi-implicit coupling to obtain similar results.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

HARMONY: Large-Scale Architecture Search for Efficient Hybrid Language Models

As large language models scale to trillions of parameters, their computational and memory requirements present critical challenges for efficient training and deployment. While Mixture of Experts (MoE) architectures enable efficient scaling through sparse parameter activation, and state-space models like Mamba offer linear-time complexity, principled methods for combining these paradigms remain undeveloped. We introduce HARMONY (Hybrid Architecture Research for Mamba, Optimized with Neural efficiencY), a multi-objective evolutionary neural architecture search framework for discovering efficient hybrid language models that integrate Transformer attention mechanisms, Mixture-of-Experts routing, and Mamba state-space components. Through large-scale distributed search using 16,384 MI250X GPUs on the Frontier supercomputer, HARMONY explores a comprehensive design space encompassing six attention variants (MHA, MQA, GQA, MLA, SWA, and Mamba-2), variable MoE configurations with both routed and shared experts, and extensive Mamba hyperparameters. Our framework discovers heterogeneous architectures that balance training performance with computational efficiency through multi-objective optimization incorporating latency penalties and fitness-based selection. Analysis of discovered architectures reveals that optimal hybrid designs favor heterogeneous component mixing rather than homogeneous patterns, with Mamba-2 and Multi-Head Latent Attention (MLA) emerging as preferred mechanisms. Discovered architectures demonstrate superior training efficiency: our best configuration achieves a final perplexity of 1.0874 with 2.38B parameters while processing 4,320 tokens/second, outperforming significantly larger manually designed models. Full-scale evaluation shows HARMONY's top architectures achieve better loss trajectories than equivalently-sized models using state-of-the-art configurations including Mixtral, Jamba, and Samba. Additionally, we demonstrate 91% weak scaling efficiency when training discovered 36B-parameter models across 1,024 GPUs. HARMONY is released as an open framework with comprehensive tools for building and training hybrid models using expert-data-pipeline parallelism, democratizing access to automated architecture design for next-generation language models.

Herron, Emily [ORNL] (ORCID:0000000273008172)↗

Enhancing the accuracy of XPS calculations: Exploring hybrid basis set schemes for CVS-EOMIP-CCSD calculations

Reliable computational methodologies and basis sets for modeling x-ray spectra are essential for extracting and interpreting electronic and structural information from experimental x-ray spectra. In particular, the trade-off between numerical accuracy and computational cost due to the size of the basis set is a major challenge, since molecular orbitals undergo extreme relaxation in the core-hole state. To gain clarity on the changes in electronic structure induced by the formation of a core-hole, the use of sufficiently flexible basis for expanding the orbitals, particularly for the core region, has been shown to be essential. This work focuses on the refinement of core-hole ionized state calculations using the equation-of-motion coupled cluster family of methods through an extensive analysis on the effectiveness of “hybrid” and mixed basis sets. In this investigation, we utilize the CVS-EOMIP-CCSD method in combination and construct hybrid basis sets piecewise from readily available Dunning’s correlation consistent basis sets in order to calculate x-ray ionization energies (IEs) for a set of small gas phase molecules. Our results provide insights into the impact of basis sets on the CVS-EOMIP-CCSD calculations of K-edge IEs of first-row p-block elements. Furthermore, these insights enable us to understand more about the basis set dependence of the core IEs computed and allow us to establish a protocol for deriving reliable and cost-effective theoretical estimates for computing IEs of small molecules containing such elements.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Multi-physics melt pool modeling and process optimization for laser direct energy deposition of Nb-based refractory C103: Defect formation, geometric precision, and process mapping

Recent developments in additive manufacturing (AM) technology have reignited interest in the fabrication of the Nb-based refractory C103 alloy offering solutions to the challenges posed by traditional manufacturing methods. However, the limited numerical and experimental studies on laser direct energy deposition (DED) of C103 have hindered the understanding of the relationships between process parameters and build quality. This has made it challenging to consistently produce parts with the desired quality and microstructure suitable for critical applications. In this study, we focus on optimizing the laser DED process for C103 by employing a hybrid approach that combines experimental techniques and computational fluid dynamics (CFD). This approach facilitates the development of process maps for defect detection and geometric precision. To achieve this, multi-layer C103 samples were fabricated using laser DED under various process parameters, enabling the creation of a process map for defect detection. Additionally, a multi-physics, multiphase simulation framework was developed within a high-performance computing (HPC) environment to establish process maps for geometric precision. Using these process maps, printability windows were identified for achieving both the desired geometric accuracy and defect-free prints. It was observed that prints with a power-to-velocity (P/V) ratio close to unity resulted in defect-free outcomes. This study provides a foundation for reducing design lead time and rejected parts, ultimately optimizing the laser DED process for C103.

Defect formation and geometric precision↗

Confinement-controlled selective CO 2 insertion into a dicopper dihydride core: A multiscale mechanistic study

CO 2 is an abundant C 1 feedstock for fuel and chemical synthesis. We have previously demonstrated experimentally a stepwise insertion of CO 2 into a [Cu 2 H 2 ] core via a solid–gas in crystallo reaction, forming formate species that are unstable and inaccessible under solution-phase conditions. This work elucidates how structural confinement within the crystal lattice enables such selective reactivity. In particular, co-crystallized tetrahydrofuran molecules induce site asymmetry around the [Cu 2 H 2 ] unit, modulating both the local electronic environment and CO 2 diffusion pathways. Using a multiscale computational approach that combines classical molecular mechanics, hybrid quantum mechanics/molecular mechanics molecular dynamics, and enhanced-sampling free energy calculations, we demonstrate how site asymmetry affects CO 2 binding affinities and reaction pathways. These results provide detailed mechanistic insight into CO 2 insertion and hydride transfer, highlighting key differences between crystal- and solution-phase pathways and offering a general framework for understanding how lattice confinement shapes chemical reactivity.

Chemical bonding↗

Er Al :Al 2 ⁢O 3 for telecom-band photonics: Electronic structure and optical properties

Er-doped Al 2 ⁢O 3 is a promising host for telecom-band integrated photonics. Here, in this study, we combine ab initio calculations with a symmetry-resolved analysis to elucidate substitutional Er on the Al site (Er Al ) in 𝛼−Al 2 ⁢O 3 . First-principles relaxations confirm the structural stability of Er Al . We then use the local trigonal crystal-field symmetry to classify the Er-derived impurity levels by irreducible representations and to derive polarization-resolved electric-dipole selection rules, explicitly identifying the symmetry-allowed 𝑓−𝑑 hybridization channels. Kubo-Greenwood absorption spectra computed from Kohn-Sham states quantitatively corroborate these symmetry predictions. Furthermore, we connect the calculated intra-4⁢𝑓 line strengths to Judd-Ofelt theory, clarifying the role of 4⁢𝑓−5⁢𝑑 admixture in enabling optical activity. Notably, we predict a characteristic absorption near 1.47 µ⁢m (telecom band), relevant for on-chip amplification and emission. To our knowledge, a symmetry-resolved first-principles treatment of Er:Al 2 ⁢O 3 with an explicit Judd-Ofelt interpretation has not been reported, providing a transferable framework for tailoring rare-earth dopants in wide-band-gap oxides for integrated photonics. Our results for the optical spectra are in good agreement with experimental data. The resulting symmetry-based selection rules translate directly to polarization-dependent coupling in Al 2 ⁢O 3 integrated photonic waveguides and resonators, enabling device-level design of TE/TM-mode interaction with Er emitters in the 1.5-µ⁢m telecom band.

Er-doped Al2O3↗

Performance Impact and Trade-Offs for Tuning Key Architectural Parameters on CPU+GPU Systems

In this work, we performed an initial design space exploration of an accelerated processing unit (APU)—a hybrid CPU+GPU architecture that integrates both compute units (CUs) and memory into a unified system. This integration aims to reduce data movement, enhance memory locality, and improve energy efficiency by enabling the CPU and GPU to share memory directly. This effort focused on the interplay of key design components—cache line size, the number of CUs, and main memory technology—and the trade-offs of each configuration were analyzed. This paper highlights the various configurations’ impact on memory accesses, data reuse, and power utilization. The results provide valuable insights that can be leveraged to optimize APU architectures for high-performance and energy-efficient computing and thus create a balanced architecture. This optimization can be achieved by adopting dynamic cache management, runtime CU scaling, and advanced memory integration, highlighting the potential of APUs to address critical challenges in compute, data movement, and memory power consumption.

Asifuzzaman, Kazi [ORNL] (ORCID:0000000240044791)↗

Putting error bars on density functional theory dataset

This dataset contains submission files and raw output files from high-throughput DFT simulations to analyze the systemic errors in lattice constant, bulk moduli and formation energy predictions for a range of binary and ternary oxides using four exchange correlation functionals (LDA, PBE, PBEsol and vdW-DF-C09). This data was then used as the basis for employing materials informatics methods to predict the expected errors in the lattice constants of the studied compounds. Predicted errors were also used to better the DFT-predicted lattice parameters. Our results emphasize the link between the computed errors and the electron density and hybridization errors of a functional. In essence, these results provide “error bars” for choosing a functional for the creation of high-accuracy, high-throughput datasets as well as avenues for the development of XC functionals with enhanced performance, thereby enabling the accelerated discovery and design of new materials.

36 MATERIALS SCIENCE↗

Quantum Sensing for Energy Applications

Quantum sensing is creating potentially transformative opportunities to exploit intricate quantum mechanical phenomena in new ways to make ultrasensitive measurements of multiple parameters. A growing interest in quantum sensing has created opportunities for its deployment to improve processes pertaining to energy production, distribution, and consumption. NETL is leveraging experimental and computational quantum tools to enhance sensitivity of hybrid quantum-classical ultrasensitive sensors for the detection of hydrocarbons and rare earth elements (REEs).

Paudel, Hari P.↗

A Survey on the Expanding Scope and Interdisciplinary Opportunities for Processing-in-Memory Techniques

Processing-in-Memory (PIM) is emerging as a practical path to overcome the limitations of traditional von Neumann architectures. At its core, PIM systems implement computing primitives such as logic operations and multiply-accumulate acceleration through compute-in-memory, near-memory processing, or hybrid designs. The role of memory cells varies widely across technologies, acting as inputs, outputs, or analog accumulators through bit-lines and sense amplifiers. This diversity creates trade-offs in precision, bandwidth, latency, and programmability, making it difficult to build a unified understanding on the progress of the field. In this survey, we organize recent advances of PIM into three areas. First, we discuss the progress on the architectural optimizations of PIM and its integration with both DRAM and emerging non-volatile memories. Second, we examine how PIM is being used to accelerate key computing domains, including generative AI workloads and high-performance kernels, along with new approaches. Third, we highlight the growing adoption of PIM in computational sciences, where it is being applied to solve interdisciplinary problems such as genome analysis, mRNA quantification, mass spectrometry, quantum circuit simulation, wave modeling, and secure computation. Finally, we synthesize the major challenges that continue to slow PIM adoption, including manufacturing constraints, power delivery, thermal reliability, data consistency, runtime and memory-management coordination, and the difficulty of building portable software abstractions without sacrificing commercial viability. This work provides an updated, structured perspective on PIM’s potential across computing and computational sciences and the barriers that must be solved for it to reach its full impact.

Asifuzzaman, Kazi [Oak Ridge National Laboratory (↗

Assessing VQLS for Fluid Dynamics on a Hybrid Quantum-HPC Stack

Recent advances in quantum linear solvers offer a promising direction for accelerating extreme scientific computations such as fluid dynamics. However, the deep and complex circuits required by many quantum algorithms limit their practical use on current quantum hardware. The Variational Quantum Linear Solver (VQLS) presents a viable alternative for near-term quantum devices (NISQ), and initial efforts have explored its application to select fluid dynamics problems. In this work, we evaluate the use of VQLS for canonical fluid dynamics problems, aiming to identify pathways for generalizing its implementation across a broader class of systems. We analyze the impact of various circuit ansatz and classical optimizers on solution quality and convergence behavior. Furthermore, we assess the algorithm's feasibility within a hybrid quantum–high-performance computing (HPC) framework by porting it to QFw, a state-of-the-art quantum-HPC software stack. 11This manuscript has been authored by UT-Battelle, LLC, under contract DE-AC05-00OR22725 with the US Department of Energy (DOE). The US government retains and the publisher, by accepting the article for publication, acknowledges that the US government retains a nonexclusive, paid-up, irrevocable, worldwide license to publish or reproduce the published form of this manuscript, or allow others to do so, for US government purposes. DOE will provide public access to these results of federally sponsored research in accordance with the DOE Public Access Plan. This research used resources of the Oak Ridge Leadership Computing Facility at the Oak Ridge National Laboratory, which is supported by the Office of Science of the US DOE under Contract No. DE-AC05-00OR22725.

Gopalakrishnan Meena, Murali [ORNL] (ORCID:0000000↗

How to Model Batteries (with PV, Stand-Alone, or Hybrids) in SAM and PySAM

This tutorial will be a deep dive into considerations for battery modeling and demonstrating how to model them in SAM, including battery chemistry, thermal modeling, degradation/lifetime, dispatch, interconnection limits and curtailment, and their associated impacts on project profits and battery lifetime. By the end of the tutorial attendees will know how to size and model both behind-the-meter and front-of-meter battery systems, including financial analysis and pairing with other PV models (including pvlib) via PySAM.

25 ENERGY STORAGE↗

Accurate point defect energy levels from non-empirical screened range-separated hybrid functionals: The case of native vacancies in ZnO

We use density functional theory (DFT) with non-empirically tuned screened range-separated hybrid (SRSH) functionals to calculate the electronic properties of native zinc and oxygen vacancy point defects in ZnO, and we predict their defect levels for thermal and optical transitions in excellent agreement with available experiments and prior calculations that use empirical hybrid functionals. Furthermore, the ability of this non-empirical first-principles framework to accurately predict quantities of relevance to both bulk- and defect-level spectroscopy enables high-accuracy DFT calculations with non-empirical hybrid functionals for defect physics, at a reduced computational cost.

Defects↗

Parallel Programming in MCNP6

Monte Carlo N-Particle (MCNP)1 is a general-purpose Monte Carlo particle transport code developed by Los Alamos National Laboratory (LANL). To efficiently handle long simulations, MCNP version 6 (MCNP6) supports parallel execution using two primary programming models: • Shared-memory task-based threading using OpenMP (Open Multi-Processing), and • Distributed-memory calculations using MPI (Message Passing Interface). The OpenMP and MPI programming models enable MCNP6 to scale from desktop systems to high-performance computing (HPC) clusters, allowing users to run MCNP in one of three parallel modes: • OpenMP-only, • MPI-only, and • Hybrid (MPI + OpenMP). The choice of parallelization mode depends on the underlying computer architecture and the characteristics of the simulation problem.

97 MATHEMATICS AND COMPUTING↗

Differentiable hybrid neural network approach for enhancing reactor dynamics simulations

Reactor dynamics simulations provide essential insights into the time-dependent behavior of nuclear reactors under various operating conditions. However, high-fidelity simulations can be computationally intensive, requiring significant computational resources. Here, to address this challenge, this study employs a differentiable hybrid model that utilizes neural networks as a corrector to enhance the performance of a low-fidelity simulation, aligning its predictions with those of a high-fidelity simulation. Low-fidelity and high-fidelity simulations were obtained by adjusting the mesh size in the System Dynamics Analysis Tool. The differentiable hybrid model was trained in two approaches: time-step-wise and sequence-wise. It was then applied to simulate various transients in a molten salt reactor. Its performance was evaluated by comparing its responses to transients against those of the high-fidelity simulation. An additional approach was performed using a data-driven model to correct the low-fidelity simulation. In comparison, the differentiable hybrid model showed significant improvements in transient prediction, effectively addressing the limitations of the low-fidelity simulations. The results highlighted the robustness of the differentiable hybrid model in both training approaches. It delivered simulations that were at least 3.8 times faster than high-fidelity models. In the time-step-wise approach, it achieved at least a 39% improvement in accuracy. In the sequence-wise approach, it showed at least an 81% accuracy improvement over the full transient. This approach offers a promising path for improving computational efficiency without compromising accuracy in nuclear reactor simulations, making it suitable for real-time digital twin applications.

42 - ENGINEERING↗

First nucleon gluon PDF from large momentum effective theory

We report the first nucleon gluon parton distribution function (PDF) using Large-Momentum Effective Theory (LaMET). We focus on the gluon operator which was demonstrated to have the best signal-to-noise in the previous attempt [1] in computing gluon PDFs using LaMET. We compute the corresponding Wilson coefficients needed for the hybrid-renormalized matrix elements and the matching kernel to convert the quasi-PDF to the lightcone one at the one-loop level. We demonstrate that with the proper Wilson coefficients in place, the counterterms for the renormalization are independent of the hadron and mass within statistical error. Using the resulting renormalization, we then compute the nucleon PDF using a HISQ ensemble generated by the MILC collaboration with N f = 2 + 1 +1, a ≈ 0.12 fm, with valence pion masses of 310 and 690 MeV and two gauge link smearing techniques. Despite the physics effects of the heavier than physical pion masses and gauge link smearing, this calculation provides excellent proof of principle and compares reasonably with selected global fit results.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Scaling kinetic Monte-Carlo simulations of grain growth with combined convolutional and graph neural networks

Graph neural networks (GNN) have emerged as a promising machine learning method for microstructure simulations such as grain growth. However, accurate modeling of realistic grain boundary networks requires large simulation cells, which GNN has difficulty scaling up to. To alleviate the computational costs and memory footprint of GNN, we suggest a hybrid architecture combining a convolutional neural network (CNN) based bijective autoencoder to compress the spatial dimensions, and a GNN that evolves the microstructure in the latent space of reduced spatial sizes. Our results demonstrate that the new design significantly reduces computational costs with using fewer message passing layer (from 12 down to 3) compared with GNN alone. The reduction in computational cost becomes more pronounced as the spatial size increases, indicating strong computational scalability. For the largest mesh evaluated (160 3 ), our method reduces memory usage and runtime in inference by 117× and 115×, respectively, compared with GNN-only baseline. More importantly, it shows higher accuracy and stronger spatiotemporal capability than the GNN-only baseline, especially in long-term testing. Such combination of scalability and accuracy is essential for simulating realistic material microstructures over extended time scales. The improvements can be attributed to the bijective autoencoder’s ability to compress information losslessly from spatial domain into a high dimensional feature space, thereby producing more expressive latent features for the GNN to learn from, while also contributing its own spatiotemporal modeling capability. Training data are generated from stochastic grain growth simulations, providing realistic variability for learning robust microstructure evolution. Comprehensive system validation confirms that the model is accurate, robust, and scalable.

36 MATERIALS SCIENCE↗