Search NASA⌕ Search

SEARCH · Search NASA

Results for “computing frameworks”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32

Small tensor product distributed active space (STP-DAS) framework for relativistic and non-relativistic multiconfiguration calculations: Scaling from 10 9 on a laptop to 10 12 determinants on a supercomputer

Despite the power and flexibility of configuration interaction (CI) based methods in computational chemistry, their broader application is limited by an exponential increase in both computational and storage requirements, particularly due to the substantial memory needed for excitation lists that are crucial for scalable parallel computing. Here, the objective of this work is to develop a new CI framework, namely, the small tensor product distributed active space (STP-DAS) framework, aimed at drastically reducing memory demands for extensive CI calculations on individual workstations or laptops, while simultaneously enhancing scalability for extensive parallel computing. Moreover, the STP-DAS framework can support various CI-based techniques, such as complete active space (CAS), restricted active space, generalized active space, multireference CI, and multireference perturbation theory, applicable to both relativistic (two- and four-component) and non-relativistic theories, thus extending the utility of CI methods in computational research. We conducted benchmark studies on a supercomputer to evaluate the storage needs, parallel scalability, and communication downtime using a realistic exact-two-component CASCI (X2C-CASCI) approach, covering a range of determinants from 10 9 to 10 12 . Additionally, we performed large X2C-CASCI calculations on a single laptop and examined how the STP-DAS partitioning affects performance.

Complete-active space self-consistent field↗

MetaHeuristic Feature Selection for Energy Group Optimization and Analysis

Energy discretization is a crucial component of deterministic neutron transport simulations. Metaheuristic (MH) optimizers are effective algorithms to determine group structures that maximize both solution accuracy and computational efficiency. This project establishes a framework for optimizing group structures for PARTISN simulations using the Python library MEALPY. Group structure optimization is formulated as a binary feature selection problem, and results are investigated with permutation and material importance techniques to determine physically relevant energy bounds. We conclude that MH optimizers find group structures that drastically improve flux calculations while preserving k-effective accuracy. Further, we find that individual energy bounds are not necessarily physically relevant, but rather specific energy ranges are.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Understanding Mixed Precision GEMM with MPGemmFI: Insights into Fault Resilience

Emerging deep learning workloads urgently need fast general matrix multiplication (GEMM). Thus, one of the critical features of machine-learning-specific accelerators such as NVIDIA Tensor Cores, AMD Matrix Cores, and Google TPUs is the support of mixed-precision enabled GEMM. For DNN models, lower-precision FP data formats and computation offer acceptable correctness but significant performance, area, and memory footprint improvement. While promising, the mixed-precision computation on error resilience remains unexplored. To this end, we develop a fault injection framework that systematically injects fault into the mixed-precision computation results. We investigate how the faults affect the accuracy of machine learning applications. Based on the characteristics of error resilience, we offer lightweight error detection and correction solutions that significantly improve the overall model accuracy by 75% if the models experience hardware faults. The solutions can be efficiently integrated into the accelerator's pipelines.

Fang, Bo↗

Position Paper - pFLogger: The Parallel Fortran Logging framework for HPC Applications

In the context of high performance computing (HPC), software investments in support of text-based diagnostics, which monitor a running application, are typically limited compared to those for other types of IO. Examples of such diagnostics include reiteration of configuration parameters, progress indicators, simple metrics (e.g., mass conservation, convergence of solvers, etc.), and timers. To some degree, this difference in priority is justifiable as other forms of output are the primary products of a scientific model and, due to their large data volume, much more likely to be a significant performance concern. In contrast, text-based diagnostic content is generally not shared beyond the individual or group running an application and is most often used to troubleshoot when something goes wrong. We suggest that a more systematic approach enabled by a logging facility (or logger) similar to those routinely used by many communities would provide significant value to complex scientific applications. In the context of high-performance computing, an appropriate logger would provide specialized support for distributed and shared-memory parallelism and have low performance overhead. In this paper, we present our prototype implementation of pFlogger a parallel Fortran-based logging framework, and assess its suitability for use in a complex scientific application.

Fortran↗

POSITION PAPER - pFLogger: The Parallel Fortran Logging Framework for HPC Applications

In the context of high performance computing (HPC), software investments in support of text-based diagnostics, which monitor a running application, are typically limited compared to those for other types of IO. Examples of such diagnostics include reiteration of configuration parameters, progress indicators, simple metrics (e.g., mass conservation, convergence of solvers, etc.), and timers. To some degree, this difference in priority is justifiable as other forms of output are the primary products of a scientific model and, due to their large data volume, much more likely to be a significant performance concern. In contrast, text-based diagnostic content is generally not shared beyond the individual or group running an application and is most often used to troubleshoot when something goes wrong. We suggest that a more systematic approach enabled by a logging facility (or 'logger') similar to those routinely used by many communities would provide significant value to complex scientific applications. In the context of high-performance computing, an appropriate logger would provide specialized support for distributed and shared-memory parallelism and have low performance overhead. In this paper, we present our prototype implementation of pFlogger - a parallel Fortran-based logging framework, and assess its suitability for use in a complex scientific application.

Clune, Thomas L.↗

Privacy Preserving Federated Learning for Advanced Scientific Ecosystems

We present a framework to provide privacy preserving (PP) federating learning (FL) across multiple computational and experimental facilities. This work joins the compute capabilities of National Energy Research Scientific Computing Center (NERSC) and Oak Ridge National Laboratory Research Cloud (ORC) with simulated experimental data, such as those produced at the SLAC National Accelerator Laboratory and Spallation Neutron Source (SNS). We describe the software infrastructure developed to provide privacy for computational and experimental networks. We developed algorithmic privacy across the federated system by embedding database security, computation, and communication into the federation architecture, utilizing scientific tools developed by the experimental community.

Archibald, Rick [ORNL] (ORCID:0000000245389780)↗

Development of an Aeroelastic Modeling Capability for Transient Nozzle Side Load Analysis

Lateral nozzle forces are known to cause severe structural damage to any new rocket engine in development during test. While three-dimensional, transient, turbulent, chemically reacting computational fluid dynamics methodology has been demonstrated to capture major side load physics with rigid nozzles, hot-fire tests often show nozzle structure deformation during major side load events, leading to structural damages if structural strengthening measures were not taken. The modeling picture is incomplete without the capability to address the two-way responses between the structure and fluid. The objective of this study is to develop a coupled aeroelastic modeling capability by implementing the necessary structural dynamics component into an anchored computational fluid dynamics methodology. The computational fluid dynamics component is based on an unstructured-grid, pressure-based computational fluid dynamics formulation, while the computational structural dynamics component is developed in the framework of modal analysis. Transient aeroelastic nozzle startup analyses of the Block I Space Shuttle Main Engine at sea level were performed. The computed results from the aeroelastic nozzle modeling are presented.

Wang, Ten-See↗

Development of an Aeroelastic Modeling Capability for Transient Nozzle Side Load Analysis

Lateral nozzle forces are known to cause severe structural damage to any new rocket engine in development during test. While three-dimensional, transient, turbulent, chemically reacting computational fluid dynamics methodology has been demonstrated to capture major side load physics with rigid nozzles, hot-fire tests often show nozzle structure deformation during major side load events, leading to structural damages if structural strengthening measures were not taken. The modeling picture is incomplete without the capability to address the two-way responses between the structure and fluid. The objective of this study is to develop a coupled aeroelastic modeling capability by implementing the necessary structural dynamics component into an anchored computational fluid dynamics methodology. The computational fluid dynamics component is based on an unstructured-grid, pressure-based computational fluid dynamics formulation, while the computational structural dynamics component is developed in the framework of modal analysis. Transient aeroelastic nozzle startup analyses of the Block I Space Shuttle Main Engine at sea level were performed. The computed results from the aeroelastic nozzle modeling are presented.

Wang, Ten-See↗

A generative artificial intelligence framework for long-time plasma turbulence simulations

Generative deep learning techniques are employed in a novel framework for the construction of surrogate models capturing the spatiotemporal dynamics of 2D plasma turbulence. The proposed Generative Artificial Intelligence Turbulence (GAIT) framework enables the acceleration of turbulence simulations for long-time transport studies. GAIT leverages a convolutional variational auto-encoder and a recurrent neural network to generate new turbulence data from existing simulations, extending the time horizon of transport studies with minimal computational cost. The application of the GAIT framework to plasma turbulence using the Hasegawa–Wakatani (HW) model is presented, evaluating its performance via various analyses. Very good agreement is found between the GAIT and the HW models in the spatiotemporal Fourier and Proper Orthogonal Decomposition spectra, the flow topology characterized by the Okubo–Weiss parameter, and the time autocorrelation function of turbulent fluctuations. Excellent agreement has also been obtained in the probability distribution function of particle displacements and the effective turbulent diffusivity. In-depth analyses of the latent space of turbulent states, choice of hyperparameters and alternative deep learning models for the time prediction are presented. Our results highlight the potential of Artificial Intelligence-based surrogate models to overcome the computational challenges in turbulence simulation, which can be extended to other situations such as geophysical fluid dynamics.

Artificial intelligence↗

Gate-Based Quantum Simulation of Gaussian Bosonic Circuits on Exponentially Many Modes

We introduce a framework for simulating, on an ( n + 1 )-qubit quantum computer, the action of a Gaussian bosonic (GB) circuit on a state over 2 n modes. Specifically, we encode the initial bosonic state’s expectation values over quadrature operators (and their covariance matrix) as an input qubit state. This is then evolved by a quantum circuit that effectively implements the symplectic propagators induced by the GB gates. We find families of GB circuits and initial states leading to efficient quantum simulations. For this purpose, we introduce a dictionary that maps between GB and qubit gates such that particle- (non-particle-) preserving GB gates lead to real- (imaginary-) time evolutions at the qubit level. For the special case of particle-preserving circuits, we present a bounded-error-quantum-polynomial time (BQP)-complete GB decision problem, indicating that GB evolutions of Gaussian states on exponentially many modes are as powerful as universal quantum computers. We also perform numerical simulations of an interferometer on ∼ 8 × 10 9 modes, illustrating the power of our framework. Published by the American Physical Society 2025

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

An End-to-End Framework for Verifying and Validating Manufacturing Design Integrity

Cyber attacks on networked automated manufacturing systems can severely impact part quality. In fact, malicious modifications may be introduced at any point during the manufacturing lifecycle. Therefore, it is vital to verify and validate that manufactured parts conform to their designs. This chapter describes a formal, end-to-end framework that verifies and validates the design integrity of manufactured parts by considering all potential points of alteration during precision manufacturing processes. The framework prevents unauthorized changes to computer-aided designs, verifies the correctness of translations from CAD models to G-code, maintains the integrity of G-code transferred to manufacturing machines, verifies the runtime execution of G-code and part geometry, and considers the contexts of manufacturing machine operations and how manufactured parts could be altered.

Jablonski, Matthew [Cybersecurity Manufacturing In↗

A modified scheil approach for nucleation-dependent solidification pathways

As-solidified microstructures of near-eutectic alloys often contain multiple primary phases that are not expected from equilibrium phase diagrams. Such microstructures are caused by cooling-rate-dependent solidification pathways, a factor not captured by the Scheil–Gulliver model or variations thereof. Here, we present a model and algorithm that incorporate the critical nucleation undercooling for each solid phase into the Scheil–Gulliver model. We hypothesize that the non-equilibrium microstructure formation is primarily governed by a nucleation-competition mechanism. This mechanism accounts for both stable/metastable phase selection and primary-phase formation within eutectic regions driven by asymmetric nucleation barriers. The model is validated against a hypereutectic Al-Fe alloy, where it successfully reproduces the observed microstructural constituents, revealing the key dependencies of solidification microstructure on nucleation kinetics. Applicability to multicomponent systems is demonstrated through a hypereutectic Al–Fe–Si ternary alloy, where the model successfully predicts divorced eutectic microstructures and the associated oscillatory solidification pathways along univariant lines. As a result, the proposed framework establishes a nucleation-dependent computational approach for interpreting and predicting solidification microstructures.

Alloy design↗

Physics-informed latent neural operator for real-time predictions of time-dependent parametric PDEs

Deep operator network (DeepONet) has shown significant promise as surrogate models for systems governed by partial differential equations (PDEs), enabling accurate mappings between infinite-dimensional function spaces. However, when applied to systems with high-dimensional input-output mappings arising from large numbers of spatial and temporal collocation points, these models often require heavily overparameterized networks, leading to long training times. Latent DeepONet addresses some of these challenges by introducing a two-step approach: first learning a reduced latent space using a separate model, followed by operator learning within this latent space. While efficient, this method is inherently data-driven and lacks mechanisms for incorporating physical laws, limiting its robustness and generalizability in data-scarce settings. Here, in this work, we propose PI-Latent-NO, a physics-informed latent neural operator framework that integrates governing physics directly into the learning process. Our architecture features two coupled DeepONets trained end-to-end: a Latent-DeepONet that learns a low-dimensional representation of the solution, and a Reconstruction-DeepONet that maps this latent representation back to the physical space. By embedding PDE constraints into the training via automatic differentiation, our method eliminates the need for labeled training data and ensures physics-consistent predictions. The proposed framework is both memory and compute-efficient, exhibiting near-constant scaling with problem size and demonstrating significant speedups over traditional physics-informed operator models. We validate our approach on a range of parametric PDEs, showcasing its accuracy, scalability, and suitability for real-time prediction in complex physical systems.

Latent representations↗

“Best” Iterative Coupled-Cluster Triples Model? More Evidence for 3CC

To follow up on the unexpectedly good performance of several coupled-cluster models with approximate inclusion of 3-body clusters we performed a more complete assessment of the 3CC method for accurate computational thermochemistry in the standard HEAT framework. New spin-integrated implementation of the 3CC method applicable to closed- and open-shell systems utilizes a new automated toolchain for derivation, optimization, and evaluation of operator algebra in many-body electronic structure. We found that with a double-ζ basis set the 3CC correlation energies and their atomization energy contributions are almost always more accurate (with respect to the CCSDTQ reference) than the CCSDT model as well as the standard CCSD(T) model. The mean absolute errors in cc-pVDZ {3CC, CCSDT, and CCSD(T)} electronic (per valence electron) and atomization energies relative to the CCSDTQ reference for the HEAT data set, were {24, 70, 122} μE h /e and {0.46, 2.00, 2.58} kJ/mol, respectively. The mean absolute errors in the complete-basis-set limit {3CC, CCSDT, and CCSD(T)} atomization energies relative to the HEAT model reference, were {0.52, 2.00, and 1.07} kJ/mol, The significant and systematic reduction of the error by the 3CC method and its lower cost than CCSDT suggests it as a viable candidate for post- CCSD(T) thermochemistry applications, as well as the preferred alternative to CCSDT in general.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Non-resonant Raman optical activity from phase-space electronic structure theory

In order to model experimental non-resonant Raman optical activity, chemists must compute a host of second-order response tensors (e.g., the electric-dipole–magnetic-dipole polarizability) and their nuclear derivatives along a set of vibrational modes. While these response functions are almost always computed within a Born–Oppenheimer (BO) framework, here we provide a natural interpretation of the electric-dipole–magnetic-dipole polarizability within phase space electronic structure theory, a beyond-BO model whereby the electronic structure depends on nuclear momentum (P) in addition to nuclear position (R). By coupling to nuclear momentum, phase space electronic structure theory is able to capture the asymmetric response of the electronic properties to an external field, in so far as for a vibrating (non-stationary) molecule, $\frac{∂μ}{∂B}$≠$\frac{∂m}{∂F}$, where μ and m are the electrical linear and magnetic dipoles, and F and B are electric and magnetic fields. As an example, for a prototypical methyloxirane molecule, we show that phase space electronic structure theory is able to deliver a reasonably good match with experimental results in a manner that is formally invariant to gauge origin G 0 —provided that one uses a complete basis or, alternatively, gauge invariant atomic orbitals.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Including the vacuum energy in stellarator coil design

Being three-dimensional, stellarators have the advantage that plasma currents are not essential for creating rotational-transform; however, the external current-carrying coils in stellarators can have strong geometrical shaping, which can complicate the construction. Reducing the inter-coil electromagnetic forces acting on strongly shaped 3D coils and the stress on the support structure while preserving the favorable properties of the magnetic field is a design challenge. In this work, we recognize that the inter-coil forces are the gradient of the vacuum magnetic energy. We introduce an objective functional built on the usual quadratic flux on a prescribed target surface together with a weighed penalty on the vacuum energy. The Euler–Lagrange equation for stationary states is derived, and numerical illustrations are computed using a modern stellarator optimization framework. A study of the effect of the energy functional on the inter-coil forces is conducted and the energy is shown to be a promising quantity in producing coils with low forces.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Hadronic light-by-light contribution to the muon anomaly from lattice QCD with infinite volume QED at physical pion mass

The hadronic light-by-light scattering contribution to the muon anomalous magnetic moment, (g–2)⁢/2, is computed in the infinite volume QED framework with lattice QCD. We report $a^{HLbL}_μ$ = 12.47⁢(1.15)⁢(0.95) ×10 –10 where the first error is statistical and the second systematic. The result is mainly based on the 2+1 flavor Möbius domain wall fermion ensemble with inverse lattice spacing a –1 = 1.73 GeV, lattice size L = 5.5 fm, and m π = 139 MeV, generated by the RBC-UKQCD collaborations. The leading systematic error of this result comes from the lattice discretization. This result is consistent with previous determinations.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Coupled Lindblad Pseudomode Theory for Simulating Open Quantum Systems

Coupled Lindblad pseudomode theory is a promising approach for simulating non-Markovian quantum dynamics on both classical and quantum platforms, with dynamics that can be realized as a quantum channel. We provide theoretical evidence that the number of coupled pseudomodes only needs to scale as polylog⁡(𝑇/𝜖) in the simulation time 𝑇 and precision 𝜖. Inspired by the realization problem in control theory, we also develop a robust numerical algorithm for constructing the coupled modes that avoid the nonconvex optimization required by existing approaches. We demonstrate the effectiveness of our method by computing population dynamics and absorption spectra for the spin-boson model. Furthermore, this Letter provides a significant theoretical and computational improvement to the coupled Lindblad framework, which impacts a broad range of applications from classical simulations of quantum impurity problems to quantum simulations on near-term quantum platforms.

Anderson impurity model↗