Search NASA⌕ Search

SEARCH · Search NASA

Results for “coded computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

High- and Mid-Fidelity Modeling Comparison for a Floating Marine Turbine System: Preprint

There is a lack of suitable numerical tools, particularly open-source tools, that can be used for designing and optimizing marine turbine systems. The National Renewable Energy Laboratory (NREL) has added features to their widely used mid-fidelity wind turbine modeling code, OpenFAST, to enable modeling of axial flow marine turbines. This necessitated the addition of several physical effects relevant to marine turbines that can be neglected for wind turbines. These include buoyant loads, added mass and inertia loads, wave-current superposition, and changes to the coordinate systems. This updated version of OpenFAST allows for the modeling of both fixed and floating marine turbine systems at a speed comparable to real time. While efficient for long simulations, large sets of load cases, and design studies, mid-fidelity codes cannot capture all of the potentially important physical phenomenon impacting marine turbine systems. High-fidelity computational fluid dynamics (CFD) simulations can capture more flow effects with fewer assumptions and provide detailed body pressure mapping and flow-field information. It is important to compare predictions between mid-fidelity and high-fidelity codes, both to verify the models and to understand the limitations. A floating marine turbine system designed by NREL was modeled both with OpenFAST and with the commercial CFD code, STARCCM+. The CFD model used a 3-D unsteady Reynolds-averaged Navier-Stokes solver for a volume-of-fluid numerical wave and current tank. The blade-resolved simulations used the sliding-interface technique for the spinning rotor and an overset grid to accommodate the rigid-body motion of the floating system. The mooring system was modelled with a custom coupling of the CFD solver with the open-source code, MoorDyn. This improves upon the existing quasi-static catenary solver in STARCCM+, which lacks seabed contact or line-to-line connections. Spatial and temporal convergence studies were conducted. The simulation results for a combined current and wave condition are compared between OpenFAST and CFD, highlighting the capabilities of the mid-fidelity code and identifying the areas where a high-fidelity approach is needed.

CFD↗

gRASPA

GPU Monte Carlo Simulation Code with a taste of RASPA We present enhancements in Monte Carlo simulation speed and functionality within an open-source code, gRASPA, which uses graphical processing units (GPUs) to achieve significant performance improvements compared to serial, CPU implementations of Monte Carlo. The code supports a wide range of Monte Carlo simulations, including canonical ensemble (NVT), grand canonical, NVT Gibbs, Widom test particle insertions, and continuous-fractional component Monte Carlo. Implementation of grand canonical transition matrix Monte Carlo (GC-TMMC) and a novel feature to allow different moves for the different components of metal-organic framework (MOF) structures exemplify the capabilities of gRASPA for precise free energy calculations and enhanced adsorption studies, respectively. The introduction of a High-Throughput Computing (HTC) mode permits many Monte Carlo simulations on a single GPU device for accelerated materials discovery. The code can incorporate machine learning (ML) potentials. The open-source nature of gRASPA promotes reproducibility and openness in science, and users may add features to the code and optimize it for their own purposes. The code is written in CUDA/C++ and SYCL/C++ to support different GPU vendors. The gRASPA code is publicly available at https://github.com/snurr-group/gRASPA.

Li, Zhao [Purdue/Northwestern/Notre Dame Universit↗

Integration of the NCRC Database and Other INL Databases

The Nuclear Computational Resource Center provides a portal by which industry professionals, educational staff, students, national laboratory employees, and others may request access to certain engineering software tools. As the tools provided through the Nuclear Computational Resource Center portal are not open-source and freely available, a set of approvals are necessary before access is granted. All code recipients must be associated with an institution that has a license with Idaho National Laboratory for the code requested. Information about these licenses is controlled by Idaho National Laboratory’s Technology Deployment organization and housed in a Technology Deployment database. Those requesting code access who are not citizens of the United States must also have a security plan, mandated by Idaho National Laboratory policy. Security plans are managed by the International Access Program and are stored in an International Access Program database known as IFacts. Granting access to software thus depends on information stored in the Technology Deployment database and IFacts. In the past, no connection between the Nuclear Computational Resource Center portal and these databases existed, making checking the status of license agreements and security plans time consuming and error prone. This report demonstrates that the Nuclear Computational Resource Center portal now connects to both the Technology Deployment database and IFacts, greatly improving the ease of use of the Nuclear Computational Resource Center system for administrators, which leads to a better overall experience for those requesting code access.

99 GENERAL AND MISCELLANEOUS↗

Results from a synthetic model of the ITER XRCS-Core diagnostic based on high-fidelity x-ray ray tracing

A high-fidelity synthetic diagnostic has been developed for the ITER core x-ray crystal spectrometer diagnostic based on x-ray ray tracing. This synthetic diagnostic has been used to model expected performance of the diagnostic, to aid in diagnostic design, and to develop engineering tolerances. The synthetic model is based on x-ray ray tracing using the recently developed xicsrt ray tracing code and includes a fully three-dimensional representation of the diagnostic based on the computer aided design. The modeled components are: plasma geometry and emission profiles, highly oriented pyrolytic graphite pre-reflectors, spherically bent crystals, and pixelated x-ray detectors. Plasma emission profiles have been calculated for Xe 44+ , Xe 47+ , and Xe 51+ , based on an ITER operational scenario available through the Integrated Modelling & Analysis Suite database, and modeled within the ray tracing code as a volumetric x-ray source; the shape of the plasma source is determined by equilibrium geometry and an appropriate wavelength distribution to match the expected ion temperature profile. All individual components of the x-ray optical system have been modeled with high-fidelity producing a synthetic detector image that is expected to closely match what will be seen in the final as-built system. Particular care is taken to maintain preservation of photon statistics throughout the ray tracing allowing for quantitative estimates of diagnostic performance.

47 OTHER INSTRUMENTATION↗

SHarmonic: A fast and accurate implementation of spherical harmonics for electronic-structure calculations

The authors present SHarmonic, a new implementation of the spherical harmonics targeted for electronic-structure calculations. Their approach is to use explicit formulas for the harmonics written in terms of normalized Cartesian coordinates. This approach results in a code that is as precise as other implementations while being at least one order of magnitude more computationally efficient. The library can run on graphics processing units as well, achieving an additional order of magnitude in execution speed. This new implementation is simple to use and is provided under an open-source license; it can be readily used by other codes to avoid the error-prone and cumbersome implementation of the spherical harmonics.

Mathematics and Computing↗

FLARE: field line analysis and reconstruction for 3D boundary plasma modeling

The FLARE code is a magnetic mesh generator that is integrated within a suite of tools for the analysis of the magnetic geometry in toroidal fusion devices. A magnetic mesh is constructed from field line segments and permits fast reconstruction of field lines in 3D boundary plasma codes such as EMC3-EIRENE. Both intrinsically non-axisymmetric configurations (stellarators) and those with symmetry breaking perturbations of an axisymmetric equilibrium (tokamaks) are supported. The code itself is written in Modern Fortran with MPI support for parallel computing, and it incorporates object-oriented programming for the definition of the magnetic field and the material surface geometry. Extended derived types for a number of different magnetohydrodynamic equilibrium and plasma response models are implemented. The core element of FLARE is a field line tracer with adaptive step-size control, and this is integrated into tools for the construction of Poincaré maps and invariant manifolds of X-points. A collection of high-level procedures that generate output files for visualization is build on top of that. The analysis modules are build with Python frontends that facilitate customization of tasks and/or scripting of parameter scans.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Advances in 3D transient plasma dynamics and control through MHD and hybrid fluid-kinetic simulations with JOREK

Transient phenomena and their control are of high relevance in magnetic confinement fusion plasmas to guarantee a stable and safe plasma operation. Interpretative simulations can maximize the insights gained from experiments on present machines and predictive simulations can help in the preparation of design, mitigation techniques and operational scenarios for future devices. In this article, we provide an overview of recent advances and novel scientific results obtained with the 3D non-linear hybrid fluid-kinetic code JOREK, covering physics of plasma transients from the core to the scrape-off layer (SOL) both for tokamak and stellarator devices. Substantial progress was made in the physics understanding, model validation with experiments and experiment interpretation, thus, giving confidence for predictions to devices like DTT, ITER and DEMO. The topics addressed comprise a wide range: the edge physics of new operation scenarios and edge localized mode suppression; major disruptions with a focus on runaway electrons and vertical displacement events as well as disruption mitigation by shattered pellet injection; the physics mechanisms and operational limits of the flux pumping regime for sawtooth control; MHD limits of stellarators and work towards incorporating advanced edge/SOL/exhaust dynamics; continuing improvements of the code for more efficient hybrid simulations on conventional and accelerated high performance computing architectures.

disruptions↗

Enabling Parallel Performance and Portability of Solid Mechanics Simulations Across CPU and GPU Architectures

Efficiently simulating solid mechanics is vital across various engineering applications. As constitutive models grow more complex and simulations scale up in size, harnessing the capabilities of modern computer architectures has become essential for achieving timely results. This paper presents advancements in running parallel simulations of solid mechanics on multi-core CPUs and GPUs using a single-code implementation. This portability is made possible by the C++ matrix and array (MATAR) library, which interfaces with the C++ Kokkos library, enabling the selection of fine-grained parallelism backends (e.g., CUDA, HIP, OpenMP, pthreads, etc.) at compile time. MATAR simplifies the transition from Fortran to C++ and Kokkos, making it easier to modernize legacy solid mechanics codes. We applied this approach to modernize a suite of constitutive models and to demonstrate substantial performance improvements across different computer architectures. This paper includes comparative performance studies using multi-core CPUs along with AMD and NVIDIA GPUs. Results are presented using a hypoelastic–plastic model, a crystal plasticity model, and the viscoplastic self-consistent generalized material model (VPSC-GMM). The results underscore the potential of using the MATAR library and modern computer architectures to accelerate solid mechanics simulations.

Morgan, Nathaniel (ORCID:0000000276118449)↗

Efficiently Computable Limits on EPR Pair Generation in Quantum Broadcast Channels

We investigate the generation of EPR pairs between three observers in a general causally structured setting, where communication occurs via a noisy quantum broadcast channel. The most general quantum codes for this setup take the form of tripartite quantum channels. Since the receivers are constrained by causal ordering, additional temporal relationships naturally emerge between the parties. These causal constraints enforce intrinsic no-signalling conditions on any tripartite operation, ensuring that it constitutes a physically realizable quantum code for a quantum broadcast channel. We analyze these constraints and, more broadly, characterize the most general quantum codes for communication over such channels. We examine the capabilities of codes that are fully no-signalling among the three parties, positive partial transpose (PPT)-preserving, or both, and derive simple semidefinite programs to compute the achievable entanglement fidelity. We then establish a hierarchy of semidefinite programming converse bounds -- both weak and strong -- for the capacity of quantum broadcast channels for EPR pair generation, in both one-shot and asymptotic regimes. Notably, in the special case of a point-to-point channel, our strong converse bound recovers and strengthens existing results. Finally, we demonstrate how the PPT-preserving codes we develop can be leveraged to construct PPT-preserving entanglement combing schemes, and vice versa.

FOS: Physical sciences↗

Attention to quantum complexity

The imminent era of error-corrected quantum computing demands robust methods to characterize quantum state complexity from limited, noisy measurements. We introduce the Quantum Attention Network (QuAN), a classical artificial intelligence (AI) framework leveraging attention mechanisms tailored for learning quantum complexity. Inspired by large language models, QuAN treats measurement snapshots as tokens while respecting permutation invariance. Combined with our parameter-efficient miniset self-attention block, this enables QuAN to access high-order moments of bit-string distributions and preferentially attend to less noisy snapshots. We test QuAN across three quantum simulation settings: driven hard-core Bose-Hubbard model, random quantum circuits, and toric code under coherent and incoherent noise. QuAN directly learns entanglement and state complexity growth from experimental computational basis measurements, including complexity growth in random circuits from noisy data. In regimes inaccessible to existing theory, QuAN unveils the complete phase diagram for noisy toric code data as a function of both noise types, highlighting AI’s transformative potential for assisting quantum hardware.

Kim, Hyejin [Cornell Univ., Ithaca, NY (United Sta↗

Playing Nonlocal Games across a Topological Phase Transition on a Quantum Computer

Many-body quantum games provide a natural perspective on phases of matter in quantum hardware, crisply relating the quantum correlations inherent in phases of matter to the securing of quantum advantage at a device-oriented task. In this Letter, we introduce a family of multiplayer quantum games for which topologically ordered phases of matter are a resource yielding quantum advantage. Unlike previous examples, quantum advantage persists away from the exactly solvable point and is robust to arbitrary local perturbations, irrespective of system size. We demonstrate this robustness experimentally on Quantinuum’s H1-1 quantum computer by playing the game with a continuous family of randomly deformed toric code states that can be created with constant-depth circuits leveraging midcircuit measurements and unitary feedback. We are thus able to tune through a topological phase transition—witnessed by the loss of robust quantum advantage—on currently available quantum hardware. This behavior is contrasted with an analogous family of deformed Greenberger-Horne-Zeilinger states, for which arbitrarily weak local perturbations destroy quantum advantage in the thermodynamic limit. Lastly, we discuss a topological interpretation of the game, which leads to a natural generalization involving an arbitrary number of players.

97 MATHEMATICS AND COMPUTING↗

ComPort: Rigorous Testing Methods to Safeguard Software Porting (Final UW Report)

This report summarizes the University of Washington’s contributions to the ComPort project, Rigorous Testing Methods to Safeguard Software Porting, funded under the U.S. Department of Energy, Office of Science, Office of Advanced Scientific Computing Research (ASCR), award number DE-SC0022081. UW’s work addressed numerics-driven correctness challenges arising in heterogeneous and rapidly evolving computing environments. Our efforts focused on developing interactive analysis tools, correctness-preserving rewriting systems, semantic-aware code-transformation infrastructure, and multi-language support for reproducible numerical experiments.

97 MATHEMATICS AND COMPUTING↗

MONKES: a fast neoclassical code for the evaluation of monoenergetic transport coefficients in stellarator plasmas

Abstract MONKES is a new neoclassical code for the evaluation of monoenergetic transport coefficients in stellarators. By means of a convergence study and benchmarks with other codes, it is shown that MONKES is accurate and efficient. The combination of spectral discretization in spatial and velocity coordinates with block sparsity allows MONKES to compute monoenergetic coefficients at low collisionality, in a single core, in approximately one minute. MONKES is sufficiently fast to be integrated into stellarator optimization codes for direct optimization of the bootstrap current and to be included in predictive transport suites. The code and data from this paper are available at https://github.com/JavierEscoto/MONKES/ .

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Automated Hybrid Variance Reduction on Advanced Architectures in the Shift Monte Carlo Code

Monte Carlo transport methods are the most accurate schemes for solving problems with complex energy and spatial features, but they come with a high computational cost. Although hybrid methods have enabled the use of Monte Carlo transport for a large class of problems, they still require significant computing resources. Modern multicore CPUs with large numbers of compute cores and graphical processing units (GPUs) provide opportunities to optimize the memory and run-time costs of hybrid Monte Carlo methods. This paper documents the development and analysis of three Monte Carlo transport algorithms that support hybrid transport using the consistent adjoint-driven importance sampling (CADIS) and forward-weighted CADIS methods in the Shift Monte Carlo code: history-based transport using static and dynamic threading on multicore CPUs and event-based transport enabling weight window tracking on GPUs. The results are shown for two challenging hybrid problems on the Frontier supercomputer at the Oak Ridge Leadership Computing Facility. The results show that all three methods yield good performance and enable solutions of difficult fixed-source transport problems in less than 2 min on 20 nodes of Frontier. Dynamic threading was observed to give up to 20% better scaling behavior than static threading. Moreover, the AMD Instinct 250X GPU was found to give 9 to 11 times greater throughput per graphics compute die than the best CPU performance. In conclusion, additional opportunities for optimization of hybrid transport on GPUs are discussed.

Denovo↗

High-Throughput Computing: Case Study of Medical Image Processing Applications

HPC is designed for large-scale simulations using monolithic codes of tightly coupled processes highly optimized to deliver decreased time to solution. Medical image processing is not a traditional field of HPC. Similar to AI applications, medical image processing parses large datasets, typically multiple times, to support a variety of studies for classification, diagnosis or monitoring purposes. The convergence of AI, HPC and Big Data encouraged more fields using image processing to transition to HPC. However, not all applications benefit from the same optimizations. In this paper we focus on high throughput medical image processing applications that analyze a huge dataset of small MRI images and that require HPC systems to decrease the time of parsing the entire dataset and not individual MRIs. We show in this research the performance of running SLANT, an image processing application for a whole brain segmentation, on large-scale systems and highlight performance limitations. We present optimizations prioritizing throughput that exhibit a 3.5x speed-up on the Summit Supercomputer that can be used as a baseline for building a high-throughput execution framework for other HPC systems.

Predescu, Maria↗

Improving the precision of forces in real-space pseudopotential density functional theory

The high-order finite difference real-space pseudopotential density functional theory (DFT) approach is a valuable method for large-scale, massively parallel DFT calculations. A significant challenge in the approach is the oscillating “egg-box” error introduced by aliasing associated with a coarse grid spacing. To address this issue while minimizing computational cost, we developed a finite difference interpolation (FDI) scheme [Roller et al., J. Chem. Theory Comput. 19, 3889 (2023)] as a means of exploiting the high resolution of the pseudopotential to reduce egg-box effects systematically. Here, we show an implementation of this method in the PARSEC code and examine the practical utility of the combination of FDI with additional methods for improving force precision and/or reducing its computational cost, including orbital-based forces, compensating charges (namely, adding and subtracting a judiciously chosen charge density such that the total density is unaltered), and a modified spatial domain in which the real-space grid is defined. Using selected small molecules, as well as metallic Li, as test cases, we show that a combination of all four aspects leads to a significant reduction in computational cost while retaining a high level of precision that supports accurate structures and vibrational spectra, as well as stable and accurate molecular dynamics runs.

Chemistry↗

Lessons Learned and Scalability Achieved When Porting Uintah to DOE Exascale Systems

A key challenge faced when preparing codes for Department of Energy (DOE) exascale systems was designing scalable applications for systems featuring hardware and software not yet available at leadership-class scale. With such systems now available, it is important to evaluate scalability of the resulting software solutions on these target systems. One such code designed with the exascale DOE Aurora and DOE Frontier systems in mind is the Uintah Computational Framework, an open-source asynchronous many-task (AMT) runtime system. To prepare for exascale, Uintah adopted a portable MPI+X hybrid parallelism approach using the Kokkos performance portability library (i.e., MPI+Kokkos). This paper complements recent work with additional details and an evaluation of the resulting approach on Aurora and Frontier. Results are shown for a challenging benchmark demonstrating interoperability of 3 portable codes essential to Uintah-related combustion research. These results demonstrate single-source portability across Aurora and Frontier with scaling characteristics shown to 3,072 Aurora nodes and 9,216 Frontier nodes. In addition to showing results run to new scales on new systems, this paper also discusses lessons learned through efforts preparing Uintah for exascale systems.

Holmen, John [ORNL] (ORCID:0000000259342641)↗

Zero-Field NMR and Millitesla-SLIC Spectra for >200 Molecules from Density Functional Theory and Spin Dynamics

NMR is usually performed at magnetic fields of 1 T and above to obtain sufficient sensitivity and spectral dispersion to identify chemicals based on chemical shifts and J couplings. At lower fields, the advent of hyperpolarization technologies and sensitive detectors can address sensitivity concerns. However, it remains disputed whether spectral signatures at zero and ultra-low fields are sufficient for chemical identification. Here, we report an all–electron DFT-based batch calculation of J-coupling constants, which are used to generate J coupling NMR spectra at zero field and 6.5 mT for over 200 small molecules. In the developed computational tool chain, we first used the all-electron FHI-aims code to calculate the molecular J couplings and chemical shifts. We then fed the calculated NMR parameters into the NMR simulation package SPINACH to simulate both heteronuclear J coupling spectra at zero-field, and homonuclear J coupling spectra as spin-lock induced crossing (SLIC) spectra at ultra-low field (6.5 mT). The resulting spectra demonstrate that zero and ultra-low field NMR spectra can represent unique identifiers of chemical structure for small molecules.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗