Search NASA⌕ Search

SEARCH · Search NASA

Results for “Advanced Computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24

The VTK-m User's Guide (V. 2.2)

High-performance computing relies on ever finer threading. Advances in processor technology include ever greater numbers of cores, hyperthreading, accelerators with integrated blocks of cores, and special vectorized instructions, all of which require more software parallelism to achieve peak performance. Traditional visualization solutions cannot support this extreme level of concurrency. Extreme scale systems require a new programming model and a fundamental change in how we design algorithms. To address these issues we created VTK-m: the visualization toolkit for multi-/many-core architectures. VTK-m supports a number of algorithms and the ability to design further algorithms through a top-down design with an emphasis on extreme parallelism. VTK-m also provides support for finding and building links across topologies, making it possible to perform operations that determine manifold surfaces, interpolate generated values, and find adjacencies. Although VTK-m provides a simplified high-level interface for programming, its template-based code removes the overhead of abstraction.

97 MATHEMATICS AND COMPUTING↗

The Viskores User's Guide, Release 1.1

High-performance computing relies on ever finer threading. Advances in processor technology include ever greater numbers of cores, hyperthreading, accelerators with integrated blocks of cores, and special vectorized instructions, all of which require more software parallelism to achieve peak performance. Traditional visualization solutions cannot support this extreme level of concurrency. Extreme scale systems require a new programming model and a fundamental change in how we design algorithms. To address these issues we created Viskores: the visualization toolkit for multi/many-core architectures. Viskores supports a number of algorithms and the ability to design further algorithms through a top-down design with an emphasis on extreme parallelism. Viskores also provides support for finding and building links across topologies, making it possible to perform operations that determine manifold surfaces, interpolate generated values, and find adjacencies. Although Viskores provides a simplified high-level interface for programming, its template-based code removes the overhead of abstraction.

97 MATHEMATICS AND COMPUTING↗

The Viskores User's Guide (V.1.0)

High-performance computing relies on ever finer threading. Advances in processor technology include ever greater numbers of cores, hyperthreading, accelerators with integrated blocks of cores, and special vectorized instructions, all of which require more software parallelism to achieve peak performance. Traditional visualization solutions cannot support this extreme level of concurrency. Extreme scale systems require a new programming model and a fundamental change in how we design algorithms. To address these issues we created Viskores: the visualization toolkit for multi-/many-core architectures. Viskores supports a number of algorithms and the ability to design further algorithms through a top-down design with an emphasis on extreme parallelism. Viskores also provides support for finding and building links across topologies, making it possible to perform operations that determine manifold surfaces, interpolate generated values, and find adjacencies. Although Viskores provides a simplified high-level interface for programming, its template-based code removes the overhead of abstraction.

97 MATHEMATICS AND COMPUTING↗

Integrated Molten Salt Reactor Modeling Capabilities in NEAMS Thermal Hydraulics Tools

The DOE Nuclear Energy Advanced Modeling and Simulation (NEAMS) program supports a full range of computational thermal fluids analysis capabilities and code developments for a broad range of advanced reactor concepts. The research and development approach under the thermal fluids technical area synergistically combines three length and time scales in a hierarchical multi-scale approach. To enable multi-scale thermal fluids capability using these codes, a key joint effort has been underway to develop an integrated system- and engineering-scale thermal fluids analysis capability, through integration of SAM and Pronghorn codes, both based on the MOOSE framework.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

The advanced evolution of massive stars

The nuclear rates for reactions involving 12 C and 16 O are key to computing the energy release and nucleosynthesis evolution of massive stars during their advanced burning phases. Ultimately, these burning rates shape the stellar structure and evolution and influence the nature of the compact objects produced at the end of the stellar life. We explore the implications of new nuclear reaction rates from both experimental and theoretical studies for 12 C(α, γ) 16 ​O, 12 C+ 12​ C, 12 C+ 16 ​O, and 16 O+ 16 ​O reactions for massive stars. Our goal is to investigate how the chemical structure and nucleosynthesis evolve from the He-exhaustion stage to the O-burning phase and how these processes influence the ultimate stellar fate. We computed rotating and non-rotating models for stars of different masses at solar metallicity. We used the stellar evolution code GENEC, which includes a large network of nuclear reactions and isotopes involved in advanced phases, as well as updated rates for 12 C(α,γ) 16 O. For the three fusion reactions involving 12 C and 16 O, we considered new rates following a data-driven fusion suppression scenario (hereafter HIN(RES)) and new theoretical rates obtained with time-dependent Hartree-Fock (TDHF) calculations. The updated 12 C(α, γ) 16 ​O rates mainly impact the chemical structure evolution changing the 12 C/ 16 O ratio at He-exhaustion and have little effect on the CO core mass. This variation in the 12 C/ 16 O ratio is in some cases critical for predicting the final fate of the model, which is very sensitive to 12 C abundance, and in particular the 20 M ⊙ remnant may change from a black hole to a neutron star. The He-burning (C-burning) lifetime is also decreased (increased) by about −2% (+15%). The combined new rates for 12 C+ 12 ​C and 16 O+ 16 ​O fusion reactions according to the HIN(RES) model lead to shorter C- and O-burning lifetimes by ≈ − 10%, and −50%, respectively, and shift the ignition conditions to higher temperatures and densities. In contrast, the theoretical TDHF rates primarily affect C-burning, increasing its duration by about 30% and lowering the ignition temperature. These changes modify the chemical structure of the core, the size and duration of C-burning shells, and hence their compactness. They also impact the central and shell nucleosynthesis (by ±1 dex and by factors of ±2–10, respectively), while 12 C+ 16 ​O reaction rates variations remain the least important.

abundances↗

Practical Probabilistic Programming

Recent advances in probabilistic programming languages (PPLs) have provided the capability for exact inference: computing a closed-form probability distribution for a given probabilistic program. In particular, the new language Roulette uses a language oriented programming (LOP) approach, wherein analysts build new programming languages on top of a set of primitives provided by Roulette, which then translates these structures into a weighted model counting problem which can be solved by automated reasoning tools. However, because Roulette provides few convenience features, developing these new languages is challenging even for expert users. We developed a standard library of common probability functions for Roulette with the goal of improved usability. This included approximation of continuous probability density functions using discrete probability mass functions. We demonstrated this approach by modeling a cosmic ray striking a RAM controller. We found that Roulette provides a powerful interface for highly expressive probabilistic programs to be generated. In collaboration with the NNSA Advanced Simulation and Computing program, which resulted in development of a tool called Circulette, we were able to model complex circuits expressed in Verilog using probabilistic programs with an expressivity not previously possible. Our research question that motivated the development of a Roulette standard library was to determine whether non-experts could use a PPL to model relevant problems regarding radiation effects on microelectronics. This standard library improved the expressivity of Roulette by implementing common probability density functions, mathematical operators on distributions, and support for empirical distributions. While Roulette is a powerful modeling language, the untyped, LOP approach makes error messages difficult to understand and requires expert aid. We recommend further research on Roulette, especially with its error messages, to enable improved usability. At the same time, this project demonstrated that for users familiar with Roulette and the LOP approach, Roulette provides powerful new capabilities that can be integrated with other Sandia modeling capabilities.

97 MATHEMATICS AND COMPUTING↗

On the Trotter Error in Many-body Quantum Dynamics with Coulomb Potentials

Efficient simulation of many-body quantum systems is central to advances in physics, chemistry, and quantum computing, with a key question being whether the simulation cost scales polynomially with the system size. Here, in this work, we analyze many-body quantum systems with Coulomb interactions, which are fundamental to electronic and molecular systems. We prove that Trotterization for such unbounded Hamiltonians achieves a 1/4-order convergence rate, with explicit polynomial dependence on the number of particles. The result holds for all initial wavefunctions in the domain of the Hamiltonian, and the 1/4-order convergence rate is optimal, as previous work has numerically demonstrated that it can be saturated by a specific initial ground state. The main challenges arise from the many-body structure and the singular nature of the Coulomb potential. Our proof strategy differs from prior state-of-the-art Trotter analyses, addressing both difficulties in a unified framework. Our analysis treats the Coulomb potential as an unbounded operator without modification or regularization, and does not rely on spatial discretization, making it compatible with both first- and second-quantized circuit constructions.

Fang, Di [Duke Univ., Durham, NC (United States)]↗

LandScan HD: a high-resolution gridded ambient population methodology for the world

Unwarned population distributions accounting for routine human activities are needed to address many global human security challenges, including disasters, conflict, and infrastructure demand. LandScan High Definition (LSHD) supports this need through gridded ambient population estimates that measure average human presence between daytime and nighttime at a high spatial resolution of 3 arcseconds (approximately 90 m). Although LSHD has traditionally been produced on a country-specific basis, advances in global foundational data and computational resources now enable scaling its methodology to the world. Combining aspects of top-down and bottom-up gridded population methods, LSHD allocates subnational population totals from authoritative statistics to built-up areas based on occupancy estimates for multiple facility types (e.g., residential, commercial) and then reaggregates these estimates to a global population grid. We scale this approach by organizing the LSHD data stack into a 1° resolution tileset of vector analytic features, enabling an efficient and repeatable workflow for all countries worldwide. Examining the Philippines as an output of the global LSHD baseline dataset, we contrast unwarned and residential (WorldPop) population distributions by (1) exploring a practical application of flood risk assessment and (2) evaluating their congruence with outcomes of collective human activities (subnational CO 2 emissions). Finally, we discuss plans to address current LSHD limitations through data/modeling and uncertainty quantification improvements and provide outlook for workflow automation and extending the model to social, demographic and economic population characteristics.

Building morphology↗

Out-of-distribution detection with non-parametric density estimation for models predicting processing history of uranium ore concentrates

The rapid advancement in machine learning (ML) and computer vision (CV) coincides with the growth of interest in deploying these ML/CV models in numerous fields from medicine to social science. Similar to those areas, we have witnessed a great number of works in materials science employing ML/CV models – neural networks in particular – in their studies in recent years. These models have proven to obtain accurate performance in various tasks. However, these models struggle to attain a similar performance when encountering test samples coming from a distribution that is different from the training set. More importantly, they fail without providing any warning to the users. Therefore, we propose a framework for detecting out-of-distribution (OOD) samples to alert users when a human intervention might be necessary in this work. Specifically, we explore the use of a non-parametric density estimation method to detect OOD samples. Here, we assess OOD detection capability of the proposed framework on ML models developed for categorizing precipitation routes of U 3 O 8 when encountering OOD datasets that contain samples (1) undergone different imaging acquisition process, (2) undergone different material synthesis process, and (3) different materials than ID set. Through those experiments, we achieve an average area under the receiver operating characteristic (AUROC) of at least 91% on average in detecting OOD samples. With minimal overhead cost and superior performance, the proposed framework enables a reliable and safe system when deploying in real-world scenarios.

Convolutional neural networks↗

Curvature Induced Modifications of Chirality and Magnetic Configuration in Perpendicular Films

Designing curvature in three-dimensional (3D) magnetic nanostructures enables controlled manipulation of local energy landscapes, allowing for the modification of noncollinear spin textures relevant for next-generation spintronic devices. In this study, we experimentally investigate 3D magnetization textures in a Co/Pd multilayer film, exhibiting strong perpendicular magnetic anisotropy (PMA), deposited onto curved Cu nanowire meshes with diameters as small as 50 nm and lengths of several microns. Utilizing magnetic soft X-ray nanotomography, we achieve reconstructions of 3D magnetic domain patterns at approximately 30 nm spatial resolution. This approach provides detailed information on both the orientation and magnitude of magnetization within the film. Our results reveal that interfacial anisotropy in the Co/Pd multilayers drives the magnetization toward the local surface normal. In contrast to typical labyrinth domains observed in planar films, the presence of curved nanowires significantly alters the domain structure, with domains preferentially aligning along the nanowire axis in close proximity, while adopting random orientations farther away. We report direct experimental observation of a curvature-induced Dzyaloshinskii-Moriya interaction (DMI), which is quantified to be approximately one-third of the intrinsic DMI in Co/Pd stacks. The curvature induced DMI enhances stability of Néel-type domain walls. These experimental observations are further supported by micromagnetic simulations. Altogether, our findings demonstrate that introducing curvature into magnetic nanostructures provides a powerful strategy for tailoring complex magnetic behaviors, paving the way for the design of advanced 3D racetrack memory and neuromorphic computing devices.

Raftrey, David↗

Nanophotonic waveguide chip-to-world beam scanning

A seamless chip-to-world photonic interface enables broad advancements in optical ranging, display, communication, computation and quantum information science. The ideal solution enables two-dimensional scanning of a diffraction-limited beam from anywhere on a photonic integrated circuit to a large number of resolvable spots. Current beam-scanning technologies are limited by a fundamental trade-off: photonic-integrated-circuits with diffractive optics offer scalability but have poor mode quality, whereas inertially limited micromechanical scanners provide high-quality beams but lack scalable integration. Here we report a photonic ski-jump—a nanoscale waveguide monolithically integrated on a piezoelectric cantilever—to overcome these limitations. It passively curls ~90° out-of-plane within a less-than-0.1 mm 2 footprint, emits a submicrometre, broadband diffraction-limited beam, and exhibits kilohertz-rate mechanical resonances with quality factors of over 10,000. Fabricated in a volume complementary metal–oxide–semiconductor (CMOS) foundry, our device enables scalable two-dimensional beam scanning. Driven on-resonance at CMOS-level voltages, it achieves a footprint-adjusted spot rate of 68.6 mega spots s –1 mm–², exceeding state-of-the-art micro-electro-mechanical systems mirrors by more than 50-fold, which is sufficient for one million pixels at 100 Hz from an approximately 1.5 mm diameter footprint. We demonstrate full-colour image and video projection, and single-photon initialization and readout from silicon vacancy centres in diamond. Finally, by demonstrating uniformity across a 64 ski-jump array, we establish a pathway to achieving greater than one gigaspot resolution at kilohertz rates within a sub-5-cm-diameter footprint, creating a seamless optical pipeline between integrated photonic processors and the free-space world.

Displays↗

Additive manufacturing of LiCoO 2 electrodes via vat photopolymerization for lithium ion batteries

Additive manufacturing has the potential to revolutionize the fabrication of lithium-ion batteries for a diversity of applications including in portable, biomedical, aerospace, and the transportation fields. Standard commercial batteries consist of stacked layers of various components (current collectors, cathode, anode, separator and electrolyte) in a two-dimensional manner. By leveraging the latest advances in additive manufacturing and computer-aided design, an improved geometric and electrochemical configuration of these batteries can maximize energy efficiency while allowing design optimization to reduce dead space for a given application. In this work, a composite UV photosensitive resin was prepared and used as feedstock in a vat photopolymerization system. The resin was loaded with LiCoO 2 acting as electrochemically active material for the cathode of a lithium-ion battery, and was further improved with the addition of conductivity-enhancing carbonaceous additives. Challenges to additive manufacturing arise from the opacity and high viscosity of the composite nature of these electrochemically-active resins, which cause light refraction during selective UV curing. Subsequently, items were printed and subjected to a thermal post-processing step to obtain an adequate compromise between electrochemical performance and mechanical integrity. Both sintered and green state 3D printed cathodes were assembled into half-cell lithium-ion batteries using lithium metal as a reference and counter electrode. Electrochemical cycling of these batteries yielded satisfactory results approaching commercial LiCoO 2 cathodes’ performance, with the potential advantages of additive manufacturing – high surface area anode–cathode configurations for power performance as well as shape conformability.

25 ENERGY STORAGE↗

Systematic Construction of Time-Dependent Hamiltonians for Microwave-Driven Josephson Circuits

Time-dependent electromagnetic drives are fundamental for controlling complex quantum systems, including superconducting Josephson circuits. In these devices, accurate time-dependent Hamiltonian models are imperative for predicting their dynamics and designing high-fidelity quantum operations. Existing numerical methods, such as black-box quantization (BBQ) and energy-participation ratio (EPR), excel at modeling the static Hamiltonians of Josephson circuits. However, these techniques do not fully capture the behavior of driven circuits stimulated by external microwave drives, nor do they include a generalized approach to account for the inevitable noise and dissipation that enter through microwave ports. Here, we introduce numerical techniques that leverage classical microwave simulations, efficiently executable in finite-element solvers, to obtain the time-dependent Hamiltonian of microwave-driven superconducting circuits with arbitrary geometries under charge, flux, or mixed electromagnetic modulation. Importantly, our techniques do not rely on a lumped-element description of the superconducting circuit, in contrast to previous approaches to tackling this problem. We demonstrate the versatility of our approach by characterizing the driven properties of realistic circuit devices in complex electromagnetic environments, including coherent dynamics due to charge and flux modulation, as well as drive-induced relaxation and dephasing. Our techniques offer a powerful toolbox for optimizing circuit designs and advancing practical applications in superconducting quantum computing.

Lu, Yao [Yale U.; Yale U. (main); Fermilab] (ORCID↗

Direct pulse-level compilation of arbitrary quantum logic gates on superconducting qutrits

Advanced simulations and calculations on quantum computers require high-fidelity implementations of quantum operations. The universal gateset approach builds complex unitaries from a small set of primitive gates, often resulting in a long gate sequence, which is typically a leading factor in the total accumulated error. Compiling a complex unitary for processors with higher-dimensional logical elements, such as qutrits, exacerbates the accumulated error per unitary, since an even longer gate sequence is required. Optimal control methods promise time- and resource-efficient compact gate sequences and, therefore, higher fidelity. These methods generate pulses that can directly implement any complex unitary on a quantum device. In this work, we demonstrate that any arbitrary qubit and qutrit gate can be realized with high fidelity, which can significantly reduce the length of a gate sequence. We generate and test pulses for a large set of randomly selected arbitrary unitaries on several quantum processing units (QPUs): the Lawrence Livermore National Laboratory Quantum Device and Integration Testbed’s (QuDIT’s) standard QPU and three of Rigetti’s QPUs: Ankaa-2, Ankaa-9Q-1, and Aspen-M-3. On the QuDIT platform’s standard QPU, the average fidelity of random qutrit gates is 97.9 ± 0.5% measured with conventional QPT and 98.8 ± 0.6% from QPT with gate folding. Rigetti’s Ankaa-2 achieves random qubit gates with an average fidelity of 98.4 ± 0.5% (conventional QPT) and 99.7 ± 0.1% (QPT with gate folding). On Ankaa-9Q-1 and Aspen-M-3, the average fidelities with conventional qubit QPT measurements were higher than 99% (see Appendix). Here we show that optimal control gates are robust to drift for at least 3 h and that the same calibration parameters can be used for all implemented gates. Our work promises that the calibration overheads for optimal control gates can be made small enough to enable efficient quantum circuits based on this technique.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Fault-tolerant optical interconnects for neutral-atom arrays

We analyze the use of photonic links to enable large-scale fault-tolerant connectivity of locally error-corrected modules based on neutral atom arrays. Our approach makes use of recent theoretical results showing the robustness of surface codes to boundary noise and combines recent experimental advances in atom-array quantum computing with logical qubits with optical quantum networking techniques. We find the conditions for fault tolerance can be achieved with local two-qubit Rydberg gate and nonlocal Bell-pair errors below 1% and 10%, respectively, without requiring distillation or space-time overheads. Realizing the interconnects with a lens, a single optical cavity, or an array of cavities enables—with sufficient multiplexing—a Bell-pair generation rate in the 1–50 MHz range. When directly interfacing logical qubits, this rate translates to error-correction cycles in the 25–2000 kHz range, satisfying all requirements for fault tolerance and in the upper range fast enough for 100 kHz logical clock cycles. Published by the American Physical Society 2025

Sinclair, Josiah (ORCID:0000000215238295)↗

Reinforcement Learning-Based Approach for EMT Automation of Large-Scale PV Plants

In the pursuit of efficient and precise modeling of large-scale power systems, particularly utility-scale photovoltaic (PV) plants, Electromagnetic Transient (EMT) simulations play a crucial role. As utility-scale PV plants increase in size and complexity, traditional computational methods become inadequate, necessitating more advanced techniques. This paper highlights the progressive efforts made to accelerate EMT simulations. A novel continuous reinforcement learning (RL) strategy is explored to automate the differentiation and categorization of stiff and non-stiff differential algebraic equations (DAEs). The use of stiff and non-stiff integration methods applied to relevant parts of the DAEs assists with the speed-up of the simulations. The paper details the data acquisition, development and offline training of the RL model, leading to its validation that demonstrates a high precision in optimizing simulation methods. The proposed RL promises to significantly enhance the efficacy of EMT simulations, offering a robust framework for the future of power system analysis.

Xia, Qianxue↗

Shifting Between Compute and Memory Bounds: A Compression-Enabled Roofline Model

In the evolving landscape of high-performance computing, especially to fight the end of Moore’s Law and Dennard’s Scaling, the ability to shift between compute-bound and memory-bound states is critical for enhancing adaptability and flexibility to diverse system and domain-specific architectures. Such capability is vital for optimizing performance across distinguished hardware configurations, such as accelerators, memory hierarchies, and cache systems. Despite that ad hoc optimization techniques, such as compressed/approximate computation, have been enabled for compute-/data-intensive computing for improved performance in distinct hardware settings, there lacks an understanding of 1) the rational behind performance improvement; 2) capability of different optimizations; 3) what optimization to respond to specific computational and memory demands. This work proposes a compression-enabled roofline model to facilitate this adaptability with data compression techniques to balance and transform between computational and memory demands. This model enables applications to adjust in response to the specific strengths and limitations of the underlying hardware and system to optimize resource utilization. The effectiveness of this approach is demonstrated with matrix multiplication kernels on different input sizes, with turning on/off various compression techniques, including 1) low-precision floating point; 2) sparse matrix formulation; and 3) compressed arrays with ZFP. By reducing memory transfer volumes and cache misses and increasing data locality and computational intensity through compression, the specific roofline model can transform between compute and memory bounds to align more efficiently with system capabilities. This advancement not only improves overall performance but also maximizes adaptability in diverse computing environments.

Naraparaju, Ramasoumya [University of Washington]↗

Fiats: Functional inference and training for surrogates

Fiats provides a platform for research on the training and deployment of neural-network surrogate models for computational science. Fiats also supports exploring, advancing, and combining functional, object-oriented, and parallel programming patterns in Fortran 2023. As such, the Fiats name has dual expansions: “Functional Inference And Training for Surrogates” or “Fortran Inference And Training for Science.” Fiats inference and training procedures are pure and therefore satisfy a language constraint imposed on procedure invocations inside Fortran’s parallel loop construct: do concurrent. Furthermore, the Fiats training procedures are built around a do concurrent parallel reduction. Several compilers can automatically parallelize do concurrent on Central Processing Units (CPUs) or Graphics Processing Units (GPUs). Fiats thus aims to achieve performance portability through standard language mechanisms.

Rouson, Damian [Lawrence Berkeley National Laborat↗