Search NASA⌕ Search

SEARCH · Search NASA

Results for “Implementation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25

Adiabatic quantum imaginary time evolution

We introduce an adiabatic state preparation protocol which implements quantum imaginary time evolution under the Hamiltonian of the system. Unlike the original quantum imaginary time evolution algorithm, adiabatic quantum imaginary time evolution does not require quantum state tomography during its runtime and, unlike standard adiabatic state preparation, the final Hamiltonian is not the system Hamiltonian. Instead, the algorithm obtains the adiabatic Hamiltonian by integrating a classical differential equation that ensures that one follows the imaginary time evolution state trajectory. We introduce some heuristics that allow this protocol to be implemented on quantum architectures with limited resources. We explore the performance of this algorithm via classical simulations in a one-dimensional spin model and highlight essential features that determine its cost, performance, and implementability for longer times, and compare to the original quantum imaginary time evolution for ground-state preparation. More generally, our algorithm expands the range of states accessible to adiabatic state preparation methods beyond those that are expressed as ground states of simple explicit Hamiltonians. Published by the American Physical Society 2024

Hejazi, Kasra (ORCID:000000032349478X)↗

Generalized Quantum Signal Processing

Quantum signal processing (QSP) and quantum singular value transformation (QSVT) currently stand as the most efficient techniques for implementing functions of block-encoded matrices, a central task that lies at the heart of most prominent quantum algorithms. However, current QSP approaches face several challenges, such as the restrictions imposed on the family of achievable polynomials and the difficulty of calculating the required phase angles for specific transformations. In this paper, we present a generalized quantum signal processing (GQSP) approach, employing general SU(2) rotations as our signal-processing operators, rather than relying solely on rotations in a single basis. Our approach lifts all practical restrictions on the family of achievable transformations, with the sole remaining condition being that | P | ≤ 1 , a restriction necessary due to the unitary nature of quantum computation. Furthermore, GQSP provides a straightforward recursive formula for determining the rotation angles needed to construct the polynomials in cases where P and Q are known. In cases where only P is known, we provide an efficient optimization algorithm capable of identifying in under a minute of GPU time, a corresponding Q for polynomials of degree on the order of 10 7 . We further illustrate GQSP simplifies QSP-based strategies for Hamiltonian simulation, offer an optimal solution to the ϵ -approximate fractional query problem that requires O ( ( 1 / δ ) + log ( 1 / ϵ ) ) queries to perform where O ( 1 / δ ) is a proved lower bound, and introduces novel approaches for implementing bosonic operators. Moreover, we propose a novel framework for the implementation of normal matrices, demonstrating its applicability through synthesis of diagonal matrices, as well as the development of a new algorithm for convolution through synthesis of circulant matrices using only O ( d log N + log 2 N ) 1 and 2-qubit gates for a filter of lengths d . Published by the American Physical Society 2024

Motlagh, Danial↗

Fast Sideband Control of a Multimode Cavity Memory with Weak Dispersive Coupling to a Transmon

Mitigating ancilla-mediated error channels is a critical challenge in controlling high-quality superconducting cavities using circuit quantum electrodynamics (cQED). We address this by weakening the dispersive coupling while demonstrating fast, high-fidelity multimode control through transmon-mediated sideband interactions. We implement transmon-cavity SWAP gates with speeds up to 30 times larger than the bare dispersive coupling. Combined with transmon rotations, this enables universal state preparation in a single mode, though achieving unitary gates and extending control to multiple modes remains a challenge. In this work, we overcome this limitation by introducing two control strategies: (i) a shelving technique that stores populations in sideband-transparent states, and (ii) a method that exploits the dispersive shift to implement photon-number-selective transmon-cavity SWAP gates. We use these protocols to prepare Fock and binomial code states across any of the ten modes of a multimode cavity with millisecond coherence times—serving as a multimode quantum memory. We demonstrate unitaries that encode and decode an arbitrary qubit state from the transmon into corresponding vacuum and Fock state superpositions, as well as entangled NOON states of cavity mode pairs—a scheme extendable to arbitrary multimode Fock encodings in the cavity modes. Furthermore, we implement a new binomial encoding gate that converts arbitrary transmon superpositions into binomial code states in any cavity mode at a rate exceeding the dispersive shifts in our system, achieving an average post-selected state fidelity of 96.3% in a 4 μ⁢s gate time. By using precalibrated transmon and sideband pulses, our work demonstrates multimode control with significantly reduced calibration overhead, enabling efficient unitary operations using sideband interactions in multimode cQED systems.

Huang, Jordan [Rutgers Univ., Piscataway, NJ (Unit↗

A general Bayesian algorithm for the autonomous alignment of beamlines

Autonomous methods to align beamlines can decrease the amount of time spent on diagnostics, and also uncover better global optima leading to better beam quality. The alignment of these beamlines is a high-dimensional expensive-to-sample optimization problem involving the simultaneous treatment of many optical elements with correlated and nonlinear dynamics. Bayesian optimization is a strategy of efficient global optimization that has proved successful in similar regimes in a wide variety of beamline alignment applications, though it has typically been implemented for particular beamlines and optimization tasks. In this paper, we present a basic formulation of Bayesian inference and Gaussian process models as they relate to multi-objective Bayesian optimization, as well as the practical challenges presented by beamline alignment. We show that the same general implementation of Bayesian optimization with special consideration for beamline alignment can quickly learn the dynamics of particular beamlines in an online fashion through hyperparameter fitting with no prior information. We present the implementation of a concise software framework for beamline alignment and test it on four different optimization problems for experiments on X-ray beamlines at the National Synchrotron Light Source II and the Advanced Light Source, and an electron beam at the Accelerator Test Facility, along with benchmarking on a simulated digital twin. We discuss new applications of the framework, and the potential for a unified approach to beamline alignment at synchrotron facilities.

47 OTHER INSTRUMENTATION↗

LibraryX: A Framework for Cross-Library-Call Optimization

Scientific applications utilize performance libraries as a software engineering concept: these libraries encapsulate important and well-understood (mathematical) operations, allow for reuse, and are implemented and tuned by experts. Domain scientists then implement complex algorithms based on these domainspecific libraries. While individual library calls are optimized, larger performance gains across sequences of calls—sometimes spanning multiple libraries—are often unrealized, forcing a trade-off between performance and implementation complexity.To overcome this issue, we propose LibraryX, an approach and a system that allows for cross-library-call optimization even when library calls stem from multiple performance libraries. LibraryX annotates library calls with semantic information and optimizes entire directed acyclic graphs (DAGs) of calls dynamically using the SPIRAL code generation system. We demonstrate its effectiveness across a range of memory bound workloads, achieving significant speedups on Nvidia, AMD, and Intel accelerators compared to code using native libraries without cross-call optimization.

Rao, Sanil [Carnegie Mellon University,Department ↗

Distributed Multi-GPU Community Detection on Exascale Computing Platforms

Community detection is a fundamental operation in graph mining, and by uncovering hidden structures and patterns within complex systems it helps solve fundamental problems pertaining to social networks, such as information diffusion, epidemics, and recommender systems. Scaling graph algorithms for massive networks becomes challenging on modern distributed-memory multi-GPU (Graphics Processing Unit) systems due to limitations such as irregular memory access patterns, load imbalances, higher communication-computation ratios, and cross-platform support. We present a novel algorithm HiPDPL-GPU (distributed parallel Louvain) to address these challenges. We conduct experiments involving different partitioning techniques to achieve optimized performance of HiPDPL-GPU on the two largest supercomputers: Frontier and Summit. Remarkably, HiPDPL-GPU processes a graph with 4.2 billion edges in less than 3 minutes using 1024 GPUs. Qualitatively performance of HiPDPL-GPU is similar or better compared to other state-of-the-art CPU- and GPU-based implementations. While prior GPU implementations have predominantly employed CUDA, our first-of-its-kind implementation for community detection is cross-platform, accommodating both AMD and NVIDIA GPUs.

graph algorithms, high performance comptuing↗

Cryogenic-Refined MOSFET Modeling for Oscillator, Frequency Divider, and Amplifier Designs Below 4 K

Capturing device characteristic changes at cryogenic temperatures is crucial for cryo-CMOS circuit designs. In this work, we present an isothermal cryogenic-refined modeling approach for CMOS transistors that is simple, low overhead, and easy to implement while offering the required accuracy for predicting circuit performance at the designated temperatures. Guided by die-level measurement data and circuit design principles, the model introduces corrections to only five critical parameters: threshold voltage, carrier mobility, elevated low-frequency flicker noise, dominant high-frequency shot noise, and subthreshold swing (SS). These refinements are implemented around the foundry-provided SPICE model, which is typically validated only down to about 200 K. With these adjustments, the proposed cryogenic-refined model achieves less than 5% error in both large-signal metrics (I–V characteristics) and small-signal parameters (e.g., transconductance) when compared with device measurements at deep-cryogenic temperatures. The methodology is validated in two advanced technologies: TSMC 40-nm CMOS and GlobalFoundries (GF) 45-nm RF-SOI. We further demonstrate its applicability in three representative RF circuits: a 30-GHz LC oscillator, a high-speed current-mode-logic (CML) frequency divider (FD), and a subthreshold Gb/s amplifier, all showing close agreement between simulated predictions and measurements performed at 4 and 2.5 K. Finally, we believe that the proposed approach is implementation-friendly and can significantly accelerate the development of cryo-CMOS integrated circuits.

circuit modeling↗

Development and Experimental Validation of a High-Power DC Distribution Testbed for Advanced Charging Infrastructure and Energy Management

This paper presents the development of a hardware testbed for DC-distributed high-power charging (HPC) stations. As DC distributed solutions emerge as a viable solution to optimize HPC site operations, challenges such as interoperability, protection, and seamless integration of distributed energy resources (DER) persist. These issues underscore the need for a robust testing facility to investigate compliance of available commercial off-the-shelf (COTS) market devices. The developed testbed features a DC-distributed charging hub including a charger, emulated energy storage system (ESS), and site level communication and controller implementation. It facilitates the testing of COTS hardware, charger prototypes, standards validation and site energy management system (SEMS) controllers at rated power. This paper details the development of the charging infrastructure platform, implementation of communication system, validation of different SEMS algorithms, and understanding improvements required for future expansion. Using the developed testbed, interoperability gaps for SEMS implementation with multi-vehicle concurrent charging via a multi-port charger are experimentally observed. Aimed at supporting the transition to large-scale EV charging infrastructure deployment and DER integration, this testbed plays a crucial role in conformity testing of COTS device interoperability.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Uncertainty propagation and sensitivity analysis for constrained optimization of nuclear waste vitrification

Abstract The vitrification of high‐level waste (HLW) by heating a mixture of glass‐forming chemicals (GFCs) with the waste can be improved using a constrained optimization problem. This study explores how different uncertainty propagation (UP) methods implemented with the optimization process can affect the glass formulation of nuclear waste glasses. UP is the effort of propagating uncertain inputs through a system to understand and quantify output distributions. Uncertainty intervals are crafted from output distributions to inform the optimization algorithm. UP is often implemented with Monte Carlo (MC) sampling for large nonlinear systems, which can be difficult to implement within a constrained optimization algorithm that requires derivative information. Other UP methods often used for optimization under uncertainty (OUU) can be designed to work within an established constrained optimization framework. Methods of UP are evaluated in this study including iterative sampling approaches, first‐order approximations, and surrogate modeling with machine learning (ML). A method of dimensional reduction based on global sensitivity analysis is introduced to support the UP methods for the large dimensionality of the problem. Analytical UP methods able to achieve similar optimums 10 times faster than the baseline MC approach, and produce 93.9% similar output distributions are reported.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

A Techno-Economic Analysis of a 50MWth Light-Trapping Cavity-Planar Solar Receiver Tower Capital Expenditures and its Cost Mitigation Strategies

To maximize thermal efficiency, the National Renewable Energy Laboratory (NREL) has proposed a light-trapping cavity-planar receiver design intended to capture energy from reradiating surfaces. This system implements a macroscale light trapping mechanism induced by panels with triangular channels, this mechanism allows for elevated temperatures on the receiver panels and in turn the temperature of the HTF; using this design coupled with the implementation of a fluidized particle flow as the HTF, it can be expected for some components to reach a peak working temperature of nearly 1000 degrees C cyclically throughout each day-night cycle. While these temperatures correlate to higher efficiency of the CSP tower they also demand intense thermomechanical properties from the materials used to make the receiver panels. In this analysis, we will list our assumptions to provide clarity on the significance of our calculation. This cost analysis will be conducted using a combination of both case study data and surveying industry to determine costs that are relevant to the current market trends. This analysis represents an early attempt to establish the Capital Expenditures required for a CSP tower of such design to determine the feasibility of implementing such a system in the industry.

concentrated solar power (CSP)↗

UNDERSTANDING THE SEMI-PROBABILISTIC APPROACHES IN STRUCTURAL RELIABILITY USED TO SET DESIGN RELIABILITY TARGETS FOR GRAPHITE COMPONENTS USING ASME BPVC METHODS

Graphite is a quasi-brittle material, resulting in random variability in tensile strength distributions. To account for the random variability in strength, HHA-3000 of the ASME BPVC provides two semi-probabilistic methods for qualifying nuclear graphite components in the design stage, the simplified and full assessments. The full and simplified assessments apply statistical methods to engineering-based design problems. This is often referred to as reliability-based design. Reliability-based design (RBD) is a method to develop reliable designs by accounting for uncertainties and result in small chances of failure when also considering safety factors. RBDs provide reliability targets using semi-probabilistic approaches. RBD is implemented in ASME BPVC HHA-3000 for nuclear graphite components, but is not specific to that application. There has been much confusion around the methods implemented in ASME BPVC HHA-3000 for qualifying nuclear graphite components. To address the confusion, this paper takes a hierarchical approach. First, the general RBD framework is presented. Then, the semi-probabilistic methods and the underlying assumptions implemented in the assessments are presented. The semi-probabilistic methods are separated from the engineering modifications that have been made to the assessments. After building the framework and underlying assumptions, the specific methods in the full and simplified assessments are explained in three steps: inputs, methods, outputs. The methods are applied to an H-451 reflector block. Tensile strength properties for other graphite grades are provided.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Fast and Accurate Intersections on a Sphere

We introduce a fast, high-precision algorithm for calculating intersections between great circle arcs and lines of constant latitude on the unit sphere. We first propose a simplified intersection point formula with improved speed and numerical robustness over the ones traditionally implemented in geoscience software. We then show how algorithms based on the concept of error-free transformations (EFT) can be applied to evaluate this formula within a relative error bound that is on the order of machine precision. Here, we demonstrate that, with a vectorized and parallelized implementation, this enhanced accuracy is achieved with no compute time overhead compared to a direct calculation in hardware floating point, making our algorithm suitable for performance-sensitive applications like regridding of high-resolution climate data. In contrast, evaluating our formula using high-precision data types like quadruple precision and arbitrary precision, or using the robust intersection computation routines from the Computational Geometry Algorithms Library, leads to significant computational overhead, especially since these alternatives inhibit vectorization. More generally, our work demonstrates how EFT techniques can be combined and extended to implement nontrivial geometric calculations with high accuracy and speed.

Environmental sciences↗

Future Circular Collider Feasibility Study Report

Volume 3 of the FCC Feasibility Report presents studies related to civil engineering, the development of a project implementation scenario, and environmental and sustainability aspects. The report details the iterative improvements made to the civil engineering concepts since 2018, taking into account subsurface conditions, accelerator and experiment requirements, and territorial considerations. It outlines a technically feasible and economically viable civil engineering configuration that serves as the baseline for detailed subsurface investigations, construction design, cost estimation, and project implementation planning. Additionally, the report highlights ongoing subsurface investigations in key areas to support the development of an improved 3D subsurface model of the region. The report describes the development of the project scenario based on the ‘avoid-reduce-compensate’ iterative optimisation approach. The reference scenario balances optimal physics performance with territorial compatibility, implementation risks, and costs. Environmental field investigations covering almost 600 hectares of terrain—including numerous urban, economic, social, and technical aspects—confirmed the project’s technical feasibility and contributed to the preparation of essential input documents for the formal project authorisation phase. The summary also highlights the initiation of public dialogue as part of the authorisation process. The results of a comprehensive socio-economic impact assessment, which included significant environmental effects, are presented. Even under the most conservative and stringent conditions, a positive benefit-cost ratio for the FCC-ee is obtained. Finally, the report provides a summary of the studies conducted to document the current state of the environment.

Benedikt, M. [European Organization for Nuclear Re↗

SmartFuse: Reconfigurable Smart Switches to Accelerate Fused Collectives in HPC Applications

Communication switches have sometimes been augmented to process collectives (e.g., the IBM BlueGene project and the Mellanox SHArP switch). In this work, we find that there is a great acceleration opportunity through the further augmentation of switches to accelerate more complex functions that combine communication with computation. We consider three types of such functions. The first is fully-fused collectives built by fusing multiple existing collectives like Allreduce with Alltoall. The second is semi-fused collectives built by combining a collective with another computation. The third we refer to as higher-order collectives built by combining multiple computations and communications, such as to perform matrix-matrix multiply (PGEMM). In this work, we propose a framework called SmartFuse to accelerate fused collective functions. The core of SmartFuse is a reconfigurable smart switch to support these operations. The semi/fully fused collectives are implemented with a CGRAlike architecture, while higher-order collectives are implemented with a more specialized computational unit that can also schedule communication. Supporting our framework is software to evaluate and translate relevant parts of the input program, compile them into a control data flow graph, and then map this graph to the switch hardware. The proposed framework, once deployed, has the strong potential to accelerate existing HPC applications transparently by encapsulation within an MPI implementation. Experimental results show that this approach improves the performance of the PGEMM kernel, MINIFE, and AMG by, on average, 94%, 15%, and 13%, respectively.

Haghi, Pouya↗

FuseIM: Fusing Probabilistic Traversals for Influence Maximization on Exascale Systems

Probabilistic breadth-first traversals (BPTs) are used in many network science and graph machine learning applications. In this paper, we are motivated by the application of BPTs in stochastic diffusion-based graph problems such as influence maximization. These applications heavily rely on BPTs to implement a Monte-Carlo sampling step for their approximations. Given the large sampling complexity, stochasticity of the diffusion process, and the inherent irregularity in real-world graph topologies, efficiently parallelizing these BPTs remains significantly challenging. In this paper, we present a new algorithm to fuse massive number of concurrently executing BPTs with random starts on the input graph. Our algorithm is designed to fuse BPTs by combining separate traversals into a unified frontier on distributed multi-GPU systems. To show the general applicability of the fused BPT technique, we have incorporated it into two state-of-the-art influence maximization parallel implementations (gIM and Ripples). Our experiments on up to 4K nodes of the OLCF Frontier supercomputer (32,768 GPUs and 196K CPU cores) show strong scaling behavior, and that fused BPTs can improve the performance of these implementations up to 34x (for gIM) and ~360x (for Ripples).

Neff, Reece W.↗

Ontologies at Work: Analyzing Information Requirements for Model Predictive Control in Buildings

Model Predictive Control (MPC) has shown significant potential for improving energy efficiency, indoor air quality and occupant comfort of buildings. MPC-based control algorithms have also shown the ability to shift loads and optimize for multiple objectives, including but not limited to reducing the green-house gas emissions, energy costs and peak demand. However, one of the main implementation challenges of these control algorithms is the integration and configuration effort needed to deploy a supervisory MPC controller in a building. By assigning standardized references to information sources and control points in buildings, existing studies have shown that semantic ontologies and corresponding queries have the potential to ease the deployment of such controllers. Yet, the use of semantic information to ease the deployment processes of MPC controllers is still limited. In this paper, we review three MPC experiments and synthesize the information requirements of these optimization problems. We then turn to existing and upcoming semantic ontologies such as Brick, SAREF and ASHRAE Standard 223 to represent these requirements, evaluating their potential to support the implementation of an MPC controller. This investigation concludes with a discussion of existing opportunities and open questions that the community should explore to support more streamlined MPC implementations.

Prakash, Anand Krishnan↗

Large-Message All-to-All Communication at Frontier Scale

Near the full scale of exascale supercomputers, latency can dominate the cost of all-to-all communication even for very large message sizes. We describe GPU-aware all-to-all implementations designed to reduce latency for large message sizes at extreme scales, and we present their performance using 65536 tasks (8192 nodes) on the Frontier supercomputer at the Oak Ridge Leadership Computing Facility. Two implementations perform best for different ranges of message size, and all outperform the vendor-provided MPI_Alltoall. Our results show promising options for improving implementations of MPI_Alltoall_init.

White, Trey [ORNL] (ORCID:000900052186075X)↗

AI-Batt (Autonomous Identification of Battery Life Models) [SWR 21-36]

Autonomous Identification of Battery Life Models (AI-Batt) AI-Batt is a MATLAB code base for developing lifetime models for batteries from accelerated aging data. The code base provides many functions for processing, visualizing, and modeling battery aging data, making the data processing, exploration, and modeling workflow substantially faster. These tools are tailored for working with battery aging data sets, which usually consist of many separate time-series for each cell, with many test conditions and possible replicates at each condition, which makes it difficult to simply process or visualize the data set. Complex modeling tasks, such as cross-validation, sensitivity analysis, and uncertainty quantification have been implemented to enable thorough statistical investigation of model predictions. Additionally, several machine-learning algorithms are implemented to autonomously identify suitable models via symbolic regression. Data processing functions automatically cast data from the struct data type, which is commonly used to store experimental data, but is not an acceptable input for most algorithms, to the table data type, which can be easily used as input to any optimization algorithm. Also, the data can be separated into time-invariant and time-variant data tables, which is helpful for exploring the data set as well as developing separate models for time-variant and time-invariant aging mechanisms. For example, in aging tests with constant temperature, temperature is a time-invariant experimental condition. Visualization tools enable plotting of data, model fits, and model simulations possible with single-line function calls, empowering data exploration of complex data sets with both time-varying and time-invariant trends. Plots can be automatically generated for the whole data set, or separated by data group (groups of test replicates) or individual data series. Data points or data series can be automatically colored by the value of a variable with a variety of color maps, and model predictions can also be colored by the value of a fit statistic. Comparisons between data sets and the predictions/simulations of different models on the same data set can be easily plotted as well. Distributions of parameter values from bootstrap resampling can be plotted to visualize the reliability of parameter estimation, or determine any correlations between parameters. Modeling tools handle the complex task of creating and parsing symbolic equations for modeling battery lifetime. Equations are parsed to grab relevant data variables, parameter values, or specified sub-models for input into optimization, evaluation, or simulation functions. Models can be optimized locally (one set of parameters for each data series), bi-level (some parameters shared across the data set), or globally (single set of parameters for all data). Functions implementing symbolic regression algorithms help users to discover effective model equations, even in poorly sampled, high-dimensional data.

Smith, Kandler [National Renewable Energy Lab. (NR↗