Search NASA⌕ Search

SEARCH · Search NASA

Results for “program processors”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Techniques for recovering from errors when executing software applications on parallel processors

In various embodiments, a software program uses hardware features of a parallel processor to checkpoint a context associated with an execution of a software application on the parallel processor. The software program uses a preemption feature of the parallel processor to cause the parallel processor to stop executing instructions in accordance with the context. The software program then causes the parallel processor to collect state data associated with the context. After generating a checkpoint based on the state data, the software program causes the parallel processor to resume executing instructions in accordance with the context.

Hukerikar, Saurabh↗

Systems and methods for single-axis tracking via sky imaging and machine leanring comprising a neural network to determine an angular position of a photovoltaic power system

A system and method is disclosed for solar tracking and controlling an angular position of a photovoltaic power system. The solar tracking system includes an imaging device for capturing images of the sky; a solar position data generating module; and a control system comprising a neural network. The neural network has multiple convolutional layers to generate a first output associated with the images, and a solar position data module. A first dense layer module receives the solar position data and generates a second output. A second dense layer module receives the first output and the second output and generates a concatenated data sequence. A processor is programmed to generate a multi-planar irradiance signal (MPIS) in response to the concatenated data sequence, and determine an angular position of the PV power system and adjust the angular position in response to an angle of maximum irradiance.

Stein, Joshua↗

Equilipy

There is high demand for high-throughput calculations of phase equilibria based on the CALPHAD method (CALculation of PHAse Diagram). High-throughput calculations are possible through HPC, currently by a commercial program called Thermo-Calc. However, the number of nodes/processors are limited to the number of purchased license (16 processors per license). This program provides a toolkit for high-throughput calculations of phase equilibria by the CALPHAD method. The program is prepared for Python environment, so that it can be easily installed and used together with other open-source programs. The program can be run in supercomputer.

Kwon, Sunyong [Oak Ridge National Laboratory (ORNL↗

Parallel hybrid quantum-classical machine learning for kernelized time-series classification

Supervised time-series classification garners widespread interest because of its applicability throughout a broad application domain including finance, astronomy, biosensors, and many others. Here, in this work, we tackle this problem with hybrid quantum-classical machine learning, deducing pairwise temporal relationships between time-series instances using a timeseries Hamiltonian kernel (TSHK). A TSHK is constructed with a sum of inner products generated by quantum states evolved using a parameterized time evolution operator. This sum is then optimally weighted using techniques derived from multiple kernel learning. Because we treat the kernel weighting step as a differentiable convex optimization problem, our method can be regarded as an end-to-end learnable hybrid quantum-classical-convex neural network, or QCC-net, whose output is a data set-generalized kernel function suitable for use in any kernelized machine learning technique such as the support vector machine (SVM). Using our TSHK as input to a SVM, we classify univariate and multivariate time-series using quantum circuit simulators and demonstrate the efficient parallel deployment of the algorithm to 127-qubit superconducting quantum processors using quantum multi-programming.

97 MATHEMATICS AND COMPUTING↗

Cache management based on access type priority

Systems, apparatuses, and methods for cache management based on access type priority are disclosed. A system includes at least a processor and a cache. During a program execution phase, certain access types are more likely to cause demand hits in the cache than others. Demand hits are load and store hits to the cache. A run-time profiling mechanism is employed to find which access types are more likely to cause demand hits. Based on the profiling results, the cache lines that will likely be accessed in the future are retained based on their most recent access type. The goal is to increase demand hits and thereby improve system performance. An efficient cache replacement policy can potentially reduce redundant data movement, thereby improving system performance and reducing energy consumption.

Yin, Jieming↗

Programmable simulations of molecules and materials with reconfigurable quantum processors

Simulations of quantum chemistry and quantum materials are believed to be among the most important applications of quantum information processors. However, realizing practical quantum advantage for such problems is challenging because of the prohibitive computational cost of programming typical problems into quantum hardware. Here we introduce a simulation framework for strongly correlated quantum systems represented by model spin Hamiltonians that uses reconfigurable qubit architectures to simulate real-time dynamics in a programmable way. Our approach also introduces an algorithm for extracting chemically relevant spectral properties via classical co-processing of quantum measurement results. We develop a digital–analogue simulation toolbox for efficient Hamiltonian time evolution using digital Floquet engineering and hardware-optimized multi-qubit operations to accurately realize complex spin–spin interactions. As an example, we propose an implementation based on Rydberg atom arrays. In addition, we show how detailed spectral information can be extracted from the dynamics through snapshot measurements and single-ancilla control, enabling the evaluation of excitation energies and finite-temperature susceptibilities from a single dataset. To illustrate the approach, we show how to use the method to compute key properties of a polynuclear transition-metal catalyst and two-dimensional magnetic materials.

74 ATOMIC AND MOLECULAR PHYSICS↗

Parallel Variable Population Multi-Objective Optimizer (pvpmoo) v1.0

This is a parallel variable population multi-objective optimizer with an adaptive unified differential evolution algorithm or a genetic algorithm. It can also be used for single objective optimization. Some features of this code include: 1) The population size varies from generation to generation to save the total # of objective function evaluations. 2) The population is uniformly distributed to a number of parallel processors for simultaneous objective function evaluation. 3) The objective function evaluation can be attained from an external simulation program with control variables in its input file and objectives calculated from its output files. 4) The optimizer includes an adaptive unified differential evolution algorithm and a real value genetic algorithm. The parameters in the unified differential evolution algorithm can be chosen to attain any mutation schemes in the published literature.

Qiang, Ji↗

Equilipy: a python package for calculating phase equilibria

The CALPHAD (CALculation of PHAse Diagram) approach (Nigel Saunders & Miodownik, 1998) provides predictions for thermodynamically stable phases in multicomponent-multiphase materials across a wide range of temperatures. Consequently, the CALPHAD calculations became an essential tool in materials and process design (Luo, 2015). Such design tasks frequently require navigating a high-dimensional space due to multiple components involved in the system. This increasing complexity demands high-throughput CALPHAD calculations, especially in the rapidly evolving field of alloy design. In response to the need, we developed Equilipy an open-source Python package designed for calculating phase equilibria of multicomponent-multiphase systems. Equilipy is specifically tailored for high-throughput CALPHAD calculations, offering parallel computations across multiple processors and nodes with the given NPT input conditions namely elemental compositions (N), pressure (P), and temperature (T). Equilipy utilizes the program structure and Gibbs energy functions from the Fortran-based program, Thermochimica (Piro et al., 2013), with incorporating a new Gibbs energy minimization algorithm. This algorithm, originally developed by Capitani and Brown in 1987 (Capitani & Brown, 1987), has been revised and implemented to enhance the stability and performance of calculations. The Fortran codes are precompiled and interfaced with Python via F2PY, ensuring high computation speed. Benchmark tests shown in Figure 1 demonstrate that Equilipy’s computation speed is comparable to those of established commercial software, TC-Python and PanPython. This result highlights its efficiency and potential applications in various scientific and industrial fields.

97 MATHEMATICS AND COMPUTING↗

Risk-informed Graded Approach for Reliability and Performance Assessment of Sensor and Instrumentation Systems within Advanced Condition Monitoring Technologies

Advanced condition monitoring (ACM) technologies, such as digital twins, are innovative strategies designed to provide real-time health insights, including the remaining useful life of components. The primary goal of ACM is to predict and alert operators to potential functional failures before they occur. ACM systems achieve this by integrating predictive models with various sensor instrumentation, analog-to-digital converters, data warehouses, and data pre-processors. These sensor and instrumentation systems (SIS) are essential for forming a comprehensive understanding of component conditions and ensuring the predictive success of ACM programs. Introducing new technologies like ACM involves varying degrees of risk that can impact plant reliability. Therefore, risk mitigation should be commensurate with the performance and reliability of the developed technology, following a risk-informed graded approach (RIGA). Establishing a RIGA process requires a clear understanding of the hazards and reliability of all subsystems, including their interdependencies and potential impacts on the overall system. Given the critical role of SIS in ACM, this work reviews hazard identification and reliability quantification methods for SIS. It also considers these methods' implications when developing a RIGA process for ACM.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Classic and Quantum Task-Based Intelligent Runtime for QIRs Running on Multiple QPUs

High-performance computing systems are rapidly evolving into heterogeneous platforms that fuse quantum accelerators with traditional classical processing units (CPUs) and graphical processing units (GPUs). This convergence calls for runtimes capable of managing both classical and quantum workloads in a unified manner. We introduce an intelligent, task-based runtime that marries the Intelligent RuntIme System (IRIS) asynchronous scheduler with a quantum programming stack through the Quantum Intermediate Representation Execution Engine (QIR-EE). Our design allows programs written in the quantum intermediate representation (QIR) to be dispatched concurrently to a variety of back-ends, including multiple quantum simulators and nascent quantum processors, enabling genuine hybrid execution on a single node. To illustrate its practicality, we partition a 4-qubit and 20-qubit circuit into three sub-circuits using quantum circuit cutting via the QCut library. Each sub-circuit is simulated independently by the QIR-EE driver within IRIS, after which a classical post-processing step merges the simulation results to recover the outcome of the original full-circuit computation. This case study demonstrates how finer task granularity can enable the parallel execution and lower the simulation burden per quantum task while preserving overall accuracy, highlighting the feasibility of our hybrid approach.

Miniskar, Narasinga Rao [ORNL] (ORCID:000000018259↗

Toward coherent quantum computation of scattering amplitudes with a measurement-based photonic quantum processor

In recent years, applications of quantum simulation have been developed to study the properties of strongly interacting theories. This has been driven by two factors: on the one hand, needs from theorists to have access to physical observables that are prohibitively difficult to study using classical computing; on the other hand, quantum hardware becoming increasingly reliable and scalable to larger systems. In this work, we discuss the feasibility of using quantum optical simulation for studying scattering observables that are presently inaccessible via lattice QCD and are at the core of the experimental program at Jefferson Laboratory, the future Electron-Ion Collider, and other accelerator facilities. We show that recent progress in measurement-based photonic quantum computing can be leveraged to provide deterministic generation of required exotic gates and implementation in a single photonic quantum processor. Published by the American Physical Society 2024

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Advancements in Multiphysics Microdepletion Analysis of an eVinci TM -like Microreactor Leveraging OpenMC-CRAB Workflow

Nuclear microreactors (MRs) are a class of nuclear reactor technology, characterized by reduced dimensions, modular design, and reduced power output in contrast to conventional Light Water Reactors (LWRs). MRs are proposed for supplying electricity and eventual process heat to remote locations, such as military installations and disaster-affected areas. Current research work sponsored by the US Department of Energy Microreactor Program (MRP) is devoted to the development of novel modeling and simulation tools to better support MR vendors and regulatory bodies. Notably, the NRC is projected to utilize the CRAB multiphysics software driver for executing both design and beyond-design-basis accident analyses. Furthermore, the NRC has been utilizing the MELCOR code to calculate mechanistic source terms during accidents. Since MELCOR relies on isotopic inventory and reactor temperature/power profiles under accident conditions, which theoretically can be derived from CRAB, the goal is to establish a comprehensive CRAB-MELCOR computational framework. Past work was focused on testing and demonstrating CRAB's capability to generate results that can be used to inform mechanistic source term calculations in MELCOR. In particular, a computational workflow leveraging OpenMC-generated microscopic cross sections and CRAB was first applied to perform multiphysics microscopic depletion calculation followed by an accident scenario for a stylized microreactor problem. In fiscal year 2024, the research work has been focused on applying the OpenMC-CRAB workflow, which was first tested in fiscal year 2023, to a realistic 3D heat-pipe cooled MR problem representative of the eVinci TM design. The latter computational problem was developed with inputs from WEC to conserve selected neutronic and thermal characteristics of the eVinci TM design without releasing proprietary data. The results of this simulation, encompassing isotopic inventory, power density distribution, and kinetic parameters, will inform both MELCOR and the WEC-developed FATE code for mechanistic source terms calculations. The results from the two codes will then be compared for code verification purposes. This report contains the design characteristics of the realist heat pipe cooled microreactor developed as a use-case for the verification exercise, and the current results for the multiphysics microscopic depletion performed with the OpenMC-CRAB workflow. The results include eigenvalue as a function of time, power distribution at EOL, in addition to nuclides inventory's time evolution and spatial distribution. Finally, we report improvements to the workflow efficiency achieved through a collaboration with the NEAMS programs. Through this collaborative effort, we were able to strongly decrease the computational time for the multiphysics microdepletion calculation (i.e., from 17.4 hours to 5.7 hours on 280 processors) in addition to simplifying the interface to generate isotopics spatial distribution utilizable by FATE and MELCOR. Future work, including the improvement of the current microscopic cross-sections' library and the simulation of an accident scenario at EOL, is also discussed.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

IceNet for FireBox - A Berkeley Warehouse-Scale Computer

Berkeley’s FireBox is a next-generation warehouse-scale computer (WSC) that utilizes the energy-efficiency and bandwidth density of integrated silicon-photonic interconnects to enable a new high-bandwidth and low-latency network fabric connecting thousands of compute nodes to petabytes of DRAM and Flash storage. The high bandwidth, low latency and high connectivity of FireBox’s WSC network fabric (IceNet) will enable dramatic improvements in the overall system energy efficiency enabling fine-grain power control on system resources (processors, links and memory/storage components). IceNet is a special 3-stage photonic Clos network architected to achieve ultra-low-latency connectivity between processor nodes and memory, drastically cutting down on the energy wasted in resource idling (processors and memory stalled due to pending network requests). This is achieved by integration of the first and last switch stages into processor/memory hub clients and by heavy over-provisioning of the high-radix middle switches (FlareSwitches). A key hardware component developed in this program is an active laser power management photonic integrated circuits called LightSpark. It interacts with the FlareSwitch and provides laser power to a subset of occupied switch ports, increasing the utilization of laser light in the photonic network by an order of magnitude. In addition to guiding the laser power where it is needed, the laser-power management module enables both wavelength and laser redundancy, significantly increasing the robustness of the system. The goal of the IceNet fabric is to enable communication between 1000s of processor nodes and PBs of memory/storage with <100ns latency, <10pJ/b wall-plug energy cost at multiple Pb/s of available connectivity bandwidth. These metrics represent two-orders of magnitude improvement with respect to the status of current data-center technology.

42 ENGINEERING↗

Field Programmable Gate Array Data Capture for Control Systems

Some Industrial Control Systems (ICS) networks are based on protocols such as Serial and Industrial Ethernet. These protocols currently have no existing cybersecurity monitoring tools, leaving a large gap in the cyber defense of critical infrastructure. In order to analyze such ICS traffic, it is first necessary to implement methods of capturing the ICS data. Whereas traditional methods of analyzing data would use microprocessors, the nature of high-speed analog data can be difficult to implement on such a versatile processor, as they are rather inefficient for doing a single task. Whereas Field Programmable Gate Arrays (FPGAs) provide an adequate tool in analyzing high speed data, as despite the lack of program versatility, Programmable Logic can implement a solution with minimal clock cycles, allowing time for each new packet of data to be captured before a new data sample is taken.

42 ENGINEERING↗

4-Clique network minor embedding for quantum annealers

Quantum annealing is a quantum algorithm for computing solutions to combinatorial optimization problems. This study proposes a method for minor embedding optimization problems onto sparse quantum annealing hardware graphs called 4-clique network minor embedding. This method is in contrast to the standard minor embedding technique of using a path of linearly connected qubits in order to represent a logical variable state. The 4-clique minor embedding is possible on Pegasus graph connectivity, which is the native hardware graph for some of the current D-Wave quantum annealers. The Pegasus hardware graph contains many cliques of size 4, making it possible to form a graph composed entirely of paths of connected 4-cliques on which a problem can be minor-embedded. The 4-clique chains come at the cost of additional qubit usage on the hardware graph, but they allow for stronger coupling within each chain, thereby increasing chain integrity, reducing chain breaks, and allow for greater usage of the available energy scale for programming logical problem coefficients on current quantum annealers. The 4-clique minor embedding technique is compared with the standard linear path minor embedding with experiments on two D-Wave quantum annealing processors with Pegasus hardware graphs. We show proof-of-concept experiments where the 4-clique minor embeddings can use weak chain strengths while successfully carrying out the computation of minimizing random all-to-all spin glass problem instances. Published by the American Physical Society 2024

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Compiling Quantum Circuits for Dynamically Field-Programmable Neutral Atoms Array Processors

Dynamically field-programmable qubit arrays (DPQA) have recently emerged as a promising platform for quantum information processing. In DPQA, atomic qubits are selectively loaded into arrays of optical traps that can be reconfigured during the computation itself. Leveraging qubit transport and parallel, entangling quantum operations, different pairs of qubits, even those initially far away, can be entangled at different stages of the quantum program execution. Such reconfigurability and non-local connectivity present new challenges for compilation, especially in the layout synthesis step which places and routes the qubits and schedules the gates. In this paper, we consider a DPQA architecture that contains multiple arrays and supports 2D array movements, representing cutting-edge experimental platforms. Within this architecture, we discretize the state space and formulate layout synthesis as a satisfiability modulo theories problem, which can be solved by existing solvers optimally in terms of circuit depth. For a set of benchmark circuits generated by random graphs with complex connectivities, our compiler OLSQ-DPQA reduces the number of two-qubit entangling gates on small problem instances by 1.7x compared to optimal compilation results on a fixed planar architecture. To further improve scalability and practicality of the method, we introduce a greedy heuristic inspired by the iterative peeling approach in classical integrated circuit routing. Using a hybrid approach that combined the greedy and optimal methods, we demonstrate that our DPQA-based compiled circuits feature reduced scaling overhead compared to a grid fixed architecture, resulting in 5.1X less two-qubit gates for 90 qubit quantum circuits. These methods enable programmable, complex quantum circuits with neutral atom quantum computers, as well as informing both future compilers and future hardware choices.

Physics↗

The VTK-m User's Guide (V. 2.2)

High-performance computing relies on ever finer threading. Advances in processor technology include ever greater numbers of cores, hyperthreading, accelerators with integrated blocks of cores, and special vectorized instructions, all of which require more software parallelism to achieve peak performance. Traditional visualization solutions cannot support this extreme level of concurrency. Extreme scale systems require a new programming model and a fundamental change in how we design algorithms. To address these issues we created VTK-m: the visualization toolkit for multi-/many-core architectures. VTK-m supports a number of algorithms and the ability to design further algorithms through a top-down design with an emphasis on extreme parallelism. VTK-m also provides support for finding and building links across topologies, making it possible to perform operations that determine manifold surfaces, interpolate generated values, and find adjacencies. Although VTK-m provides a simplified high-level interface for programming, its template-based code removes the overhead of abstraction.

97 MATHEMATICS AND COMPUTING↗