Time-parallel multiple-shooting for multi-qubit optimal control
Explore the source record for details and available documents.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
A metalens array is disclosed for controllably modifying a phase of a wavefront of an optical beam. The metalens array may have a substrate having at least first and second metalens unit cells, and forming a single integrated structure with no stitching being required of the first and second metalens unit cells. The first metalens unit cell has a first plurality of nanoscale features and is configured to modify a phase of a first portion of a wavefront of an optical signal incident thereon in accordance with a first predetermined phase pattern to create at least one first focal voxel within an image plane. The second metalens unit cell has a second plurality of nanoscale features configured to modify the phase of a second portion of the wavefront of the optical signal incident thereon, in accordance with a second predetermined phase pattern, to simultaneously create at least one second focal voxel within the image plane. Each metalens unit cell also has an overall diameter of no more than about 200 microns.
Explore the source record for details and available documents.
Bayesian inference enables informationally efficient quantum state tomography (QST) yet is challenging to scale computationally. We demonstrate a parallelizable Bayesian QST method that, although unorthodox, proves remarkably practical, attaining significant speedups in multiqubit state estimation.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
The coming massive parallelism of exascale computing presents a pressing challenge for the many DOE simulations of time-dependent partial differential equations (PDEs), which typically use traditional sequential time stepping methods. Since this traditional approach is inherently serial, it presents a sequential bottleneck when moving to exascale computing, because future performance gains will come through greater concurrency, not faster clock speeds. Thus, the goal of this work is to research parallelism in time, i.e., methods that compute multiple time values simultaneously, not sequentially. The focus will be on hyperbolic and chaotic problems of interest to DOE, with the goal of enabling scalable simulations of time-dependent hyperbolic and chaotic problems on future architectures. The chosen methodology for solving these problems parallel-in-time is multigrid, because multigrid (when it works) is a powerful, optimal, and scalable solver for discretized PDEs. Multigrid is already commonly used in many DOE simulations for scalably and optimally solving space-only PDE problems. The areas of hyperbolic and chaotic problems are chosen because of their relevance to problems of programmatic interest to DOE. However, these problems are also well-known to be difficult for parallel-in-time methods, with the most common method, parareal, diverging in many cases. The current state of-the-art for parallel-in-time at LLNL is the multigrid reduction in time (MGRIT) XBraid package, which also struggles for such problems, while still showing some improvement over parareal. In summary, new methods are needed for an efficient parallel-in-time scheme for hyperbolic and chaotic problems, and this work shall research promising new multigrid methods in this area. In particular, this work shall continue researching the directions from the current collaboration with Dr. Falgout, which are laid out in the work Toward Parallel in Time for Chaotic Dynamical Systems and showed the first known results of a parallel-in-time speedup for a chaotic problem. This work outlines two key improvements to XBraid for chaotic problems, the so-called “theta” and “delta-correction” methods. Here, these two improvements will be further researched and improved (including with a new relaxation method inspired by on Least Squares Shadowing (LSS)) and explored for more complicated problems.
Finding frequent subgraph patterns in a big graph is an important problem with many applications such as classifying chemical compounds and building indexes to speed up graph queries. Since this problem is NP-hard, some recent parallel and distributed systems have been developed to accelerate the mining. However, they often have a huge memory cost, very long running time, suboptimal load balancing, poor scale-out capability, and possibly inaccurate results. In this article, we propose an efficient system called T-FSM for parallel mining of frequent subgraph patterns in a big graph. T-FSM supports a new anti-monotonic frequentness measure called Fraction-Score, which is more accurate than the widely used MNI measure. The execution engine of T-FSM supports both intra-machine parallelism and inter-machine parallelism. For intra-machine parallelism, T-FSM adopts a novel task-based execution model to ensure high multithreading concurrency, bounded memory consumption, and effective load balancing. For inter-machine parallelism, T-FSM ensures good scale-out performance with a lightweight pattern rebalancing approach that reduces workload skewness of pattern evaluations among machines. To avoid recomputing the contexts for migrated patterns, we design a novel context cache table to support concurrent and asynchronous requesting and caching of remote context data, which can timely evict and garbage collect used pattern contexts that are no longer needed to keep memory consumption bounded. Extensive experiments show that T-FSM is orders of magnitude faster than existing state-of-the-art parallel systems (more than 10×, 51×, 131×, 55× speedup over ScaleMine, DistGraph, Pangolin and Peregrine, respectively) and distributed systems (more than 42× and 88× over ScaleMine and DistGraph, respectively) for frequent subgraph pattern mining, and it scales out satisfactorily to 512 CPU cores on the Polaris supercomputer at Argonne National Laboratory.
Controlling the orientation of nanostructured block copolymer (BCP) thin films is essential for their use in templating, transport, and pattern transfer. Conventional efforts mainly focus on adjusting enthalpic interactions between the blocks and interfaces, while entropic contributions are often overlooked. Here, we show that the morphology of BCP thin films can be precisely tuned by the architectural design of star BCPs. Specifically, we synthesized multiarm star BCPs with a polystyrene (PS) core and poly(2-vinylpyridine) (P2VP) corona, which exhibits a lamellar microdomain morphology. The entropic penalty associated with a parallel orientation of the microdomains to the substrate is controlled by varying the number of arms comprising the star BCPs, from 2-arms (triblock) to 3-arms and 4-arms. We systematically investigated the thin film morphology at different depths using grazing incidence small-angle X-ray and neutron scattering (GISAXS and GISANS), atomic force microscopy (AFM), water contact angle (WCA), and interference microscopy. The results show that 2-arm star BCPs show a parallel orientation, the 3-arm star BCPs form a uniform PS film at the air surface with a vertical orientation of the microdomains underneath, and the 4-arm star BCPs exhibit a parallel microdomain orientation at the air surface with mixed parallel and perpendicular microdomain orientation in the bulk. Additionally, we found that the inclination angle of microdomains at the edges of islands and holes, resulting from the incommensurability between film thickness and the characteristic period of the microdomain morphology, increases with a higher number of arms. This suggests a greater grain boundary tilt angle in the microdomains of star-shaped block copolymers (BCPs). When silicon substrates were modified with PS homopolymer, the selective interaction between substrate and core blocks promotes a parallel orientation for the 4-arm star BCPs. In conclusion, this work shows that control of arm number in star BCPs affords diverse BCP thin film morphologies, offering insights into the star BCP conformations in thin films across different depths.
This study employs electron-scale gyrokinetic simulations to investigate the electron temperature gradient (ETG) driven instabilities, turbulence, and transport in the pedestal region of the National Spherical Torus Experiment, comparing non-lithiated (narrow pedestal) and lithiated (wide pedestal) scenarios. Our findings reveal that, in the non-lithiated case, a branch of strongly unstable ETG modes exhibiting finite parallel magnetic field fluctuations ($\delta B_{\parallel} \neq 0$) emerges at the pedestal top and upper density pedestal region. This branch is uncovered only when $\delta B_{\parallel}$ is retained in the simulations and is associated with substantial electrostatic electron heat flux. This region of strong ETG transport corresponds to the only region in the plasma where the pressure gradient is far below the critical gradient for kinetic ballooning modes. We investigated the origin of this finite $\delta B_{\parallel}$ ETG branch by analyzing the gyrokinetic field equations. Nonlinear saturation is also analyzed and contrasted for simulations with and without $\delta B_{\parallel}$. In contrast with the nonlithiated case, ETG modes in the lithiated case produce substantial transport in the steep gradient region, but are negligible at the pedestal top.
Many polyatomic astrophysical plasmas are compressible and out of chemical and thermal equilibrium, introducing a bulk viscosity into the plasma via the internal degrees of freedom of the molecular composition, directly impacting the decay of compressible modes, $\mathrm{{\boldsymbol {\mathit {v}}}}_{\parallel }(\boldsymbol {k})$. This is especially important for small-scale, turbulent dynamo processes in the interstellar medium (ISM), which are known to be sensitive to the effects of compression. To control the viscous properties of $\mathrm{{\boldsymbol {\mathit {v}}}}_{\parallel }(\boldsymbol {k})$, we perform trans-sonic, visco-resistive dynamo simulations with additional bulk viscosity $\nu _{\text{bulk}}$, deriving a new $\nu _{\text{bulk}}$ Reynolds number $\text{Re}_{\text{bulk}}$, and viscous Prandtl number $\text{P}\nu \equiv \text{Re}_{\text{bulk}}/ \text{Re}_{\text{shear}}$, where $\text{Re}_{\text{shear}}$ is the shear viscosity Reynolds number. We derive a framework for decomposing $E_{\rm mag}$ growth rates into incompressible and compressible terms via orthogonal tensor decompositions of $\boldsymbol {\nabla }\otimes \mathrm{{\boldsymbol {\mathit {v}}}}$, where $\mathrm{{\boldsymbol {\mathit {v}}}}$ is the fluid velocity. We find that $\mathrm{{\boldsymbol {\mathit {v}}}}_{\parallel }(\boldsymbol {k})$ play a dual role, growing and decaying $E_{\rm mag}$, and that field-line stretching is the main driver of growth, even in compressible dynamos. In the absence of $\nu _{\text{bulk}}$ ($\text{P}\nu \rightarrow \infty$), $\mathrm{{\boldsymbol {\mathit {v}}}}_{\parallel }(\boldsymbol {k})$ pile up on small-scales, creating a spectral bottleneck, which disappears for $\text{P}\nu \approx 1$. As $\text{P}\nu$ decreases, $\mathrm{{\boldsymbol {\mathit {v}}}}_{\parallel }(\boldsymbol {k})$ are dissipated at increasingly larger scales, in turn suppressing incompressible modes through a coupling between high-k modes. We emphasize the importance of further understanding the role of $\nu _{\text{bulk}}$ in compressible astrophysical plasmas, which we estimate could be as strong as the shear viscosity in the cold ISM, and highlight that compressible direct numerical simulations without bulk viscosity have unresolved compressible mode dissipation scales.
The new version of HydraGNN v4.0 provides additional core capabilities, such as: Inclusion of multi-body atomistic cluster expansion MACE, polarizable atom interaction neural network PAINN, and equivariant principal neighborhood aggregation (PNAEq) among the message passing layers supported -Inclusion of graph transformers to directly model long-range interactions between nodes that are distant in the graph topology Integration of graph transformers with message passing layers by combining the graph embedding generated by the two mechanisms, which allows for an improved expressivity of the HydraGNN architecture Improved re-implementation of multi-task learning (MTL) to allow its use for stabilized training across imbalanced, multi-source, multi-fidelity data Introduction of multi-task parallelism, a newly proposed type of model parallelism specifically for MTL architectures, which allows to dispatch different output decoding heads to different GPU devices Integration of multi-task parallelism with pre-existing distributed data parallelism to enable a 2D parallelization for distributed training Improved portability of the distributed training across Intel GPUs, which has been testes on ALCF exascale supercomputer Aurora Inclusion of 2-level fine-grained energy profilers portable across NVIDIA, AMD, and Intel GPUs to monitor the power and energy consumption associated with different functions executed by the HydraGNN code during data pre-load and training Restructuring of previous examples and inclusion of new sets of examples to illustrate the download, preprocess, and training of HydraGNN models on new large-scale open-source datasets for atomistic materials modeling (e.g., Alexandria, Transition1x, OMat24, OMol25)
This is a modification to 'cp' and 'mv' commands to make them multi-threaded. Simple benchmarks showed that multi-threading could reduce the time to copy a large Linux source directory by over 2x. The 'cp' and 'mv' utilities are part of the existing Coreutils (https://www.gnu.org/software/coreutils/) software package that get installed on all Linux distros. Changes: * Add '-j|--parallel ' flags to 'cp' and 'mv'. This allows the utilities to recursively copy regular files in directories in parallel. This does NOT parallelize multiple single file copies to a destination (like 'cp file2 file2 file3 dst/'). Along with this, add in new 'CP_NUM_THREADS' and 'MV_NUM_THREADS' environment variables to set the number of threads. This can be useful when you want to enable parallelism by default in /etc/profile. The maximum number of threads is internally capped to the number of CPUs. * Add a '-j' flag to 'sort' to complement its existing '--parallel' flag. This is only done for consistency with 'cp' and 'mv'. * Add test cases for the new flags. Also, run each 'cp' and 'mv' test both in single-threaded and multithreaded modes for extra coverage.
ion Library for Parallel Kernel Acceleration) is a header-only C++ library that provides performance portability across different back-ends, abstracting the underlying levels of parallelism. It supports serial and parallel execution on CPUs, and extremely parallel execution on NVIDIA, AMD and Intel GPUs.This contribution will show how alpaka is used in the CMS software to develop and maintain a single code base; to use different toolchains to build the code for each supported back-end, and link them into a single application; to seamlessly select the best backend at runtime, and implement portable reconstruction algorithms that run efficiently on CPUs and GPUs from different vendors. It will describe the validation and deployment of the alpaka-based implementation in the CMS High Level Trigger, and highlight how it achieves near-native performance.
Achieving core-edge compatible divertor detachment is a critical requirement for stable operation in future fusion devices. This study compares nitrogen and neon seeding in KSTAR H-mode plasmas with carbon walls, combining experiments and SOLPS-ITER modelling to evaluate their radiative dissipation and core-edge compatibility. Experimentally, N seeding achieved stronger divertor detachment, with larger reductions in target particle and heat fluxes, a higher divertor radiation fraction, and a lower core radiation fraction than Ne. In contrast, Ne seeding triggered a significant rise in core radiation followed by H–L back transitions, limiting the maximum total radiated power fraction to roughly half that of N. SOLPS-ITER simulations reproduced the experimental trends and revealed that the better core-edge compatibility of N arises from its higher divertor retention in addition to its higher cooling factor. The relative positions of the stagnation points of impurity poloidal velocity and ionization sources did not explain the different divertor compression. Instead, in the present modelling, the higher impurity parallel particle flux, resulting from the higher parallel impurity velocity, explains the stronger nitrogen impurity compression in the divertor region. The parallel temperature distribution with N was more favourable for achieving higher impurity parallel velocity than with Ne, because the impurity velocity is governed by modifications of the main ion flow due to friction and thermal forces, both of which strongly depend on the temperature. Ultimately, this behaviour is attributed to the strongly divertor-localized radiation of N. These results demonstrate that N is more effective than Ne in achieving radiative divertor detachment while maintaining low core contamination in KSTAR, consistent with observations in other present tokamaks.
Entanglement is a resource to improve the sensitivity of quantum sensors. In an ideal case, using an entangled state as a probe to detect target fields, we can beat the standard quantum limit by which all classical sensors are bounded. However, since entanglement is fragile against decoherence, it is unclear whether entanglement-enhanced metrology is useful in a noisy environment. Its benefit is indeed limited when estimating the amplitude of dc magnetic fields under the effect of parallel Markovian decoherence, where the noise operator is parallel to the target field. In this paper, on the contrary, we show an advantage to using an entanglement over the classical strategy under the effect of parallel Markovian decoherence when we try to detect ac magnetic fields. We consider a scenario to induce a Rabi oscillation of the qubits with the target ac magnetic fields. Although we can, in principle, estimate the amplitude of the ac magnetic fields from the Rabi oscillation, the signal becomes weak if the qubit frequency is significantly detuned from the frequency of the ac magnetic field. We show that, by using the Greenberger-Horne-Zeilinger (GHZ) states, we can significantly enhance the signal of the detuned Rabi oscillation even under the effect of parallel Markovian decoherence. Further, our method is based on the fact that the interaction time between the GHZ states and ac magnetic fields scales as 1/L to mitigate the decoherence effect, where L is the number of qubits, which contributes to improving the bandwidth of the detectable frequencies of the ac magnetic fields. Our results pave the way for new applications of entanglement-enhanced ac magnetometry.
We propose an algorithm to efficiently perform latency-bound communication scenarios that consist of many small messages. In these parallel scenarios, processes typically pass around a lot of small-sized messages of a few KBs of size. Performing communication operations with P2P MPI routines or collective MPI routines (including neighborhood collectives) in such scenarios may not always yield the optimal results and may not resolve the latency bottleneck. To this end, we develop a regular structure called virtual process topology (VPT) on which the messages can be communicated in a structured and controlled manner. Using parameters of this topology, one can tune the rate of aggression in tackling the latency costs. We demonstrate that our communication algorithm is preferable to MPI P2P and collective routines for latency-bound communication and it can easily be adapted only by replacing calls to MPI routines in a parallel application. We show how to adapt existing topology-aware mapping heuristics to address the volume overhead due to communicating messages on the VPT. Moreover, we propose a novel swap-based mapping heuristic to address this overhead by optimizing the maximum volume handled by a process. Experiments on synthetic communication graphs as well as real-world applications such as parallel Canonical Polyadic sparse tensor decomposition and parallel sparse matrix-dense matrix multiplication show that our approach is a powerful way of overcoming the bottlenecks posed by sparse and latency-bound irregular communication.