A Framework to Demonstrate a DNP3 Interface with a CIM-Based Data Integration Platform
This poster was presented at the 2024 IEEE Power & Energy Society General Meeting, July 21-25, 2024, Seattle, Washington.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
This poster was presented at the 2024 IEEE Power & Energy Society General Meeting, July 21-25, 2024, Seattle, Washington.
Existing electromagnetic transient (EMT) simulation tools face challenges in accelerating EMT simulations, especially for very large-scale power networks. To tackle this issue, next generation EMT simulation tools such as RE-INTEGRATE EMT are being researched upon. Such tools should be equipped with automation capabilities and advanced numerical differential-algebraic equation (DAE) solvers. In this paper, the DAE solvers incorporated within the RE-INTEGRATE EMT simulation tool are discussed. In particular, a modified ODEINT-based DAE solver and the ARKODE solver from SUN-DIALS are leveraged within RE-INTEGRATE EMT. In addition, the automation implemented within RE-INTEGRATE EMT to automate the DAE generation (replacing the need of manual discretization and assembling DAEs) is discussed. Different use cases were implemented using the RE-INTEGRATE EMT tool and were validated with respect to baseline simulations.
The realization of quantum networks requires the development of devices and methods unprecedented in conventional networks, and yet they critically depend on the latter for implementing foundational blocks and essential operations. We describe a testbed to support the development and testing of their functionality and performance by providing quantum and conventional data planes and devices, together with a secure conventional control plane. It incorporates a variety of entangled photon sources, qubit technologies, detector technologies, photonic components, and supporting conventional switches and workstations. It implements a novel fiber telescoping scheme that provides suites of connections using fiber spools and inground-aerial fiber loops. We briefly summarize a variety of experiments conducted over this testbed including: (i) flex-grid quantum connection experiments, (ii) quantum state and channel tomography, (iii) utilization of quantum key distribution keys to secure conventional encryption and firewall devices, (iv) comparative study of analytical capacity estimates and entanglement throughput, (v) deployed squeezing coexisting with conventional communications, and (iv) measurement of polarization time variation.
Graphics Processing Units (GPUs) offer significant potential for accelerating various computational tasks, including Breadth-First Search (BFS). Numerous efforts have been made to deploy BFS on GPUs effectively. To address the dynamic nature of BFS, XBFS, the state-of-the-art work, employs an adaptive strategy that leverages different optimized frontier queue generation designs, accommodating the varying characteristics of levels in BFS. While XBFS demonstrates excellent performance on NVIDIA Quadro P6000 GPUs, it faces challenges when deployed on AMD GPUs. In this work, we present our efforts to implement XBFS’s adaptive approach on Frontier, the most powerful supercomputer system, by porting XBFS to AMD MI250X GPUs. Through targeted optimizations tailored to the unique features of AMD GPUs, our implementation achieves an average performance of 43 Giga-Traversed Edges Per Second (GTEPS) per Graphics Compute Dies (GCD). Based on these results, we observe potential for surpassing the performance of the official Frontier results from the Graph500 benchmark released in June 2024.
We explore the development of a performance-portable CPU/GPU ecosystem to integrate two of the US Department of Energy’s (DOE’s) largest scientific instruments, the Oak Ridge Leadership Computing facility and the Spallation Neutron Source (SNS), both of which are housed at Oak Ridge National Laboratory. We select a relevant data reduction workflow use-case to obtain the differential scattering cross-section from data collected by SNS’s CORELLI and TOPAZ instruments. We compare the current CPU-only production implementation using the Garnet Python multiprocess package based on the Mantid C++ framework against our proposed CPU/GPU implementation that uses the LLVM-based, just-in-time Julia scientific language and the JACC.jl performance-portable package. Two proxy apps were developed: (i) an app for extracting relevant Mantid kernels (MDNorm) in C++ and (ii) the Julia MiniVATES.jl miniapp. We present performance results for NVIDIA A100 and AMD MI100 GPUs and AMD EPYC 7513 and 7662 CPUs. The results provide insights for future generations of data reduction software that can embrace performance portability for an integrated research infrastructure across DOE’s experimental and computational facilities.
The Lawrence Berkeley National Laboratory (LBNL) and the National High Magnetic Field Laboratory (NHMFL) have published results on Bi-2212 superconductive magnets realized and tested in the canted cosine-theta and solenoid designs, respectively. Fermilab is now preparing for the assembly of the first Bi-2212 stress-managed cosine-theta insert magnet. The insert will be part of the first hybrid cosine-theta magnet made of Nb$_3$Sn outer layers within the US-MDP effort to reach a 20 T bore field. This paper presents the analytical analysis of the cosine-theta Nb$_3$Sn/Bi-2212 hybrid magnet. We report the parameters, logic, and implementation method of the 2D electromagnetic and mechanical FEM analysis of the LTS/HTS hybrid magnet. Results from a detailed heterogeneous model are compared to the homogeneous model implemented in the past. A Python code has been developed to simulate the current degradation due to stresses in the detail-modeled conductor areas. The current degradation has been introduced in the simulation dynamics for the HTS conductor as an iteration process, updating the input load of Lorentz forces of the energization at each step. The magnetic and mechanical analysis results of the 2D cosine-theta LTS/HTS dipole magnet have been described and analyzed.
Superconducting detectors with sub-eV energy resolution have demonstrated success setting limits on Beyond the Standard Model (BSM) physics due to their unique sensitivity to low-energy events. G4CMP, a Geant4-based extension for condensed matter physics, provides a comprehensive toolkit for modeling phonon and charge dynamics in cryogenic materials. This paper introduces a technical formalism to support the superconducting qubit and low-threshold detector community in implementing phonon simulations in custom materials into the G4CMP. As a case study, we present the results of a detailed analysis of silica phonon transport properties relevant for simulating substrate backgrounds in Beryllium Electron capture in Superconducting Tunnel junctions (BeEST)-style experiments using G4CMP. Additionally, Python-based tools were developed to aid users in implementing their own materials and are available on the G4CMP repository.
Large-scale battery energy storage systems (BESS) have found ever-increasing use across industry and society to accelerate clean energy transition and improve energy supply reliability and resilience. However, their optimal power management poses significant challenges: the underlying high-dimensional nonlinear nonconvex optimization lacks computational tractability in real-world implementation, and the uncertainty of the exogenous power demand makes exact optimization difficult. This paper presents a new solution framework to address these bottlenecks. The solution pivots on introducing power-sharing ratios to specify each cell’s power quota from the output power demand. To find the optimal power-sharing ratios, we formulate a nonlinear model predictive control (NMPC) problem to achieve power-loss-minimizing BESS operation while complying with safety, cell balancing, and power supply-demand constraints. We then propose a parameterized control policy for the power-sharing ratios, which utilizes only three parameters, to reduce the computational demand in solving the NMPC problem. This policy parameterization allows us to translate the NMPC problem into a Bayesian inference problem for the sake of 1) computational tractability, and 2) overcoming the nonconvexity of the optimization problem. We leverage the ensemble Kalman inversion technique to solve the parameter estimation problem. Concurrently, a low-level control loop is developed to seamlessly integrate our proposed approach with the BESS to ensure practical implementation. This low-level controller receives the optimal power-sharing ratios, generates output power references for the cells, and maintains a balance between power supply and demand despite uncertainty in output power. We conduct extensive simulations and experiments on a 20-cell prototype to validate the proposed approach.
Wind turbine active wake mixing (AWM) is an exciting new field of research where dynamic actuation, usually on the blade pitch angles, is used to increase wind farm-wide power production. From a controls perspective, the current state-of-the-art AWM strategies are very simple: a dynamic, usually periodic, reference signal is prescribed to actuators in an open-loop (OL) fashion. The actuation is then presumed to have a certain desired effect on the system, i.e., the wind turbine and the flow it affects. However, this system is highly nonlinear and experiences disturbances in the form of wind variations that are not known a priori. As a result, the OL approach might not yield optimal results. In this article, a novel approach is presented, which closes the loop on AWM controllers. A feedback loop is implemented, which uses measurements of the individual blade bending moments that are widely available on modern wind turbines to determine the individual blade pitch angles. A proof of concept of this implementation is presented in this article, and a thorough comparison with the OL method is executed using high-fidelity flow simulations. These simulations show that the novel closed-loop controller does not substantially increase wake mixing but achieves similar performance as the OL controller while reducing fatigue loads on the controlled turbine.
Taking advantage of the extreme stability of the pulsar period, it can serve as the timing source for grid synchronization to compensate for the timing drift instigated by the loss of GPS signal. Nevertheless, the real-time transmission and processing of the pulsar data suffer from its high-frequency data rate, varying from megahertz to gigahertz, resulting in reduced computing speed and increased time delay. To mitigate this issue, the hardware and software frameworks are implemented for the high-density pulsar data transmission and processing for grid synchronization in this research. Initially, the high-density pulsar data is transferred using open-source software. The complementary duty cycle timing module is designed to coordinate the operation of the dual-channel high-speed interface and software. Subsequently, the multiple-threading is applied to the receiving, parsing, and splicing pulsar data. Next, the pulsar signal extraction method is implemented based on the polyphase filterbank and time of arrival estimation. Ultimately, real-time performance verification experiments are carried out for different components under two hardware platforms. Finally, the results demonstrate that only 0.482 s is required for processing 4 Gigabyte data through multiple-threading, which is 3.8 times faster than the single thread. The pulsar signal extraction can also be executed within 707 ms for 4.8 seconds of data, thereby indicating that real-time requirements can be met.
Volt-VAR control (VVC) methods based on deep reinforcement learning (DRL) can effectively control distribution grid voltage and minimize power loss by implementing corrective and preventive control measures on the reactive power output of inverter-based distributed energy resources (DERs). However, model-free DRL-based VVC approaches usually cannot capture the important topological feature of the power system since they use a fully-connected network (FCN) to deliver the action. Therefore, this paper proposes a graph convolutional network (GCN)-based DRL approach that can employ the topological information of the network to take better control action for regulating the voltage. Our implementation allows for both centralized and decentralized configurations, utilizing a single agent and multiple agents respectively. Although the centralized GCN-based DRL approach has its advantages of minimizing voltage fluctuation and power loss, it is not suitable for large scale power systems due to its challenges in terms of scalability, computation speed and potential single points of failure. Therefore, these problems can be resolved using the decentralized GCN-based DRL approach. Moreover, to ensure the safe operation of the model, our proposed approach incorporates an exponential barrier function while formulating the reward function for each agent. To validate performance of the proposed approaches, the proposed model is tested on modified IEEE test systems and the performances are measured in terms on voltage fluctuation reduction, minimization of power loss and computational speed. Finally, the results show that the proposed topology-aware approach outperforms the FCN-based DRL approach in terms of reducing voltage fluctuation and minimizing power loss of the network. Moreover, it is shown that the decentralized GCN-based DRL has faster computational speed than other approaches.
Designing efficient and noise-tolerant quantum computation protocols generally begins with an understanding of quantum error-correcting codes and their native logical operations. The simplest class of native operations are transversal gates, which are naturally fault-tolerant. Here, in this paper, we aim to characterize the transversal gates of quantum Reed–Muller (RM) codes by exploiting the well-studied properties of their classical counterparts. We start our work by establishing a new geometric characterization of quantum RM codes via the Boolean hypercube and its associated subcube complex. More specifically, a set of stabilizer generators for a quantum RM code can be described via transversal X and Z operators acting on subcubes of particular dimensions. This characterization leads us to define subcube operators composed of single-qubit π/2 k Z -rotations that act on subcubes of given dimensions. We first characterize the action of subcube operators on the code space: depending on the dimension of the subcube, these operators either (1) act as a logical identity on the code space, (2) implement non-trivial logic, or (3) rotate a state away from the code space. Second, and more remarkably, we uncover that the logic implemented by these operators corresponds to circuits of multi-controlled-Z gates that have an explicit and simple combinatorial description. Overall, this suite of results yields a comprehensive understanding of a class of natural transversal operators for quantum RM codes.
On-ramp merging is a critical bottleneck in freeway traffic flow, contributing to congestion, accidents, and excessive fuel consumption. Although traditional ramp metering provides macroscopic control, it lacks the granularity for optimizing an individual vehicle’s trajectory. Cooperative merging, enabled by connected and automated vehicles, can potentially enhance traffic efficiency, safety, and fuel economy. However, existing research often neglects the influence of heterogeneous vehicle dynamics, unreliable vehicle-to-vehicle (V2V) communication, and real-time implementation challenges. Here, this paper introduces novel model-free online speed planners for cooperative on-ramp merging. The planners address these limitations by being agnostic to vehicle dynamics, effectively compensating for V2V communication packet drops and incurring only a light computational burden. Comprehensive evaluation, conducted on a real-time traffic-vehicle-communication co-simulation platform integrating high-fidelity vehicle dynamics, a traffic simulator, and recorded V2V communication footprints, demonstrates the effectiveness of the proposed speed planners. Simulation results reveal that the proposed method yields accurate tracking of desired speed and inter-vehicle distance, maintaining low fuel consumption even under high packet drop ratios, and demonstrating real-time implementation efficiency.
The semantics of HPC storage systems are defined by the consistency models to which they abide. Storage consistency models have been less studied than their counterparts in memory systems, with the exception of the POSIX standard and its strict consistency model. The use of POSIX consistency imposes a performance penalty that becomes more significant as the scale of parallel file systems increases and the access time to storage devices, such as node-local solid storage devices, decreases. While some efforts have been made to adopt relaxed storage consistency models, these models are often defined informally and ambiguously as by-products of a particular implementation. Here in this work, we establish a connection between memory consistency models and storage consistency models and revisit the key design choices of storage consistency models from a high-level perspective. Further, we propose a formal and unified framework for defining storage consistency models and a layered implementation that can be used to easily evaluate their relative performance for different I/O workloads. Finally, we conduct a comprehensive performance comparison of two relaxed consistency models on a range of commonly seen parallel I/O workloads, such as checkpoint/restart of scientific applications and random reads of deep learning applications. We demonstrate that for certain I/O scenarios, a weaker consistency model can significantly improve the I/O performance. For instance, in small random reads that are typically found in deep learning applications, session consistency achieved a 5x improvement in I/O bandwidth compared to commit consistency, even at small scales.
Workflow and serverless frameworks have empowered new approaches to distributed application design by abstracting compute resources. However, their typically limited or one-size-fits-all support for advanced data flow patterns leaves optimization to the application programmer—optimization that becomes more difficult as data become larger. The transparent object proxy, which provides wide-area references that can resolve to data regardless of location, has been demonstrated as an effective low-level building block in such situations. Here we propose three high-level proxy-based programming patterns—distributed futures, streaming, and ownership—that make the power of the proxy pattern usable for more complex and dynamic distributed program structures. We motivate these patterns via careful review of application requirements and describe implementations of each pattern. As a result, we evaluate our implementations through a suite of benchmarks and by applying them in three meaningful scientific applications, in which we demonstrate substantial improvements in runtime, throughput, and memory usage.
Plasma–wall interaction is one of the key research topics on the way to controlled fusion. To study the best operational designs with reduced heat and particle fluxes onto tokamak plasma facing components (PFCs) comprehensive plasma simulations are required. A recent implementation of a hybridized discontinuous Galerkin scheme into a new version of SolEdge code has the advantage of using magnetic equilibrium-free mesh. This allows us to conduct pioneering 2-D transport simulations of a full discharge in the WEST tokamak. In this work, we implemented plasma transport coefficients as functions of coordinate in the poloidal plane and neutral diffusion as a function of neutral mean free path. Moreover, the perpendicular convection flux terms were added to the code. Using the new features, a few test cases were investigated. Finally, the influence of nonconstant transport coefficients on the simulated particle and heat fluxes onto the WEST tokamak PFCs are demonstrated.
On-road wireless charging of electric vehicles (EVs) in-motion could potentially reduce range anxiety or battery size with wide-spread deployment. The planning and implementation of such systems are greatly complicated due to their susceptibility to load variation inherent to traffic flow. Here, this paper proposes a method for derisking the potential for traffic slowdowns by compensating for reduced vehicle speed and investigates how implementation may affect system performance. A load modeling case study is presented at 200kW for a mile of high-speed roadway employing speed-based power regulation with results indicating average power usage and maximum car hosting capability can be reduced by 20% and increased by 30% respectively. An 85kHz power electronics model is developed based on designs and prototypes for an 11kW, 190m airgap static system and a 200kW dynamic wireless track. The simulation is validated in the 11kW experimental prototype and modified for 200kW operation to compare with simulated performance. Sensitivity studies are performed in MATLAB/Simulink to evaluate how parameters influence system performance and confirm the capability to reduce output power and maintain efficiency at 11 and 200kW. The static 11kW experimental system operates at 93.6% efficiency and multiple options exist to reduce power while maintaining efficiency greater than 90%. The capability to dynamically modify power output from WPT coils, in an experimentally validated simulation, enables techniques to significantly mitigate load variability due to reductions in vehicle speed.
This paper presents a novel end-to-end framework for closed-form computation and visualization of critical point uncertainty in 2D uncertain scalar fields. Critical points are fundamental topological descriptors used in the visualization and analysis of scalar fields. The uncertainty inherent in data (e.g., observational and experimental data, approximations in simulations, and compression), however, creates uncertainty regarding critical point positions. Uncertainty in critical point positions, therefore, cannot be ignored, given their impact on downstream data analysis tasks. Here, in this work, we study uncertainty in critical points as a function of uncertainty in data modeled with probability distributions. Although Monte Carlo (MC) sampling techniques have been used in prior studies to quantify critical point uncertainty, they are often expensive and are infrequently used in production-quality visualization software. We, therefore, propose a new end-to-end framework to address these challenges that comprises a threefold contribution. First, we derive the critical point uncertainty in closed form, which is more accurate and efficient than the conventional MC sampling methods. Specifically, we provide the closed-form and semianalytical (a mix of closed-form and MC methods) solutions for parametric (e.g., uniform, Epanechnikov) and nonparametric models (e.g., histograms) with finite support. Second, we accelerate critical point probability computations using a parallel implementation with the VTK-m library, which is platform portable. Finally, we demonstrate the integration of our implementation with the ParaView software system to demonstrate near-real-time results for real datasets.