Hardware Implementation and Validation of a Protection Scheme Based on Local Measurements for Self-Healing Microgrids
Explore the source record for details and available documents.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
While transactive energy, which is defined as an allocation of electricity based on dynamically discovered values or prices, has been extensively studied, its uptake and use has been slow. This report describes a tool, the transactive network template, which should hasten the creation and uptake of transactive energy networks. Some basic principles of transactive energy are familiar from existing wholesale electricity markets. Locational prices are calculated today for zones within bulk electric transmission systems. Locational prices differ while accounting for the locational costs of electricity generation and the losses and constraints incurred when electricity is transmitted from generators and distributed to consumers. A transactive energy network might include these transmission zones. However, current research strives to apply transactive energy also in electricity distribution circuits, buildings, and even for individual generating and consuming devices. At the same time, researchers explore how to apply transactive energy in real time during increasingly shorter time intervals. Automated computational agents become necessary as transactive energy becomes applied to smaller circuit zones and at faster dynamic timescales. A transactive energy network is an example of a multi-agent system. Each zone in the network is represented by its transactive agent, which makes decisions for and acts on behalf of a business entity that is responsible for and manages one of the circuit regions. A transactive energy network is also an example of a decentralized, distributed control system. Control decisions and responsibilities are distributed among the network’s transactive agents. The transactive agents are independent; that is, there typically is no centralized authority or oversight function. Instead, transactive agents exchange transactive signals and thereby negotiate the prices and quantities of electricity that they will exchange. Initially, the circuit regions and responsibilities of transactive agents appear to be very dissimilar. Each circuit region may comprise transmission, distribution, or building-level circuits. Each has a unique position and electrical connectivity within the transactive energy network. Each possesses unique assets that either generate or consume electricity, and these (e.g., renewable energy generator, diesel generator, aggregate utility load, building load, space conditioning, refrigerator, etc.) may further differ in their price flexibility and in their strategies for responding to dynamic electricity prices. Given such diversity, an implementer’s first inclination might be to start from scratch to define all these devices and to engineer their seemingly unique interactions. Given that each implementer’s perspective may be narrow within a transactive energy network, it is unlikely that uniquely engineered systems would interact well. This is where the transactive network template is applicable. The transactive network template is a metamodel that has been developed to guide implementers as they configure their own transactive agent within a network of such agents. The object-oriented design of the transactive network template provides basic code object types that may be used and extended by implementers to represent each of the assets in their circuit region. These objects further facilitate the transactive agent’s necessary computations, which are divided among responsibilities to schedule power usage, balance electric supply and demand, and coordinate the exchange of electricity with the other transactive agents. This report addresses the conceptual transactive network template design. Implementers are directed to more formal design documents and reference implementations. A Python™-based1 reference implementation of the transactive network template has been coded, and three implementations have been configured to represent a national laboratory and two university campuses. Version 2 of the transactive node template generalizes the market class and its methods to facilitate multiple, and more diverse market coordination mechanisms than were facilitated by and demonstrated using Version 1. Version 3 includes new Appendix B, which addresses the designs of methods that would make dynamic prices track approved electricity rates. In the future, the author wishes to make the transactive network template more generally applicable to networks that require more accurate power flow. Development of the transactive network template is jointly funded by the U.S. Department of Energy (DOE) Energy Efficiency and Renewable Energy and the DOE Office of Electricity. In late 2015, one of the first projects to be funded by the DOE Grid Laboratory Modernization Laboratory Consortium was the Clean Energy and Transactive Campus project, led by Pacific Northwest National Laboratory. DOE funds were matched by an investment by the Washington Department of Commerce through its Clean Energy Fund. The transactive network template was developed to guide the implementation of transactive energy networks within this project’s scope.
With the recent diversification of the hardware landscape in the high-performance computing (HPC) community, performance-portability solutions are becoming more and more important. One of the most popular choices is Kokkos, which recently became a Linux Foundation project. Most of its development is supported by the US Department of Energy and the French Alternative Energies and Atomic Energy Commission. Kokkos is implemented as a C++ library with multiple backends to support CPUs as well as various GPU architectures. These backends include OpenMP, CUDA, HIP, and also SCYL. This approach enables users to leverage the preferred vendor toolchain for the respective platform (e.g. CUDA, ROCm, OneAPI). The SYCL backend is used to target Intel GPUs, in particular to support the Aurora exascale supercomputer. However, SYCL itself also offers a large degree of portability, and in fact Kokkos’ CI for SYCL has been running on NVIDIA hardware due to a lack of access to Intel GPUs. In this report, we describe our experience with using Kokkos SYCL backend on AMD GPUs targeting the Frontier supercomputer at Oak Ridge National Laboratory. The two major SYCL implementations are DPC++ and AdaptiveCpp. While the Kokkos SYCL backend has been implemented using the former, the latter was the first implementation to target AMD GPUs. We will discuss the experience with both of these SYCL implementations in terms of functionality and performance. Using Kokkos to evaluate SYCL toolchains has a number of benefits. Kokkos’ use of SYCL is fairly complex, exercising features such as graphs, relocatable device functions, atomics – including for non-arithmetic types, as well as pinned and page migratable memory allocations. Kokkos also needs to implement capabilities such as Kokkos’ hierarchical parallelism that are not a straight-forward mapping to SYCL capabilities. Furthermore, a large number of libraries and applications that represent diverse use cases are implemented in Kokkos, providing readily available test cases for a toolchain evaluation. Preliminary results show that support for AMD GPUs in DPC++ is much less mature than for NVIDIA GPUs or Intel GPUs. While the situation has improved significantly over the last year, we still encounter many runtime failures, dispatching problems, and code generation issues. With AdaptiveCpp the challenges arise even earlier in the evaluation process. Since Kokkos’ SYCL implementation is largely focused on supporting Intel GPUs, we opted to leverage SYCL extensions which are available in DPC++ but not in AdaptiveCpp. Furthermore, AdaptiveCpp appears to be less conformant with the SYCL2020 standard which Kokkos relies on. In some cases, we are able to work around the lack of feature support, in other cases we have to disable certain Kokkos capabilities to evaluate the toolchain. Our evaluation will leverage Kokkos’ unit tests to establish basic functionality and feature completeness. We then use simple benchmarks for components of a CG implementation as a measure of usability and performance of the SYCL toolchains.
The incorporation of remote operations into reactor operations is a topic of high interest among advanced and small modular reactor (A/SMR) vendors, with some considering it essential to the success of their business models. However, remote operations are a concept novel to the nuclear industry. While various technical aspects of remote operations have been explored, a significant gap remains in understanding the security implications of integrating remote operations into reactor designs, particularly concerning the security requirements for remote-operations facilities and infrastructure. This report aims to address this gap by first defining classes of remote operation based on the extent of remote access to reactor control systems and grounded in the existing regulatory framework with compatible terminology. Secondly, the report outlines the physical and cybersecurity requirements applicable to remote-operations facilities and infrastructure at each defined class. These requirements are based on existing licensing frameworks provided by 10 Code of Federal Regulations (CFR) Part 50 and 10 CFR Part 52, as well as the upcoming A/SMR licensing framework in the proposed Part 53. The assessment focuses specifically on security regulations, such as 10 CFR Part 73, which includes provisions for both cybersecurity (§ 73.54) and physical security (§ 73.55). This report proposes five classes of remote reactor operations. Class 1 involves remote monitoring only, with no control over reactor systems. Class 2 allows for the remote issuance of allowlisted commands to the reactor facility. Class 3 extends control to non-safety-significant, non-safety-related, or not important to safety systems and equipment. Class 4 permits remote control of safety-significant systems. Finally, Class 5 allows remote control of safety-related systems. It is important to note that these classes were defined purely with functionality in mind, without considering the practicality or feasibility of implementation for each class under current or upcoming regulatory guidance. The intention behind this approach is to enable an assessment of which security requirements apply to each class, allowing readers to evaluate the implementation possibilities for their specific use cases. Following the definition of remote-operation classes, the report assesses the specific physical and cybersecurity requirements applicable to the remote-operations facility and infrastructure within each defined class. This includes defining the types and locations of operators that are possible at each class of operation and, based on operator type and location, as well as functionality within each class, outlining the physical and cybersecurity requirements. By detailing the security requirements by class, the report provides readers with the information needed to determine the type of security program they may need to implement for their desired concept of operation. The next contribution of this report was to assess the practicality of implementing each proposed class of remote operations based upon the security requirement assessment. In short, three of the five proposed remote-operation classes were found to possibly have a practical path forward to implementation under the U.S. regulatory framework. Class 1 remote operations are currently in use in the U.S. while Class 2 and 3 remote operations may be logistically possible to implement under the U.S. regulatory framework. The final two Classes, 4 and 5, would likely be logistically difficult, if not infeasible to implement within the current U.S. physical- and cybersecurity regulatory framework. Given the results of the feasibility assessment, an example architecture is proposed for both Class 2, remote allowlisted commands, and Class 3, remote control of non-safety systems as well as security implication assessments of each architecture. These example implementations are not meant to be prescriptive in terms of how Class 2 or Class 3 remote operations should be deployed; instead, they are intended to be informative to stakeholders on how Class 2 or Class 3 could potentially be applied in order to inform their system design. An example architecture for Class 1 remote monitoring was not provided as Class 1 in already in use in U.S. nuclear operations. Example architectures for Class 4 and Class 5 were not provided due to their assessment of being likely infeasible to implement. The final contribution is an assessment of the physical- and cybersecurity implications of introducing autonomous operations into an A/SMR. What was found was that the security implications can be separated into two cases. Autonomous operations supported by SSCs located only at the reactor site, and autonomous operations supported by SSCs outside of the reactor site. For the first case, the introduction of autonomous systems will likely not change the facility’s requirement to comply with existing cyber and physical security regulation
Here, this work describes a crystal plasticity formulation combining several mathematical, numerical, and implementation choices to produce a highly efficient model. Specifically, the key choices in the implementation are (1) representing orientations with modified Rodrigues parameters, (2) implementing a fully coupled implicit time integration for the elastic stretch, the crystal orientations, and the model internal variables, (3) implementing the model in the NEML2 constitutive modeling framework, based on PyTorch, to vectorize the calculations and port the computation to GPUs and other hardware accelerators, and (4) an exact implementation of the consistent tangent matrix, even for arbitrary coupling to other field variables beyond the displacements, like temperature, neutron fluence, etc. The first two features of the model are, to our knowledge, novel. The paper considers each of these choices individually as well as the final model as a whole. This includes a full description of modified Rodrigues parameters, their advantages over other representations of orientations, the mathematical formulae and tools required to implement a model with modified Rodrigues parameters, and a detailed description of the geometry of the space of modified Rodrigues parameters (in an appendix). It also includes a description of a fully implicit time integration scheme for the orientations and the advantages in representing orientations with modified Rodrigues parameters in implementing such a model. The work then assess, via numerical examples, the advantages of fully coupled implicit time integration versus more common decoupled and explicit time integration schemes. These studies demonstrate the computational advantages of fully coupled integration versus other time integration algorithms, though the performance of the competing models depends on the complexity of the underlying single crystal model. The study concludes by demonstrating that the choice of time integration method affects the sharpness of the predicted texture, with explicit methods for integrating the orientations overestimating texture sharpness and implicit methods underestimating texture sharpness.
In the HPC area, both hardware and software move quickly. Often new hardware is developed and deployed, the corresponding software stack, including compilers and other tools, are under active development while leading edge software developers are working to port and tune their applications, all at the same time. While the software ecosystem is in flux, one of the key challenges for users is obtaining insight into the state of implementation of key features in the programming languages and models their applications are using – whether they have been implemented, and whether the implementation conforms to the specification, especially for newly implemented features (less tested by widespread use). OpenMP is one of the most prominent shared memory programming models used for on-node programming in HPC. With the shift towards accelerators (such as GPUs and FPGAs) and heterogeneous programming OpenMP features are getting more complex. It is natural to ask whether generative AI approaches, and large language models (LLMs) in particular, can help in producing validation and verification test suites to allow users better and faster insights into the availability and correctness of OpenMP features of interest. In this work, we explore the use of ChatGPT-4 to generate a suite of tests for OpenMP features. We have chosen a set of directives and clauses, a total of 78 combinations, which first appeared in OpenMP 3.0 (released in May 2008) but are also relevant for accelerators. We prompted ChatGPT to generate tests in the C and Fortran languages, for both host (CPU) and device (accelerator). On the Summit super-computer using the GNU implementation, we found that, of the 78 generated tests 67 C tests and 43 Fortran tests compiled successfully and fewer than those executed to completion. On further analysis we show that not all generated tests are valid. We document the process, results, and provide detailed analysis regarding the quality of tests generated. With the aim of providing input to a production quality validation and verification suite, we manually implement the corrections required to make the tests valid according to the current OpenMP specification. We quantify this effort as small, medium, or large, and record the lines of code changed to correct the invalid tests. With the corrected tests we validate recent implementations from HPE, AMD, and GNU on the Frontier supercomputer. Our experiment and subsequent analysis show that although LLMs are capable of producing HPC specific codes, they are limited by their understanding of the deeper semantics and restrictions of programming models such as OpenMP. Unsurprisingly more commonly used features have better support, while some OpenMP 3.0 directives such as sections and tasking are not universally supported on accelerators. We demonstrate that successful compilation and execution to completion are inadequate metrics for evaluating generated code and that, at this time, commodity LLMs require expert intervention for code verification. This points to gaps in the training data that is currently available for HPC. We demonstrate that with "small" effort 37% of generated invalid C tests and 63% of generated invalid Fortran tests could be corrected. This improves productivity of test generation as we circumvent writing from scratch and the common programming errors associated with it.
Genome-scale metabolic models (GSMM) are commonly used to identify gene deletion sets that result in growth coupling and pairing product formation with substrate utilization and can improve strain performance beyond levels typically accessible using traditional strain engineering approaches. However, sustainable feedstocks pose a challenge due to incomplete high-resolution metabolic data for non-canonical carbon sources required to curate GSMM and identify implementable designs. Here we address a four-gene deletion design in the Pseudomonas putida KT2440 strain for the lignin-derived non-sugar carbon source, p-coumarate (p-CA), that proved challenging to implement. We examine the performance of the fully implemented design for p-coumarate to glutamine, a useful biomanufacturing intermediate. In this study glutamine is then converted to indigoidine, an alternative sustainable pigment and a model heterologous product that is commonly used to colorimetrically quantify glutamine concentration. Through proteomics, promoter-variation, and growth characterization of a fully implemented gene deletion design, we provide evidence that aromatic catabolism in the completed design is rate-limited by fumarase hydratase (FUM) enzyme activity in the citrate cycle and requires careful optimization of another fumarate hydratase protein (PP_0897) expression to achieve growth and production. A double sensitivity analysis also confirmed a strict requirement for fumarate hydratase activity in the strain where all genes in the growth coupling design have been implemented. Metabolic cross-feeding experiments were used to examine the impact of complete removal of the fumarase hydratase reaction and revealed an unanticipated nutrient requirement, suggesting additional functions for this enzyme. While a complete implementation of the design was achieved, this study highlights the challenge of completely inactivating metabolic reactions encoded by under-characterized proteins, especially in the context of multi-gene edits.
DUNE is an international experiment dedicated to addressing some of the questions at the forefront of particle physics and astrophysics, including the mystifying preponderance of matter over antimatter in the early universe. The dual-site experiment will employ an intense neutrino beam focused on a near and a far detector as it aims to determine the neutrino mass hierarchy and to make high-precision measurements of the PMNS matrix parameters, including the CP-violating phase. It will also stand ready to observe supernova neutrino bursts, and seeks to observe nucleon decay as a signature of a grand unified theory underlying the standard model. The DUNE far detector implements liquid argon time-projection chamber (LArTPC) technology, and combines the many tens-of-kiloton fiducial mass necessary for rare event searches with the sub-centimeter spatial resolution required to image those events with high precision. The addition of a photon detection system enhances physics capabilities for all DUNE physics drivers and opens prospects for further physics explorations. Given its size, the far detector will be implemented as a set of modules, with LArTPC designs that differ from one another as newer technologies arise. In the vertical drift LArTPC design, a horizontal cathode bisects the detector, creating two stacked drift volumes in which ionization charges drift towards anodes at either the top or bottom. The anodes are composed of perforated PCB layers with conductive strips, enabling reconstruction in 3D. Light-trap-style photon detection modules are placed both on the cryostat's side walls and on the central cathode where they are optically powered. This Technical Design Report describes in detail the technical implementations of each subsystem of this LArTPC that, together with the other far detector modules and the near detector, will enable DUNE to achieve its physics goals.
As the design complexity of modern accelerators grows, there is more interest in using advanced simulations that have fast execution time or yield additional insights like gradients. The FAST/IOTA facility has been working on implementing and experimentally validating an end-to-end digital twin that is both fast and gradient-aware, allowing for rapid prototyping of new software and experiments with minimal beam time costs. Our framework integrates physics and ML codes for linac and ring simulation through a set of generic interfaces between surrogate and physics-based sections. To reproduce device inputs and outputs, system state is exposed as a deterministic event loop in a specialized discrete event simulator architecture. Because Fermilab is undergoing control system transition, several APIs were implemented as final user interfaces - a fully asynchronous EPICS soft IOC, a gRPC-based Data Pool Manager (DPM), and legacy ACNET protocols. We discuss implementation details as well as challenges handling live data assimilation and future plans to extend modelling to main complex proton accelerators like PIPII and Booster.
The Pantex Plant Ogallala Aquifer and Perched Groundwater Contingency Plan has been developed in accordance with the requirements identified in the: • Interagency Agreement for the Pantex Superfund Site, Article 8.5 Work to be Performed, • Compliance Plan Provision of Hazardous Waste Permit No. 50284, and • Record of Decision for Groundwater, Soil, and Associated Media, Pantex Plant. A Long‐Term Monitoring System Design has been designed to monitor conditions in the perched groundwater including changes in the perched aquifer as a result of implementing the response actions. Monitoring is required for verifying the effectiveness of perched groundwater response actions (i.e., conditions in the perched aquifer are being affected as intended) and for confirming that the perched aquifer and Ogallala Aquifer characterization as defined in the Resource Conservation and Recovery Act Facility Investigation Report and the Corrective Measure Studies/Feasibility Study remains accurate. If monitoring results obtained through the monitoring network identify an unexpected condition or deviation, contingent actions will be considered and implemented as necessary to ensure continued protection of the Ogallala Aquifer and human health and the environment. Potential deviations to expected technology performance may be encountered for each of the four primary response actions that compose the selected remedy for perched groundwater; Playa 1 Pump and Treat System, Southeast Area Pump and Treat System, Southeast Area In‐Situ Bioremediation System (comprised of the Southeast In‐Situ Bioremediation System Original System, Southeast Area In‐Situ Bioremediation System Extension System, Offsite In‐Situ Bioremediation System, Perchlorate/Chromium ISB, Northeast ISB and County Road 8 ISB), and Zone 11 In‐Situ Bioremediation System. Monitoring will also be conducted to determine if there are deviations to the expected characterization, e.g., contaminants not expected as a result of the RCRA Facility Investigation characterization. Deviations to expected conditions in the Ogallala Aquifer could also be encountered if the response actions in the perched groundwater are not performing as expected, i.e., preventing contaminants from migrating to the Ogallala Aquifer. Currently, Pantex has begun investigation of detections of high explosives above groundwater protection standards in wells on the Texas Tech University property and a plume that is moving to the northeast from that area. Due to those detections, this Plan recognizes the fact that future detections in the Ogallala will be focused on first‐ time detections of analytes. After a remedy is determined, this Plan will require modification to address-deviations and contingent actions. This Plan was developed to identify the contingent actions necessary to mitigate impacts resulting from deviations to site conditions or response action performance. The Plan defines the environmental problem being addressed by the response actions, clarifies the expected conditions and objectives of the response actions, and identifies the potential deviations to the response actions (due to site conditions or technology performance) that could be encountered. The deviations were evaluated to determine the likelihood of occurrence, potential impact, and time to respond to avoid impact. The Plan also identifies the monitoring outlined in the Long‐Term Monitoring System Design Report (Consolidated Nuclear Security, 2024) and Sampling Analysis Plan (PanTeXas Deterrence, 2024) that will be used to detect the deviations. Lastly, the Plan specifies the contingent actions that could be implemented in response to the deviations. Because each response focuses on a discrete portion of the perched aquifer and contaminant plume, each response action has a different set of expected conditions, and therefore differing impacts from deviations to the site and technology expectations. As a result, the contingent actions are identified for each response action and potential deviation including specific constituents, location, and conditions. If deviations are encountered that impact the ability of the response action to meet performance objectives, the contingent actions will be focused on ensuring the response action can meet the performance objective. Contingent actions may be implemented as interim actions (ISMs/removal actions) in accordance with the Record of Decision, Interagency Agreement, and Hazardous Waste Permit‐50284, if warranted by the specific circumstances. For deviations to site characterization expected conditions, the contingent action will focus on determination of the source of the deviation, determination of the appropriate response, and evaluation of additional work to be completed. However, if the deviation to characterization impacts the performance of the response action, the contingent action will again focus on ensuring performance objectives can be met. Early source term removals and cleanup actions have been implemented to protect the Ogallala Aquifer. Because of these actions and based on modeling results, the expected conditions in the Ogallala Aquifer are that constituents of concern will not be detected above the Groundwater Protection Standards (GWPSs) nor will they reach potential points of exposure above the GWPS. The primary deviation of concern for the Ogallala is if constituents are detected in the Ogallala Aquifer near or above GWPSs. If it occurs, this change in expected conditions would require further evaluation of site and contaminant characteristics to determine an appropriate course of action. The evaluation would include additional monitoring, source identification, implementation of interim protective measures (if necessary), and delineation of extent. These evaluations are necessary to determine an appropriate response action for the Ogallala. The primary goal of the Plan is to provide for the continued protection of the Ogallala Aquifer and the health of its consumers. In recognition, this Plan presents a flexible and rational approach for making future decisions associated with confirming the change in perched and Ogallala aquifer conditions and identifying a response (technical activities, changes to response actions, regulatory oversight, and public involvement).
We present the implementation of a specialized version of our previously published unified embedding model, SpeckleNN, for real-time speckle pattern classification in X-ray Single-Particle Imaging (SPI), using the SLAC Neural Network Library (SNL) on an FPGA platform. This hardware realization transitions SpeckleNN from a prototypic model into a practical edge solution, optimized for running inference near the detector in high-throughput X-ray free-electron laser (XFEL) facilities, such as those found at the Linac Coherent Light Source (LCLS). To address the resource constraints inherent in FPGAs, we developed a more specialized version of SpeckleNN. The original model, which was designed for broader classification across multiple biological samples, comprised ~5.6 million parameters. The new implementation, while reducing the parameter count to 64.6K (a 98.8% reduction), focuses on maintaining the model's essential functionality for real-time operation, achieving an accuracy of 90%. Furthermore, we compressed the latent space from 128 to 50 dimensions. This implementation was demonstrated on the KCU1500 FPGA board, utilizing 71% of available DSPs, 75% of LUTs, and 48% of FFs, with an average power consumption of 9.4W according to the Vivado post-implementation report. The FPGA performed inference on a single image with a latency of 45.015 microseconds at a 200 MHz clock rate. In comparison, running the same inference on an NVIDIA A100 GPU resulted in an average power consumption of ~73W and an image processing latency of around 400 microseconds. Our FPGA-accelerated version of SpeckleNN demonstrated significant improvements, achieving an 8.9 × speedup and a 7.8 × reduction in power consumption compared to the GPU implementation. Key advancements include model specialization and dynamic weight loading through SNL, which eliminates the need for time-consuming FPGA design re-synthesis, allowing fast and continuous deployment of models (re)trained online. These innovations enable real-time adaptive classification and efficient vetoing of speckle patterns, making SpeckleNN more suited for deployment in XFEL facilities. This implementation has the potential to significantly accelerate SPI experiments and enhance adaptability to evolving experimental conditions.
A key step in Neural Networks is activation. Among the different types of activation functions, sigmoid, tanh, and others involve the usage of exponents for calculation. From a hardware perspective, exponential implementation implies the usage of Taylor series or repeated methods involving many addition, multiplication, and division steps, and as a result are power-hungry and consume many clock cycles. We implement a piecewise linear approximation of the sigmoid function as a replacement for standard sigmoid activation libraries. This approach provides a practical alternative by leveraging piecewise segmentation, which simplifies hardware implementation and improves computational efficiency. In this paper, we detail piecewise functions that can be implemented using linear approximations and their implications for overall model accuracy and performance gain. Our results show that for the DenseNet, ResNet, and GoogLeNet architectures, the piecewise linear approximation of the sigmoid function provides faster execution times compared to the standard TensorFlow sigmoid implementation while maintaining comparable accuracy. Specifically, for MNIST with DenseNet, accuracy reaches 99.91% (Piecewise) vs. 99.97% (Base) with up to 1.31x speedup in execution time. For CIFAR-10 with DenseNet, accuracy improves to 98.97% (Piecewise) vs. 99.40% (Base) while achieving 1.24x faster execution. Similarly, for CIFAR-100 with DenseNet, the accuracy is 97.93% (Piecewise) vs. 98.39% (Base), with a 1.18x execution time reduction. These results confirm the proposed method’s capability to efficiently process large-scale datasets and computationally demanding tasks, offering a practical means to accelerate deep learning models, including LSTMs, without compromising accuracy.
Trapped ions are a promising modality for quantum systems, with demonstrated utility as the basis for quantum processors and optical clocks. However, traditional trapped-ion systems are implemented using complex free-space optical configurations, whose large size and susceptibility to vibrations and drift inhibit scaling to large numbers of qubits. In recent years, integrated-photonics-based systems have been demonstrated as an avenue to address the challenge of scaling trapped-ion systems while maintaining high fidelities. While these previous demonstrations have implemented both Doppler and resolved-sideband cooling of trapped ions, these cooling techniques are fundamentally limited in efficiency. In contrast, polarization-gradient cooling can enable faster and more power-efficient cooling and, therefore, improved computational efficiencies in trapped-ion systems. While free-space implementations of polarization-gradient cooling have demonstrated advantages over other cooling mechanisms, polarization-gradient cooling has never previously been implemented using integrated photonics. In this paper, we design and experimentally demonstrate key polarization-diverse integrated-photonics devices and utilize them to implement a variety of integrated-photonics-based polarization-gradient-cooling systems, culminating in the first experimental demonstration of polarization-gradient cooling of a trapped ion by an integrated-photonics-based system. By demonstrating polarization-gradient cooling using an integrated-photonics-based system and, in general, opening up the field of polarization-diverse integrated-photonics-based devices and systems for trapped ions, this work facilitates new capabilities for integrated-photonics-based trapped-ion platforms.
The sensitivity analysis algorithms that have been developed by the radiation transport community in multiple neutron transport codes, such as MCNP and SCALE, are extensively used by fields such as the nuclear criticality community. However, these techniques have seldom been considered for electron transport applications. In the past, the differential-operator method with the single scatter capability has been implemented in Sandia National Laboratories’ Integrated TIGER Series (ITS) coupled electron-photon transport code. This work is meant to extend the available sensitivity estimation techniques in ITS by implementing an adjoint-based sensitivity method, GEAR-MC, to strengthen its sensitivity analysis capabilities. To ensure the accuracy of this method being extended to coupled electron-photon transport, it is compared against the central-difference and differential-operator methodologies to estimate sensitivity coefficients for an experiment performed by McLaughlin and Hussman. Energy deposition sensitivities were calculated using all three methods, and the comparison between them has provided confidence in the accuracy of the newly implemented method. Unlike the current implementation of the differential-operator method in ITS, the GEAR-MC method was implemented with the option to calculate the energy-dependent energy deposition sensitivities, which are the sensitivity coefficients for energy deposition tallies to energy-dependent cross sections. The energy-dependent cross sections could be the cross sections for the material, elements in the material, or reactions of interest for the element. Further, these sensitivities were compared to the energy-integrated sensitivity coefficients and exhibited a maximum percentage difference of 2.15%.
SPADES (Solver for PArallel Discrete Event Simulation) is an open-source parallel discrete event simulation (PDES) package built on the AMReX library. Targeted at solving discrete event systems in parallel, this software package aims to be performance portable and scalable on heterogeneous computing architectures, e.g., graphic processing units (GPU). SPADES implements optimistic synchronization with rollback through an implementation of the Time Warp algorithm. An alternative conservative synchronization approach is also implemented using the Lower Bound on Incoming Time Stamp. In our implementation, logical processes are represented as cells in a grid and event messages are represented as particles. SPADES supports various parallel decomposition strategies, including the use of the Message Passing Interface (MPI) and OpenMP threading. All major GPU architectures (e.g., Intel, AMD, NVIDIA) are supported through the use of performance portability functionalities implemented in AMReX. The SPADES software is released in NREL Software Record SWR-24-99 “SPADES (Scalable Parallel Discrete Events Simulation)”.