Search NASASearch

SEARCH · Search NASA

Results for “asynchronous methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Modifying the Asynchronous Jacobi Method for Data Corruption Resilience

Moving scientific computation from high-performance computing (HPC) and cloud computing (CC) environments to devices on the edge, i.e., physically near instruments of interest, has received tremendous interest in recent years. Such edge computing environments can operate on data in situ, offering enticing benefits over data aggregation to HPC and CC facilities that include avoiding costs of transmission, increased data privacy, and real-time data analysis. Because of the inherent unreliability of edge computing environments, new fault-tolerant approaches must be developed before the benefits of edge computing can be realized. Motivated by algorithm-based fault tolerance, a variant of the asynchronous Jacobi (ASJ) method is developed that achieves resilience to data corruption by rejecting solution approximations from neighbor devices according to a bound derived from convergence theory. Numerical results on a two-dimensional Poisson problem show that the new rejection criterion, along with a novel approximation to the shortest path length on which the criterion depends, restores convergence for the ASJ variant in the presence of certain types data corruption. Numerical results are obtained for when the singular values in the analytic bound are approximated. Additional linear systems are also explored, one with a more dense sparsity pattern and one that includes advection. All results indicate that successful resilience to data corruption depends on whether the bound tightens fast enough to reject corrupted data before the iteration evolution deviates significantly from that predicted by the convergence theory defining the bound. This observation generalizes to future work on algorithm-based fault tolerance for other asynchronous algorithms, including upcoming approaches that leverage Krylov subspaces.

97 MATHEMATICS AND COMPUTING

Calibration-free analysis of Li isotope ratios using laser ablation and laser absorption spectroscopy

We introduce a rapid, calibration-free, all-optical method for high-precision lithium isotope ratio measurements in solid materials using laser ablation combined with tunable laser absorption spectroscopy. A new asynchronous acquisition method is used to acquire time-resolved, high-resolution spectra of the 6Li and 7Li D1 and D2 transitions near 671 nm. Isotope ratios and atomic column densities are extracted from measured spectra via a physics-based fitting model including hyperfine structure. Under 1 Torr air, spectra recorded = 0.75 ms after plasma onset exhibit narrow linewidths corresponding to Doppler temperatures = 400 K, enabling resolution of the isotope peaks with high signal-to-noise ratios. Analysis of LiAlO2 samples with varying 6Li:7Li ratios demonstrates isotopic precisions of 0.6–1.7% for spectra acquired in 30 s. Isotope ratios determined from the spectral fits show accuracy within –0.3% to –1.7% of reference ICP-MS measurements without requiring calibration to external standards. By eliminating sample preparation and enabling spatially resolved isotopic mapping, this method offers a rapid analysis approach to lithium isotope determination in solid materials relevant to nuclear energy, safeguards, and geochemistry.

Phillips, Mark C.

Beam Synchronous for the Rest of Us!

Fermilab’s Tevatron Clock (TCLK) infrastructure has been an integral part of the accelerator control network since the 1980’s. This 10MHz Manchester encoded protocol has enabled flexible, real-time event distribution for thousands of devices connected to the timing network with a high degree of reliability. Forthcoming upgrades to the Fermilab complex (PIP-II, LBNF, ACORN) necessitate higher levels of precision to maintain inter-bunch timing for Instrumentation and Control purposes. This presents as an opportunity to refine the event distribution protocol for tighter synchronization between machines, experiments, and eventually far-site operations. This paper outlines a method by which beam-synchronous events may be distributed through asynchronous serial protocols via integration with local LLRF and global PPS reference signals. This method is ideal for synchrotron machines with aggressive frequency sweeps (such as Fermilab's 38~53MHz Booster) and allows for precision timing to be maintained across machines without specialized hardware.

43 PARTICLE ACCELERATORS

Stage-local partitioned two-step runge-kutta methods for large systems of ordinary differential equations

We introduce stage-local partitioned two-step Runge-Kutta methods are an extension of standard two-step Runge-Kutta methods, which are an alternative to the standard additive two-step Runge-Kutta methods currently existing in the literature. Furthermore, these new schemes are designed with an eye towards truly N-partitioned systems and leverage local stage approximations to make several computationally interesting approximations viable. Specifically, the focus on local stage approximations makes possible the construction of truly asynchronous schemes, in the parallel sense, possible. In addition, we show that an implicit-explicit approach to these schemes can lead to methods that require the inversion of only local nonlinear systems.

Applied Dynamical Systems

Coincident learning for beam-based rf station fault identification using phase information at the SLAC linac coherent light source

Anomalies in radio-frequency (rf) stations can result in unplanned downtime and performance degradation in linear accelerators such as SLAC’s Linac Coherent Light Source (LCLS). Detecting these anomalies is challenging due to the complexity of accelerator systems, high data volume, and scarcity of labeled fault data. Prior work identified faults using beam-based detection, combining rf amplitude and beam position monitor data. Due to the simplicity of the rf amplitude data, classical methods are sufficient to identify faults, but the recall is constrained by the low-frequency and asynchronous characteristics of the data. In this work, we leverage high-frequency, time-synchronous rf phase data to enhance anomaly detection in the LCLS accelerator. Due to the complexity of phase data, classical methods fail, and we instead train deep neural networks within the Coincident Anomaly Detection (CoAD) framework. We find that applying CoAD to phase data detects nearly 3 times as many anomalies as when applied to amplitude data, while achieving broader coverage across rf stations. Furthermore, the rich structure of phase data enables us to cluster anomalies into distinct physical categories. Through the integration of auxiliary system status bits, we link clusters to specific fault signatures, providing additional granularity for uncovering the root cause of faults. We also investigate interpretability via Shapley values, confirming that the learned models focus on the most informative regions of the data and providing insight for cases where the model makes mistakes. This work demonstrates that phase-based anomaly detection for rf stations improves both diagnostic coverage and root cause analysis in accelerator systems and that deep neural networks are essential for effective analysis.

Accelerator Physics (physics.acc-ph)

OpenSn: A massively parallel, open-source simulation environment for discrete ordinates radiation transport

OpenSn is an open-source, massively parallel deterministic radiation transport code for solving the discrete-ordinates ( S N ) form of the Boltzmann transport equation on unstructured, arbitrary polyhedral meshes. It supports high-fidelity simulations involving steady-state, eigenvalue, and adjoint problems for neutral particles (e.g., neutrons, photons, multi-particles), using the multigroup approximation in energy. OpenSn combines angular discretization via discrete ordinates with a discontinuous Galerkin finite element method (DGFEM) in space, enabling accurate resolution of transport physics on arbitrary polyhedral cells, included locally refined spatial grids. It includes multiple angular quadrature types, including locally refined angular quadratures. Written in modern C++ with a Python API, OpenSn runs efficiently on platforms ranging from laptops to supercomputers. The transport sweep algorithm is implemented using a task-based, directed-acyclic-graph (DAG) approach for each angle and supports asynchronous parallelism across thousands of MPI ranks. Group-set aggregation improves compute intensity, and synthetic acceleration techniques (e.g., diffusion synthetic acceleration, second-moment method) enhance solver convergence. OpenSn has been verified on reactor physics problems and demonstrated excellent weak and strong scaling performance on more than 32,768 processes, making it a versatile and robust platform for large-scale transport simulations in complex geometries.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Breaking the Million-Electron and 1 EFLOP/s Barriers: Biomolecular-Scale Ab Initio Molecular Dynamics Using MP2 Potentials

The accurate simulation of complex biochemical phenomena has historically been hampered by the computational requirements of high-fidelity molecular-modeling techniques. Quantum mechanical methods, such as ab initio wave-function (WF) theory, deliver the desired accuracy, but have impractical scaling for modeling biosystems with thousands of atoms. Combining molecular fragmentation with MP2 perturbation theory, this study presents an innovative approach that enables biomolecular-scale ab initio molecular dynamics (AIMD) simulations at WF theory level. Leveraging the resolution-of-the-identity approximation for Hartree-Fock and MP2 gradients, our approach eliminates computationally intensive four-center integrals and their gradients, while achieving near-peak performance on modern GPU architectures. The introduction of asynchronous time steps minimizes time step latency, overlapping computational phases and effectively mitigating load imbalances. Utilizing up to 9,400 nodes of Frontier and achieving 59% (1006.7 PFLOP/s) of its double-precision floating-point peak, our method enables us to break the million-electron and 1EFLOP/s barriers for AIMD simulations with quantum accuracy.

Kurzak, Jakub

SAGIPS: a physics-inspired scalable asynchronous generative inverse-problem solver

Abstract Solving large-scale inverse problems using deep-learning algorithms have become an essential part of modern research and industrial applications. The complexity of the underlying inverse problem may require the utilization of high performance computing systems which poses a challenge on the algorithmic design of the inverse problem solver. Most deep learning algorithms require, due to their design, custom parallelization techniques in order to be resource efficient while showing a reasonable convergence. In this paper we introduce a S calable A synchronous G enerative I nverse P roblem S olver (SAGIPS) on high-performance computing systems. We present a workflow that utilizes an asynchronous ring-allreduce algorithm to transfer the gradients of the generator network across multiple GPUs. Experiments with a scientific proxy application demonstrate that SAGIPS shows near linear weak scaling, together with a convergence quality that is comparable to traditional methods. The approach presented here allows leveraging Generative Adverserial Network across multiple GPUs, promising advancements in solving complex inverse problems at scale.

97 MATHEMATICS AND COMPUTING

Customized Bayesian optimization for efficient beam tuning at the facility for rare isotope beams

Bayesian optimization (BO) has recently emerged as a powerful approach for on-line beam tuning, and it is rapidly gaining adoption across accelerator facilities due to its flexibility and efficiency in handling complex optimization tasks. At the Facility for Rare Isotope Beams, rapid and reliable tuning is essential to support the delivery of diverse ion species. To improve the practicality of BO in this setting, we implemented several enhancements, including scalarized composite objective construction for multicriteria optimization, asynchronous evaluation for better resource utilization, prior-mean-assisted optimization to accelerate convergence, and a local search strategy for rapid completion of the task. We present the details of these methods, discuss challenges-encountered, and share our experience applying them to specific beam-tuning tasks.

Accelerators & storage rings

SVM-Based Synchronized Fault Detection for 100% Renewable Microgrids

Traditional protection schemes face significant challenges when applied to microgrids with high penetrations of renewables with inverter-based resources (IBRs). The proliferation of advanced sensing and communication technologies has generated copious data, offering an opportunity to overcome these limitations using data-driven machine learning approaches. This work proposes a novel approach based on a support vector machine (SVM) for detecting faults within a 100% renewable microgrid. The approach encompasses a systematic offline training stage for the development of a linear SVM-based fault detection algorithm. This process covers offline data collection from the microgrid under study, the extraction of features such as positive- and negative-sequence components and the total harmonic distortion of the voltage and current measurements of the relays, and the design of the linear SVM-based classifier. During the online implementation, however, different classifiers can exhibit asynchronicity in detecting the fault inception at different subcycle-to-cycle period-level delays. To circumvent this asynchronicity issue, a separate algorithm is developed for each relay to estimate the fault inception time as close to the real fault time. The performance of the proposed SVM-based synchronized fault detection method is evaluated using online time-domain simulation studies on a microgrid test system. The results corroborate the reliability of the fault detection scheme when tested under various fault cases (fault types, locations, and impedances) and non-fault cases during both grid-tied and islanded operation modes.

100% microgrid

SVM-Based Synchronized Fault Detection for 100% Renewable Microgrids: Preprint

Traditional protection schemes face significant challenges when applied to microgrids with high penetrations of renewables with inverter-based resources (IBRs). The proliferation of advanced sensing and communication technologies has generated copious data, offering an opportunity to overcome these limitations using data-driven machine learning approaches. This work proposes a novel approach based on a support vector machine (SVM) for detecting faults within a 100% renewable microgrid. The approach encompasses a systematic offline training stage for the development of a linear SVM-based fault detection algorithm. This process covers offline data collection from the microgrid under study, the extraction of features such as positive- and negative-sequence components and the total harmonic distortion of the voltage and current measurements of the relays, and the design of the linear SVM-based classifier. During the online implementation, however, different classifiers can exhibit asynchronicity in detecting the fault inception at different subcycle-to-cycle period-level delays. To circumvent this asynchronicity issue, a separate algorithm is developed for each relay to estimate the fault inception time as close to the real fault time. The performance of the proposed SVM-based synchronized fault detection method is evaluated using online time-domain simulation studies on a microgrid test system. The results corroborate the reliability of the fault detection scheme when tested under various fault cases (fault types, locations, and impedances) and non-fault cases during both grid-tied and islanded operation modes.

100% microgrid

Electrode strain dynamics in layered intercalation battery cathodes

Rechargeable batteries using electrodes based on intercalation chemistry exhibit notable cyclability, yet their performance still suffers from chemomechanical degradation. In this study, by combining a suite of operando microscopy methods, we explored electrode strain evolution and observed intricate particle cluster rearrangement under electrochemical stimuli. We show that early-stage strain accumulation in intercalation cathodes occurs during the period of interparticle charge transfer and redox reactions stemming from asynchronous coupling and decoupling between chemical (de)intercalation and physical grain motion. This interplay drives heterogeneous redox activity, localized charge equilibration, and multiscale strain cascades that propagate through an asynchronous network of chemical-mechanical interactions. Together, these findings reveal how collective particle dynamics and hierarchical strain transmission dictate electrode deformation and degradation in intercalation cathodes.

25 ENERGY STORAGE

An Educational Program on Concentrated Solar Power and Heliostats for Power Generation and Industrial Processes

The objective of this project was to design and implement a comprehensive educational and applied research program in Concentrated Solar Thermal Power (CSTP) and heliostat technologies at Northeastern University. In alignment with the U.S. Department of Energy's Heliostat Consortium (HelioCon) goals, the project aimed to expand student and public understanding of CSTP systems while simultaneously contributing to workforce development and the broader decarbonization strategy. A particular emphasis was placed on integrating hands-on student design projects and publicly disseminating educational content relevant to CSTP systems. The project addressed a critical gap in renewable energy education: CSTP and heliostats, despite their importance in utility-scale solar energy, are rarely included in standard mechanical engineering programs. This project established new pathways for students to engage with the topic through the creation of a 4-credit graduate/senior elective course, development of five industry-facing short courses, and the inclusion of CSTP-based capstone design projects. Over two academic years, 36 students across six senior design teams developed and tested technologies such as deformable heliostats, beacon-based tracking systems, and solar-powered pyrolizers for biomass-to-biochar conversion. Concurrently, 30 undergraduate and graduate students were enrolled in the new academic course centered around CSTP principles. To ensure the relevance and accessibility of the short course content, the project team engaged with industry professionals, technical policy stakeholders, and potential course participants through structured surveys and informal consultations. Feedback from 28 respondents guided the structure, length, and delivery format of the courses - resulting in a modular design broken into five workshops. The feedback emphasized the need for flexible, asynchronous delivery and practical case studies, particularly in areas such as heliostat control, thermal storage, and solar fuel production. This engagement helped align the courses with the evolving knowledge demands of the renewable energy workforce and ensured that participants from both technical and policy backgrounds could meaningfully benefit from the material. The research and educational activities advanced the understanding of heliostat control systems, optical performance under misalignment, and thermal system integration in solar-driven pyrolysis applications. Methods and designs explored in this project proved to be both technically effective and economically feasible at the lab scale. Prototypes were constructed using commercially available components and custom-fabricated elements, demonstrating that meaningful performance improvements can be achieved with modest material and fabrication costs, supporting the feasibility of student-led research in this field. The public benefit of this project is twofold. First, it cultivates a pipeline of engineers trained to be familiar with CSTP principles, an essential workforce need identified by the Department of Energy for achieving its 2030 cost and deployment targets. Second, it contributes openly accessible educational materials, course content, and experimental frameworks to the broader community, enabling other institutions to adopt or adapt similar programming. Through outreach activities, curriculum integration, and technical exposure, this project contributes to a more informed and capable renewable energy workforce while supporting innovation in heliostat and CSTP system design. A new technical report is being prepared to document the development of the course and its outcomes, with plans to publish it in the ASME Open Access Journal of Engineering to ensure global accessibility, free of cost.

14 SOLAR ENERGY

Radiation Imaging with Event Camera

Neuromorphic or event-based imaging is a new, commercially available sensor technology inspired by how the human eye works. Instead of measuring frames at a fixed rate, the camera measures changes in pixel intensity asynchronously. This difference in readout architecture results in a high dynamic range and low latency. Event-based cameras have been used in a variety of applications, including object tracking, navigation, and lidar technologies. However, event-based cameras have not been adequately researched for their ability to image high-energy particles. This report explores the use of an event camera for imaging alpha, beta, and X-rays particles, when coupled with scintillator screens to convert high-energy particles into visible light. Methods to process event data were developed and are presented here, along with the results. The event camera can measure alpha and beta particles with comparable performance to that of a conventional camera. Event cameras can also image higher-activity sources and offer the possibility of discriminating particle interaction types on the basis of timing differences, which typical cameras cannot do. Additionally, event cameras can image objects with an X-ray source when the source strength dynamically changes but does not create a high-contrast image during static X-ray measurement.

47 OTHER INSTRUMENTATION

Myna: Connecting powder bed fusion build data to simulation tools for digital twin applications

Additive manufacturing (AM), as a digital process, can generate a detailed digital thread linking a part’s design and manufacturing to its operational performance. As AM systems advance, an increasing amount of process data is stored in manufacturing databases. In principle, this data can be utilized by simulation-based digital twin approaches, such as real-time process control and asynchronous post-processing guidance. However, few tools currently exist for systematically integrating digital thread data with computational tools. Here, in this study, we propose a software package, called Myna, for connecting data from powder bed fusion processes to simulation tools. The utility of such a platform is demonstrated using build data from the Oak Ridge National Laboratory Manufacturing Demonstration Facility “Peregrine v2023-10” public dataset to automatically configure and run 54 semi-analytical 3DThesis melt pool simulations, 78 numerical Additive FOAM melt pool simulations, and 3 ExaCA microstructure simulations. The simulated, spatially registered microstructures are then compared directly with electron backscatter diffraction characterization of the corresponding as-built part locations. The resulting simulated microstructure showed variation as a function of process parameters, particularly stripe width; however, the experimental data had little variation between the microstructure texture and grain size resulting from different processing conditions. Analysis of the discrepancies suggest that it is possible a two-phase ferritic-austenitic solidification model is needed to accurately predict grain size and texture for certain stainless steel 316L feedstock compositions under powder bed fusion conditions, providing direction for future research. As illustrated here, due to the number and complexity of the simulations involved in AM process-structure–property predictions, automated methods to connect process data and simulations will remain necessary tools for testing hypotheses and implementing digital twin applications.

Knapp, Gerald L. [Oak Ridge National Laboratory (O

Robustness of Deep Learning Classification to Adversarial Input on GPUs: Asynchronous Parallel Accumulation Is a Source of Vulnerability

The ability of machine learning (ML) classification models to resist small, targeted input perturbations—known as adversarial attacks—is a key measure of their safety and reliability. We show that floating-point non associativity (FPNA) coupled with asynchronous parallel programming on GPUs is sufficient to result in misclassification, without any perturbation to the input. Additionally, we show that this misclassification is particularly significant for inputs close to the decision boundary and that standard adversarial robustness results may be overestimated up to 4.6 when not considering machine-level details. We first study a linear classifier, before focusing on standard Graph Neural Network (GNN) architectures and datasets used in robustness assessments. We develop a novel black-box attack using Bayesian optimization to discover external workloads that can change the instruction scheduling which bias the output of reductions on GPUs and reliably lead to misclassification. Motivated by these results, we present a new learnable permutation (LP) gradient-based approach to learning floating-point operation orderings that lead to misclassifications. The LP approach provides a worst-case estimate in a computationally efficient manner, avoiding the need to run identical experiments tens of thousands of times over a potentially large set of possible GPU states or architectures. Finally, using instrumentation-based testing, we investigate parallel reduction ordering across different GPU architectures under external background workloads, when utilizing multi-GPU virtualization, and when applying power capping. Our results demonstrate that parallel reduction ordering varies significantly across architectures under the first two conditions, substantially increasing the search space required to fully test the effects of this parallel scheduler-based vulnerability. These results and the methods developed here can help to include machine-level considerations into adversarial robustness assessments, which can make a difference in safety and mission critical applications.

Shanmugavelu, Sanjif [Maxeler Technologies, a Groq