Search NASASearch

SEARCH · Search NASA

Results for “asynchronous methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Method and system for redundancy management of distributed and recoverable digital control system

A method and system for redundancy management is provided for a distributed and recoverable digital control system. The method uses unique redundancy management techniques to achieve recovery and restoration of redundant elements to full operation in an asynchronous environment. The system includes a first computing unit comprising a pair of redundant computational lanes for generating redundant control commands. One or more internal monitors detect data errors in the control commands, and provide a recovery trigger to the first computing unit. A second redundant computing unit provides the same features as the first computing unit. A first actuator control unit is configured to provide blending and monitoring of the control commands from the first and second computing units, and to provide a recovery trigger to each of the first and second computing units. A second actuator control unit provides the same features as the first actuator control unit.

Stange, Kent

Learning to train neural networks for real-world control problems

Over the past three years, our group has concentrated on the application of neural network methods to the training of controllers for real-world systems. This presentation describes our approach, surveys what we have found to be important, mentions some contributions to the field, and shows some representative results. Topics discussed include: (1) executing model studies as rehearsal for experimental studies; (2) the importance of correct derivatives; (3) effective training with second-order (DEKF) methods; (4) the efficacy of time-lagged recurrent networks; (5) liberation from the tyranny of the control cycle using asynchronous truncated backpropagation through time; and (6) multistream training for robustness. Results from model studies of automotive idle speed control serve as examples for several of these topics.

Feldkamp, Lee A.

Advancement of Deep Learning and Geometric Methods for Active Terrain Relative Navigation

To enhance NASA’s precision landing capabilities, in conjunction with the development of a novel active terrain relative navigation (ATRN) and terrain mapping system, denoted SHERIF, this work performed a comparative analysis between both deep-learning (DL) based and geometric approaches to hazard detection (HD) and safe-site-identification (SSI) through hardware-in-the loop testing on the Six degree-of-freedom Tendon Actuated Robot (STAR). The Standalone Hazard Evaluation and Refinement using Instrument Findings (SHERIF) system is capable of ingesting sensor data at an asynchronous rate, stitching successive terrain scans together to yield a high-resolution digital elevation map (DEM), performing absolute and relative localization using novel 3D feature extraction and matching methods, and HD/SSI activities. The DL-based HD/SSI algorithm provides a modular alternative to classical geometric approaches which have performance times that scale with map resolution. As the adoption of AI solutions become more prevalent for autonomous system decision making, it is prudent to explore the utility of such solutions in applications where they traditionally excel, such as image classification. Along with the development of a DL-based HD system, this work performed the first comparative analysis between DL and geometric approaches to HD/SSI using real sensor data from real-time testing in a relevant environment.

Davis Adams

Alternative majority-voting methods for real-time computing systems

Two techniques that provide a compromise between the high time overhead in maintaining synchronous voting and the difficulty of combining results in asynchronous voting are proposed. These techniques are specifically suited for real-time applications with a single-source/single-sink structure that need instantaneous error masking. They provide a compromise between a tightly synchronized system in which the synchronization overhead can be quite high, and an asynchronous system which lacks suitable algorithms for combining the output data. Both quorum-majority voting (QMV) and compare-majority voting (CMV) are most applicable to distributed real-time systems with single-source/single-sink tasks. All real-time systems eventually have to resolve their outputs into a single action at some stage. The development of the advanced information processing system (AIPS) and other similar systems serve to emphasize the importance of these techniques. Time bounds suggest that it is possible to reduce the overhead for quorum-majority voting to below that for synchronous voting. All the bounds assume that the computation phase is nonpreemptive and that there is no multitasking.

Shin, Kang G.

Dynamic grid refinement for partial differential equations on parallel computers

The fast adaptive composite grid method (FAC) is an algorithm that uses various levels of uniform grids to provide adaptive resolution and fast solution of PDEs. An asynchronous version of FAC, called AFAC, that completely eliminates the bottleneck to parallelism is presented. This paper describes the advantage that this algorithm has in adaptive refinement for moving singularities on multiprocessor computers. This work is applicable to the parallel solution of two- and three-dimensional shock tracking problems.

Mccormick, S.

Performance Optimization Methods for a Memory-Bound, Unstructured-Grid CFD Application on Massively Parallel GPU Platforms

Computational performance of the FUN3D unstructured-grid computational fluid dynamics (CFD) application on massively parallel GPU environments is memory-bound and highly dependent upon efficient reads from and atomic updates to the irregular cell-, edge-, and node-based data structures. In this talk, we present recent efforts into optimizing select performance-critical kernels on NVIDIA Tesla V100 and A100 GPUs and AMD CDNA MI100 GPUs. A novel use of L2 cache residency controls and asynchronous loads into on-chip shared memory are explored on the A100 GPU for the sparse iterative solver, which is dominated by mixed-precision, sparse matrix vector multiplication. Demonstrations show that these methods improve global memory bandwidth utilization by 13.5% on the A100 GPU. Several techniques are also presented that use registers and/or shared memory to facilitate array transposition and aggregation which combine to reduce the frequency and increase the cache efficiency of floating-point atomic updates to the irregular data structures. These methods are demonstrated to improve the kernel throughput by nearly 500% on select kernels on the AMD MI100 over atomic updates directly to global memory. Overall, both V100 and A100 GPUs outperformed the MI100 GPU on kernels dominated by double-precision atomic updates; however, the techniques demonstrated here reduced the performance gap and improved the MI100 performance.

GPU CPU unstructured CFD memory

Breaking the Million-Electron and 1 EFLOP/s Barriers: Biomolecular-Scale Ab Initio Molecular Dynamics Using MP2 Potentials

The accurate simulation of complex biochemical phenomena has historically been hampered by the computational requirements of high-fidelity molecular-modeling techniques. Quantum mechanical methods, such as ab initio wave-function (WF) theory, deliver the desired accuracy, but have impractical scaling for modeling biosystems with thousands of atoms. Combining molecular fragmentation with MP2 perturbation theory, this study presents an innovative approach that enables biomolecular-scale ab initio molecular dynamics (AIMD) simulations at WF theory level. Leveraging the resolution-of-the-identity approximation for Hartree-Fock and MP2 gradients, our approach eliminates computationally intensive four-center integrals and their gradients, while achieving near-peak performance on modern GPU architectures. The introduction of asynchronous time steps minimizes time step latency, overlapping computational phases and effectively mitigating load imbalances. Utilizing up to 9,400 nodes of Frontier and achieving 59% (1006.7 PFLOP/s) of its double-precision floating-point peak, our method enables us to break the million-electron and 1EFLOP/s barriers for AIMD simulations with quantum accuracy.

Kurzak, Jakub

Boolean differentiation and integration using Karnaugh maps

Algorithms are presented for differentiation and integration of Boolean functions by means of Karnaugh maps. The algorithms are considered simple when the number of variables is six or less; in this case Boolean differentiation and integration is said to be as easy as the Karnaugh map method of simplifying switching functions. It is suggested that the algorithms would be useful in the analysis of faults in combinational systems and in the synthesis of asynchronous sequential systems which utilize edge-sensitive flip-flops.

Tucker, J. H.

Synchronous Control Effort Minimized for Magnetic-Bearing-Supported Shaft

Various disturbances that are synchronous with the shaft speed can complicate radial magnetic bearing control. These include position sensor target irregularities (runout) and shaft imbalance. The method presented here allows the controller to ignore all synchronous harmonics of the shaft position input (within the closed-loop bandwidth) and to respond only to asynchronous motions. The result is reduced control effort.

Brown, Gerald V.

Real-time transmission of full-motion echocardiography over a high-speed data network: impact of data rate and network quality of service

With high-resolution network transmission required for telemedicine, education, and guided-image acquisition, the impact of errors and transmission rates on image quality needs evaluation. METHODS: We transmitted clinical echocardiograms from 2 National Aeronautics and Space Administration (NASA) research centers with the use of Motion Picture Expert Group-2 (MPEG-2) encoding and asynchronous transmission mode (ATM) network protocol over the NASA Research and Education Network. Data rates and network quality (cell losses [CLR], errors [CER], and delay variability [CVD]) were altered and image quality was judged. RESULTS: At speeds of 3 to 5 megabits per second (Mbps), digital images were superior to those on videotape; at 2 Mbps, images were equivalent. Increasing CLR caused occasional, brief pauses. Extreme CER and CDV increases still yielded high-quality images. CONCLUSIONS: Real-time echocardiographic acquisition, guidance, and transmission is feasible with the use of MPEG-2 and ATM with broadcast quality seen above 3 Mbps, even with severe network quality degradation. These techniques can be applied to telemedicine and used for planned echocardiography aboard the International Space Station.

NASA Discipline Cardiopulmonary

SAGIPS: a physics-inspired scalable asynchronous generative inverse-problem solver

Abstract Solving large-scale inverse problems using deep-learning algorithms have become an essential part of modern research and industrial applications. The complexity of the underlying inverse problem may require the utilization of high performance computing systems which poses a challenge on the algorithmic design of the inverse problem solver. Most deep learning algorithms require, due to their design, custom parallelization techniques in order to be resource efficient while showing a reasonable convergence. In this paper we introduce a S calable A synchronous G enerative I nverse P roblem S olver (SAGIPS) on high-performance computing systems. We present a workflow that utilizes an asynchronous ring-allreduce algorithm to transfer the gradients of the generator network across multiple GPUs. Experiments with a scientific proxy application demonstrate that SAGIPS shows near linear weak scaling, together with a convergence quality that is comparable to traditional methods. The approach presented here allows leveraging Generative Adverserial Network across multiple GPUs, promising advancements in solving complex inverse problems at scale.

97 MATHEMATICS AND COMPUTING

Customized Bayesian optimization for efficient beam tuning at the facility for rare isotope beams

Bayesian optimization (BO) has recently emerged as a powerful approach for on-line beam tuning, and it is rapidly gaining adoption across accelerator facilities due to its flexibility and efficiency in handling complex optimization tasks. At the Facility for Rare Isotope Beams, rapid and reliable tuning is essential to support the delivery of diverse ion species. To improve the practicality of BO in this setting, we implemented several enhancements, including scalarized composite objective construction for multicriteria optimization, asynchronous evaluation for better resource utilization, prior-mean-assisted optimization to accelerate convergence, and a local search strategy for rapid completion of the task. We present the details of these methods, discuss challenges-encountered, and share our experience applying them to specific beam-tuning tasks.

Accelerators & storage rings

Using Pipelined XNOR Logic to Reduce SEU Risks in State Machines

Single-event upsets (SEUs) pose great threats to avionic systems state machine control logic, which are frequently used to control sequence of events and to qualify protocols. The risks of SEUs manifest in two ways: (a) the state machine s state information is changed, causing the state machine to unexpectedly transition to another state; (b) due to the asynchronous nature of SEU, the state machine's state registers become metastable, consequently causing any combinational logic associated with the metastable registers to malfunction temporarily. Effect (a) can be mitigated with methods such as triplemodular redundancy (TMR). However, effect (b) cannot be eliminated and can degrade the effectiveness of any mitigation method of effect (a). Although there is no way to completely eliminate the risk of SEU-induced errors, the risk can be made very small by use of a combination of very fast state-machine logic and error-detection logic. Therefore, one goal of two main elements of the present method is to design the fastest state-machine logic circuitry by basing it on the fastest generic state-machine design, which is that of a one-hot state machine. The other of the two main design elements is to design fast error-detection logic circuitry and to optimize it for implementation in a field-programmable gate array (FPGA) architecture: In the resulting design, the one-hot state machine is fitted with a multiple-input XNOR gate for detection of illegal states. The XNOR gate is implemented with lookup tables and with pipelines for high speed. In this method, the task of designing all the logic must be performed manually because no currently available logic synthesis software tool can produce optimal solutions of design problems of this type. However, some assistance is provided by a script, written for this purpose in the Python language (an object-oriented interpretive computer language) to automatically generate hardware description language (HDL) code from state-transition rules.

Le, Martin

SVM-Based Synchronized Fault Detection for 100% Renewable Microgrids

Traditional protection schemes face significant challenges when applied to microgrids with high penetrations of renewables with inverter-based resources (IBRs). The proliferation of advanced sensing and communication technologies has generated copious data, offering an opportunity to overcome these limitations using data-driven machine learning approaches. This work proposes a novel approach based on a support vector machine (SVM) for detecting faults within a 100% renewable microgrid. The approach encompasses a systematic offline training stage for the development of a linear SVM-based fault detection algorithm. This process covers offline data collection from the microgrid under study, the extraction of features such as positive- and negative-sequence components and the total harmonic distortion of the voltage and current measurements of the relays, and the design of the linear SVM-based classifier. During the online implementation, however, different classifiers can exhibit asynchronicity in detecting the fault inception at different subcycle-to-cycle period-level delays. To circumvent this asynchronicity issue, a separate algorithm is developed for each relay to estimate the fault inception time as close to the real fault time. The performance of the proposed SVM-based synchronized fault detection method is evaluated using online time-domain simulation studies on a microgrid test system. The results corroborate the reliability of the fault detection scheme when tested under various fault cases (fault types, locations, and impedances) and non-fault cases during both grid-tied and islanded operation modes.

100% microgrid

SVM-Based Synchronized Fault Detection for 100% Renewable Microgrids: Preprint

Traditional protection schemes face significant challenges when applied to microgrids with high penetrations of renewables with inverter-based resources (IBRs). The proliferation of advanced sensing and communication technologies has generated copious data, offering an opportunity to overcome these limitations using data-driven machine learning approaches. This work proposes a novel approach based on a support vector machine (SVM) for detecting faults within a 100% renewable microgrid. The approach encompasses a systematic offline training stage for the development of a linear SVM-based fault detection algorithm. This process covers offline data collection from the microgrid under study, the extraction of features such as positive- and negative-sequence components and the total harmonic distortion of the voltage and current measurements of the relays, and the design of the linear SVM-based classifier. During the online implementation, however, different classifiers can exhibit asynchronicity in detecting the fault inception at different subcycle-to-cycle period-level delays. To circumvent this asynchronicity issue, a separate algorithm is developed for each relay to estimate the fault inception time as close to the real fault time. The performance of the proposed SVM-based synchronized fault detection method is evaluated using online time-domain simulation studies on a microgrid test system. The results corroborate the reliability of the fault detection scheme when tested under various fault cases (fault types, locations, and impedances) and non-fault cases during both grid-tied and islanded operation modes.

100% microgrid

Cascaded clocks measurement and simulation findings

This paper will examine aspects related to network synchronization distribution and the cascading of timing elements. Methods of timing distribution have become a much debated topic in standards forums and among network service providers (both domestically and internationally). Essentially these concerns focus on the need to migrate their existing network synchronization plans (and capabilities) to those required for the next generation of transport technologies (namely, the Synchronous Digital Hierarchy (SDH), Synchronous Optical Networks (SONET), and Asynchronous Transfer Mode (ATM). The particular choices for synchronization distribution network architectures are now being evaluated and are demonstrating that they can indeed have a profound effect on the overall service performance levels that will be delivered to the customer. The salient aspects of these concerns reduce to the following: (1) identifying that the devil is in the details of the timing element specifications and the distribution of timing information (i.e., small design choices can have a large performance impact); (2) developing a standardized method of performance verification that will yield unambiguous results; and (3) presentation of those results. Specifically, this will be done for two general cases: an ideal input, and a noisy input to a cascaded chain of slave clocks.

Chislow, Don

Electrode strain dynamics in layered intercalation battery cathodes

Rechargeable batteries using electrodes based on intercalation chemistry exhibit notable cyclability, yet their performance still suffers from chemomechanical degradation. In this study, by combining a suite of operando microscopy methods, we explored electrode strain evolution and observed intricate particle cluster rearrangement under electrochemical stimuli. We show that early-stage strain accumulation in intercalation cathodes occurs during the period of interparticle charge transfer and redox reactions stemming from asynchronous coupling and decoupling between chemical (de)intercalation and physical grain motion. This interplay drives heterogeneous redox activity, localized charge equilibration, and multiscale strain cascades that propagate through an asynchronous network of chemical-mechanical interactions. Together, these findings reveal how collective particle dynamics and hierarchical strain transmission dictate electrode deformation and degradation in intercalation cathodes.

25 ENERGY STORAGE

A software architecture for multidisciplinary applications: Integrating task and data parallelism

Data parallel languages such as Vienna Fortran and HPF can be successfully applied to a wide range of numerical applications. However, many advanced scientific and engineering applications are of a multidisciplinary and heterogeneous nature and thus do not fit well into the data parallel paradigm. In this paper we present new Fortran 90 language extensions to fill this gap. Tasks can be spawned as asynchronous activities in a homogeneous or heterogeneous computing environment; they interact by sharing access to Shared Data Abstractions (SDA's). SDA's are an extension of Fortran 90 modules, representing a pool of common data, together with a set of Methods for controlled access to these data and a mechanism for providing persistent storage. Our language supports the integration of data and task parallelism as well as nested task parallelism and thus can be used to express multidisciplinary applications in a natural and efficient way.

Chapman, Barbara