Search NASA⌕ Search

SEARCH · Search NASA

Results for “communication and synchronization reducing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

xSDK: Building an ecosystem of highly efficient math libraries for exascale

Current efforts to build increasingly powerful computer architectures are opening up new avenues for more complex and higher fidelity simulations coupled with data analytics and learning, leading to new scientific insights and deeper understanding. At one extreme, exascale computers will be much faster than previous computer generations (performing 10 18 operations per second—that is, 1,000 times faster than petascale). To achieve these performance improvements, computer architectures are becoming increasingly complex, with deep memory hierarchies, very high node and core counts, and heterogeneous features such as graphics processing units (GPUs). Such architectural changes impact the full breadth of computing scales, as heterogeneity pervades even current-generation laptops, workstations, and moderate-sized clusters. While emerging advanced architectures provide unprecedented opportunities, they also present significant challenges for developers of scientific applications, such as multiphysics and multiscale codes, who must adapt their software to handle disruptive changes in architectures and new programming models that have not yet stabilized. Developers must consider increasing concurrency while reducing communication and synchronization, and other complexities such as the potential for using mixed precision to leverage the compute power available in low-precision tensor cores. On one hand, developers must implement new scientific capabilities, which in turn increase code complexity. On the other hand, the codes must be ported to new architectures, requiring the inclusion of new programming models and the restructuring of code to achieve good performance. Addressing these issues is beyond the capability of any single person or team—leading to the need for collaboration among many teams, who encapsulate their expertise in reusable software and work together to create sustainable software ecosystems.

97 MATHEMATICS AND COMPUTING↗

Designing a parallel Feel-the-Way clustering algorithm on HPC systems

This paper introduces a new parallel clustering algorithm, named Feel-the-Way clustering algorithm, that provides better or equivalent convergence rate than the traditional clustering methods by optimizing the synchronization and communication costs. Our algorithm design centers on how to optimize three factors simultaneously: reduced synchronizations, improved convergence rate, and retained same or comparable optimization cost. To compare the optimization cost, we use the Sum of Square Error (SSE) cost as the metric, which is the sum of the square distance between each data point and its assigned clusters. Compared with the traditional MPI k-means algorithm, the new Feel-the-Way algorithm requires less communications among participating processes. As for the convergence rate, the new algorithm requires fewer number of iterations to converge. As for the optimization cost, it obtains the SSE costs that are close to the k-means algorithm. In the paper, we first design the full-step Feel-the-Way k-means clustering algorithm that can significantly reduce the number of iterations that are required by the original k-means clustering method. Next, we improve the performance of the full-step algorithm by adopting an optimized sampling-based approach, named reassignment-history-aware sampling. Our experimental results show that the optimized sampling-based Feel-the-Way method is significantly faster than the widely used k-means clustering method, and can provide comparable optimization costs. More extensive experiments with several synthetic datasets and real-world datasets (e.g., MNIST, CIFAR-10, ENRON, and PLACES-2) show that the new parallel algorithm can outperform the open source MPI k-means library by up to 110% on a high-performance computing system using 4,096 CPU cores. In addition, the new algorithm can take up to 51% fewer iterations to converge than the k-means clustering algorithm.

97 MATHEMATICS AND COMPUTING↗

Lion Cub: Minimizing Communication Overhead in Distributed Lion

Communication overhead is a key challenge in distributed deep learning, especially on slower Ethernet intercon nects, and given current hardware trends, communication is likely to become a major bottleneck. While gradient compression techniques have been explored for SGD and Adam, the Lion optimizer has the distinct advantage that its update vectors are the output of a sign operation, enabling straightforward quantization. However, simply compressing updates for communication and using techniques like majority voting fails to lead to end-to-end speedups due to inefficient communication algorithms and reduced convergence. We analyze three factors critical to distributed learning with Lion: optimizing communication methods, identifying effective quantization methods, and assessing the necessity of momentum synchronization. Our findings show that quantization techniques adapted to Lion and selective momentum synchronization can significantly reduce communication costs while maintaining convergence. We combine these into Lion Cub, which enables up to 5x speedups in end-to-end training compared to Lion. This highlights Lion’s potential as a communication-efficient solution for distributed training.

97 MATHEMATICS AND COMPUTING↗

Optimization of Asynchronous Communication Operations through Eager Notifications

UPC++ is a C++ library implementing the Asynchronous Partitioned Global Address Space (APGAS) model. We propose an enhancement to the completion mechanisms of UPC++ used to synchronize communication operations that is designed to reduce overhead for on-node operations. Our enhancement permits eager delivery of completion notification in cases where the data transfer semantics of an operation happen to complete synchronously, for example due to the use of shared-memory bypass. This semantic relaxation allows removing significant overhead from the critical path of the implementation in such cases. We evaluate our results on three different representative systems using a combination of microbenchmarks and five variations of the the HPCChallenge RandomAccess benchmark implemented in UPC++ and run on a single node to accentuate the impact of locality. We find that in RMA versions of the benchmark written in a straightforward manner (without manually optimizing for locality), the new eager notification mode can provide up to a 25% speedup when synchronizing with promises and up to a 13.5x speedup when synchronizing with conjoined futures. We also evaluate our results using a graph matching application written with UPC++ RMA communication, where we measure overall speedups of as much as 11% in single-node runs of the unmodified application code, due to our transparent enhancements.

Kamil, Amir↗

Utilization of FSK communications for time

The method of time control at a radio station for time signal transmission using frequency shift keying, is discussed. Methods for extracting a recovered clock at the receiver are described. Results of a test involving the methods of processing time signals are presented. The requirements for adapting Precise Time and Time Interval control to a communications system are analyzed. A method for reducing the time required for synchronization or identification of messages and the fallout from the communication system to passive timekeeping users are reported.

Stone, R. R., Jr.↗

Rapid sync acquisition system Patent

System designed to reduce time required for obtaining synchronization in data communication with spacecraft utilizing pseudonoise codes

Anderson, T. O.↗

torch-einshard v1.0

torch-einshard is a Python library for describing local and distributed PyTorch tensor computations with compact, einsum-like notation. Its expressions name logical axes, specify how they are sharded across a PyTorch DeviceMesh, and represent partial reductions. The library automatically performs contractions, permutations, reshaping, splitting, gathering, reduction, reduce-scatter, and repartitioning while preserving autograd. Additional features include sharding-aware FFTs, tensor rolls, halo exchange, sliding windows, 1D–3D convolutions, uneven-shard handling, parameter initialization and gradient management, and cost-based execution planning. It is designed for scientific machine learning and large-model workloads, including tensor-, sequence-, and spatial-parallel MLPs, attention, convolutions, and spectral operations. Compared with manually combining torch.einsum and distributed collectives, torch-einshard expresses both the mathematical operation and data placement in one readable formula. This reduces boilerplate and synchronization errors, keeps forward and backward communication consistent, and allows the library to select optimized collective strategies without changing model code.

Morozov, Dmitriy [Lawrence Berkeley National Labor↗

Narrow-bandwidth receiver

Synchronous switching circuit reduces bandwidth and improves sensitivity of communications receiver. With modified receiver, signals 35 db below level can be detected.

Manus, E. A.↗

Scalable In Situ Computation of Lagrangian Representations via Local Flow Maps

In situ computation of Lagrangian flow maps to enable post hoc time-varying vector field analysis has recently become an active area of research. However, the current literature is largely limited to theoretical settings and lacks a solution to address scalability of the technique in distributed memory. To improve scalability, we propose and evaluate the benefits and limitations of a simple, yet novel, performance optimization. Our proposed optimization is a communication-free model resulting in local Lagrangian flow maps, requiring no message passing or synchronization between processes, intrinsically improving scalability, and thereby reducing overall execution time and alleviating the encumbrance placed on simulation codes from communication overheads. To evaluate our approach, we computed Lagrangian flow maps for four time-varying simulation vector fields and investigated how execution time and reconstruction accuracy are impacted by the number of GPUs per compute node, the total number of compute nodes, particles per rank, and storage intervals. Our study consisted of experiments computing Lagrangian flow maps with up to 67M particle trajectories over 500 cycles and used as many as 2048 GPUs across 512 compute nodes. In all, our study contributes an evaluation of a communication-free model as well as a scalability study of computing distributed Lagrangian flow maps at scale using in situ infrastructure on a modern supercomputer.

Sane, Sudhanshu↗

Reducing Communication Overhead in Federated Learning for Network Anomaly Detection with Adaptive Client Selection

Communication overhead in federated learning (FL) poses a significant challenge for network anomaly detection systems, where the myriad of client configurations and network conditions can severely impact system efficiency and detection accuracy. While existing approaches attempt to address this through individual optimization techniques, they often fail to maintain the delicate balance between reduced overhead and detection performance. This paper presents an adaptive FL framework that dynamically combines batch size optimization, client selection, and asynchronous updates to achieve efficient anomaly detection. Through extensive profiling and experimental analysis on two distinct datasets-UNSW-NBIS for general network traffic and ROAD for automotive networks-our framework reduces communication overhead by 97.6%; (from 700.0s to 16.8s) compared to synchronous baseline approaches while maintaining comparable detection accuracy (95.10%; vs. 95.12%;). Statistical validation using Mann-Whitney U test confirms significant improvements (p < 0.05) over existing FL approaches across both datasets, demonstrating the framework's adaptability to different network security contexts. Detailed profiling analysis reveals the efficiency gains through dramatic reductions in GPU operations and memory transfers while maintaining robust detection performance under varying client conditions.

Marfo, William [University of Texas at El Paso]↗

NASA Tech Briefs, May 2009

Topics covered include: Valve-"Health"-Monitoring System; Microstrip Antenna for Remote Sensing of Soil Moisture and Sea Surface Salinity; Biomedical Wireless Ambulatory Crew Monitor; Wireless Avionics Packet to Support Fault Tolerance for Flight Applications; Aerobot Autonomy Architecture; Submillimeter Confocal Imaging Active Module; Traveling-Wave Maser for 32 GHz; System Synchronizes Recordings from Separated Video Cameras; Piecewise-Planar Parabolic Reflectarray Antenna; Reducing Interference in ATC Voice Communication; EOS MLS Level 1B Data Processing, Version 2.2; Auto-Generated Semantic Processing Services; Geospatial Authentication; Maneuver Automation Software; Event Driven Messaging with Role-Based Subscriptions; Estimating Relative Positions of Outer-Space Structures; Fabricating PFPE Membranes for Capillary Electrophoresis; Linear Actuator Has Long Stroke and High Resolution; Installing a Test Tap on a Metal Battery Case; Fabricating PFPE Membranes for Microfluidic Valves and Pumps; Room-Temperature-Cured Copolymers for Lithium Battery Gel Electrolytes; Catalysts for Efficient Production of Carbon Nanotubes; Amorphous Silk Fibroin Membranes for Separation of CO2; "Zero-Mass" Noninvasive Pressure Transducers; Radial-Electric-Field Piezoelectric Diaphragm Pumps; Ejector-Enhanced, Pulsed, Pressure-Gain Combustor; Suppressing Ghost Diffraction in E-Beam-Written Gratings; Target-Tracking Camera for a Metrology System; Polarimetric Imaging using Two Photoelastic Modulators; Miniature Wide-Angle Lens for Small-Pixel Electronic Camera; Modal Filters for Infrared Interferometry; Mo(3)Sb(7-x)Te(x) for Thermoelectric Power Generation; Two-Dimensional Quantum Model of a Nanotransistor; Scanning Miniature Microscopes without Lenses; Manipulating Neutral Atoms in Chip-Based Magnetic Traps; Expansion Compression Contacts for Thermoelectric Legs; Processing Electromyographic Signals to Recognize Words; Physical Principle for Generation of Randomness; DSN Beowulf Cluster-Based VLBI Correlator; Hybrid NN/SVM Computational System for Optimizing Designs; Criteria for Modeling in LES of Multicomponent Fuel Flow; Computerized Machine for Cutting Space Shuttle Thermal Tiles; Orbiting Depot and Reusable Lander for Lunar Transportation; FPGA-Based Networked Phasemeter for a Heterodyne Interferometer; Aquarius Digital Processing Unit; Three-Dimensional Optical Coherence Tomography; Benchtop Antigen Detection Technique using Nanofiltration and Fluorescent Dyes; Isolation of Precursor Cells from Waste Solid Fat Tissue; Identification of Bacteria and Determination of Biological Indicators; Further Development of Scaffolds for Regeneration of Nerves; Chemically Assisted Photocatalytic Oxidation System; Use of Atomic Oxygen for Increased Water Contact Angles of Various Polymers for Biomedical Applications; Crashworthy Seats Would Afford Superior Protection; Open-Access, Low-Magnetic-Field MRI System for Lung Research; Microfluidic Mixing Technology for a Universal Health Sensor; Microfluidic Extraction of Biomarkers using Water as Solvent; Microwell Arrays for Studying Many Individual Cells; Droplet-Based Production of Liposomes; and Identifying and Inactivating Bacterial Spores

Source record↗

PTTI-aided ephemeris calculation and rapid data link acquisition for manned space flight

Complexity of future manned space flight mission control can be significantly reduced by integrating GPS, the PTTI source, into telemetry, tracking, and command (TT&C). Future telecommunications, space tracking electronic intelligence, metrology, navigation, and data acquisition will thereby be served, including: on-board ephemeris determination, reduced synchronization time for time division multiple access (TDMA) links, and in-flight clock calibration, increasing on-board autonomy and reducing ground support costs. Manned space transportation through the first quarter of the 21st century will probably depend on a mix of vehicles, including the Advanced Manned Launch System (AMLS), the Personnel Launch System (PLS), and continued use of the Shuttle Fleet. Precise Ephemeris is important on-board for mission success, status monitoring, also for rendezvous and docking. Use of GPS can eliminate ground based tracking/processing, enhancing autonomy and reducing communications bandwidth. GPS time can simplify complicated functions used in bandwidth efficient time division multiple access (TDMA) communications, such as: precise and realtime synchronization of receive reference timing, transmit-timing and acquisition control, unique synchronization word (UW) detection, and elastic buffering. High clock accuracy provides increased signal-to-noise (S/N) ratio during acquisition, permitting narrower acquisition frequency and time windows. Spaceborne systems requirements to provide capabilities such as: refinement of the GEM-72 gravity model based on satellite tracking observations from ATS-6 to GEOS-3, relativistic clock experiments, NASA crustal dynamics program for developing space geodetic techniques to study the earth's crust, its gravity field, and earthquake mechanisms, and multi-disciplinary space geodetic tracking for studying global climatic changes are also reviewed.

Anderman, Alfred↗

Synchrophasor spoofing detection and remediation for wide-area damping control

Evolving cyber-attack threats put at risk automatic closed-loop systems to be incorporated in the smart grid. Wide-area control systems are particularly vulnerable to signal spoofing attacks due to sensor remoteness and dependence on satellite communication for time synchronization. A successful cyber-attack on a wide-area controller has the potential to reduce relative stability of the power system or worse, destabilize it. As such, detection algorithms must be deployed as defense against such attacks with the ability to autonomously correct for detected tampering or misoperation. The Spoof Catch and Restore Routine (SCR 2 ), a combination of three real-time spoof detectors, each requiring limited information about the plant, is reported here. Nonlinear simulations of a compromised wide-area control system deployed in the Western Interconnection show the effectiveness of SCR 2 in detecting both delay-type and counterfeit-type spoofing attacks on wide-area sensors.

42 ENGINEERING↗

Constellation Program Electrical Ground Support Equipment Research and Development

At the Kennedy Space Center, I engaged in the research and development of electrical ground support equipment for NASA's Constellation Program. Timing characteristics playa crucial role in ground support communications. Latency and jitter are two problems that must be understood so that communications are timely and consistent within the Kennedy Ground Control System (KGCS). I conducted latency and jitter tests using Alien-Bradley programmable logic controllers (PLCs) so that these two intrinsic network properties can be reduced. Time stamping and clock synchronization also play significant roles in launch processing and operations. Using RSLogix 5000 project files and Wireshark network protocol analyzing software, I verified master/slave PLC Ethernet module clock synchronization, master/slave IEEE 1588 communications, and time stamping capabilities. All of the timing and synchronization test results are useful in assessing the current KGCS operational level and determining improvements for the future.

McCoy, Keegan S.↗

A synchronized computational architecture for generalized bilateral control of robot arms

This paper describes a computational architecture for an interconnected high speed distributed computing system for generalized bilateral control of robot arms. The key method of the architecture is the use of fully synchronized, interrupt driven software. Since an objective of the development is to utilize the processing resources efficiently, the synchronization is done in the hardware level to reduce system software overhead. The architecture also achieves a balaced load on the communication channel. The paper also describes some architectural relations to trading or sharing manual and automatic control.

Bejczy, Antal K.↗

Crew Health and Performance Integrated Data Architecture (CHP-IDA) TechPort May 2024

Future exploration missions to Mars will have increased need for crew autonomy. Crew Health & Performance (CHP) related data on the ISS is currently, manually downlinked and in disparate locations, which limits crew autonomy for future missions. The CHP-IDA project is developing a backend data system platform that grants the ability to seamlessly collect, store, process, and display CHP-related data for exploration missions. This platform allows for integration of data and advanced analytics that offer crew and ground teams better insight into the crew’s health and performance. It also enables applications that can improve the crew’s ability to provide more autonomous medical care during exploration missions. Data will be collected automatically to reduce crew and ground team time and effort and will synchronize across all in-mission vehicles, habitats, and ground as communication delay permits. The Human Research Program’s (HRP) Medical Data Architecture (MDA) project focused on this backend data architecture but for medical data only. The CHP-IDA project, a joint effort between HRP’s Exploration Medical Capability (ExMC) element and the Exploration Medical Integrated Product Team (XMIPT), expands this capability to all relevant CHP-related data. The additional inputs from nutrition, environment, exercise, radiation, and any other relevant sources will give more insight into crew’s health and performance. Currently, the Human Systems Engineering and Integration Division at Johnson Space Center (JSC) is designing the system. The team completed a system requirements review (SRR) in FY22 and now the focus is on core software development, testbed buildup, and use case scenario demonstration. An end-to-end demonstration with multiple data sources across CHP domains is schedule for the end of FY24 where all three focus areas will be displayed. Following this ground demo, the software will be completed, tested, and validated for flight.

Courtney M Schkurko↗

CORE-BFS: Communication-Optimized REctangular-partitioned BFS Achieving 160.845 TeraTEPS on Frontier Supercomputer

Distributed Breadth-First Search (BFS) is fundamental to many large-scale graph applications, but its performance on parallel systems is often limited by high communication overhead. This paper presents CORE-BFS, an extremely scalable GPU-based BFS implementation that introduces a unique rectangular 2D partitioning-based design for Frontier supercomputer. To further improve performance, we propose four key optimizations: (1) Rectangular 2D-partition specific data formats that use two compressed row and one compressed column status array bitmaps combined with a Double Compressed Sparse Row (DCSR) format per partition, reducing memory footprint and inter-rank traffic; (2) Adaptive frontier & communication strategy that unifies top-down and bottom-up traversal on the rectangular layout, uses lazy synchronization in top-down levels, and switches variants based on frontier size to minimize communication overhead; (3) Frontier-split degree-aware update that maps frontier vertices to thread-centric, wavefront-centric, and block-centric kernels based on their degree to improve GPU utilization and memory coalescing; (4) Row-reduction pipeline that overlaps bottom-up adjacency list processing with row-wise bitmap reduction to hide inter-rank latency. Together, these techniques increase parallelism while reducing memory and communication overhead. On the Graph500 benchmark, CORE - BFS scales up to 9,248 Frontier nodes with scale-42 graphs and reaches 160.845 TTEPS, delivering a 5.42 × speedup over our previous Frontier implementation.

Yang, Haoshen [Rutgers University]↗