Search NASA⌕ Search

SEARCH · Search NASA

Results for “COMMUNICATION”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Enhancing Network Anomaly Detection Using Graph Neural Networks

In the world of Internet of Things (IoT) networks, where devices are constantly communicating, keeping them secure from cyber threats is critical. This paper introduces a novel approach to detecting unusual and potentially harmful activities in these networks using graph neural networks (GNNs). We combine two specific types of GNNs-GraphSAGE and graph attention networks (GAT)-to create a model that understands and represents the behaviors and interactions in a network. GraphSAGE creates an embedding of network activities by examining local data interactions, while GAT directs the model's focus to the most critical interactions. By integrating these two methods in a single model that considers different types of interactions (both host and flow nodes), we aim to create a system that accurately represents the current state of a network and can also spot anomalies effectively while reducing false positives and negatives. Our innovative approach has demonstrated promising results, achieving an accuracy of 98% on the UNSW-NB15 dataset, significantly outperforming standalone GraphSAGE and GAT models. This underscores its potential as a robust framework for securing IoT networks against cyber threats and anomalies.

Marfo, William↗

A Novel Authentication Management for the Data Security of Smart Grid

Bidirectional wireless communication is employed in various smart grid components such as smart meters and control and monitoring applications where security is vital. The Trusted Third Party (TTP) and wireless connectivity between the smart meter and the third party in the key management-based encryption techniques for the smart grid are expected to be totally trustworthy and dependable. In a wired/wireless medium, however, a man-in-the-middle may seek to disrupt, monitor and manipulate the network, or simply execute a replay attack, revealing its vulnerability. Recognizing this, this study presents a novel authentication management (model) comprised of two layer security schema. The first layer implements an efficient novel encryption method for secure data exchange between meters and control center with the help of two partially trusted simple servers (constitutes the TTP). In this setting, one server handles the data encryption between the meter and control center/central database, and the other server administers the random sequence of data transmission. The second layer monitors and verifies exchanged data packets among smart meters. It detects abnormal packets from suspicious sources. To implement this node-to-node authentication, One class support vector machine algorithm is proposed which takes advantages of the location information as well as the data transmission history (node identification, packet size, and data transmission frequency). This schema secures data communication, and imposes a comprehensive privacy throughout the system without considerably extending the complexity of the conventional key management scheme.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Architecture Design for Remote Operation of Microreactors: Poster

The nuclear sector is pursuing a number of advanced reactor concepts, one of which is the microreactor, an advanced reactor characterized by a power output of less than 20 MWth. These reactors are intended for use in applications where traditional power solutions are economically or logistically impractical. One key feature required for the successful deployment of microreactors is a remote operation capability, which can dramatically cut staffing expenses by eliminating the need for licensed operators to be present at the site of each reactor and instead concentrate in a centrally located operations center. In moving to a remote operations framework for nuclear reactors, the number of potential attack surfaces for a cyber adversary looking to cause harm or disruption increases. Therefore, a robust cybersecurity architecture is required to mitigate these potential vulnerabilities. This paper explores the functional infrastructure, security, and communication requirements to adapt remote operations to nuclear applications. It then presents a reference architecture for remote operations of microreactors that applies best practices in cybersecurity and remote communications in the context of a digital twin remote operation system.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Modeling Grid Data Flows for Transmission and Distribution Operations: Review, Design, Next Steps

Operational scenarios of the power grids grow multifold to accommodate the diverse needs of both the utilities and end consumers, and the various other stakeholders in-between. To comprehensively model and apply analytics to support objectives and business functions of grid sectors, a reliable approach to characterize and design data flows is crucial. The flows bridge business functions with communications protocols, stakeholders such as the grid actors, and data interfaces comprising different data objects. Additionally, constraints applied to the flow such as cybersecurity, trust, privacy, and ownership among others intersect these entities, requiring the delineation of their interactions under different scenarios. This paper aims to not only highlight relevant research in the space of grid data flows, but also proposes, for the transmission-distribution sector, a novel modeling approach that marries the aforementioned entities: objectives, business functions, data interfaces, communication protocols, data stakeholders, and flow constraints. It elaborates on the design philosophy and the significance of each entity within the model and applies it to an example function of fault location, isolation and service restoration (FLISR). Finally, the next steps to extend the application of this data flow model for other practical operational scenarios are discussed.

Sundararajan, Aditya [ORNL] (ORCID:000000033577854↗

An Evaluation of the Effect of Network Cost Optimization for Leadership Class Supercomputers

Dragonfly-based networks are an extensively deployed network topology in large-scale high-performance computing due to their cost-effectiveness and efficiency. The US will soon have three Exascale supercomputers for leadership class workloads deployed using dragonfly networks. Compared to indirect networks of similar scale, the dragonfly network has considerably reduced cable lengths, cable counts, and switch counts, resulting in significant network cost savings for a given system size, however, these cost reductions result in reduced global minimal paths and more challenging routing. Additionally, large scale dragonfly networks often require a taper at the global link level, resulting in less bisection bandwidth than is achievable in other traditional non-blocking topologies of equivalent scale. While dragonfly networks have been extensively studied, they have yet to be fully evaluated in an extreme scale (i.e., exascale) system that targets capability workloads. In this paper, we present the results of the first large scale evaluation of a dragonfly network on an exascale system (Frontier) and compare its behavior to a similar scale fat-tree network on a previous generation TOP500 system (Summit). This evaluation aims to determine the effect of network cost optimizations by measuring a tapered topology’s impact on capability workloads. Our evaluation is based on a collection of synthetic microbenchmarks, mini-apps, and full scale applications. It compares the scaling efficiencies of each benchmark between the dragonfly-based Frontier and the fat-tree-based Summit systems. Our results show that a dragonfly network is $\sim \mathbf{3 0 \%}$ more cost efficient than a fat-tree topology, which amortizes to $\sim 3 \%$ of an exascale system cost. Furthermore, while tapered dragonfly networks impose significant tradeoffs, the impacts are not as broad as initially thought and are mostly seen in applications with global communication patterns, particularly all-to-all (e.g., FFT-based algorithms), but also local communication patterns (e.g., nearest-neighbor algorithms) that are sensitive to network performance variability.

Khan, Awais↗

Reconfigurable Network Slicing Orchestration in Network Function Virtualization Compatible Operational Technology Environment

The ongoing transition to Industry 4.0, which is characterized by increased inter-connectivity of cyber-physical systems, requires having time-sensitive, high throughput, and secure transfer of critical data in industrial sites. In this context, network slicing emerges as a critical tool to ensure timely data delivery by provisioning the network resources to cater to specific applications’ requirements and mitigating potential cyber attacks. To address these challenges, this paper aims to tackle two key questions essential for the successful implementation of network slicing in industrial environments. First, it investigates architectural considerations for developing a network infrastructure capable of supporting network slicing functionalities effectively. The proposed approach significantly improves deployment efficiency over traditional manual configurations. Second, it delves into the automated orchestration process, elucidating the steps and components involved in transitioning from a static network management approach to dynamically leverage network function virtualization schemes for creating network slices in ad-hoc manner. The system demonstrates high throughput suitable for production-level solutions and maintains exceptionally low latency, making it ideal for ultra-reliable low-latency communications. Even with increased network demands, the system remains stable, with effective Quality of Service (QoS) management, ensuring reliable performance under varying conditions. The proposed architecture outlines the necessary components, services, and communication protocols required for a production-level orchestrator for network segmentation in SCADA environments.

Rodiles Delgado, Brian G.↗

Traffic Shaping to Traffic Engineering in Time-Sensitive OT Network

Modern industrial automation systems increasingly depend on network infrastructures for time-critical communication, driving the need for solutions that guarantee timely and reliable data delivery. IEEE 802.1 Time-Sensitive Networking (TSN) holds significant promise for converging Information Technology (IT) and Operational Technology (OT) networks, enabling interoperability and supporting the coexistence of mixed-critical traffic crucial for Industry 4.0 and IIoT. To achieve deterministic communication, TSN employs various traffic shapers such as the Time-Aware Shaper (TAS), Asynchronous Traffic Shaper (ATS), and Credit-Based Shaper (CBS). However, the effective deployment of TSN in industrial automation faces several challenges. These include the non-trivial mapping of diverse industrial traffic types to specific shapers, the complexity of optimizing shaper configurations. We present a model for effective traffic engineering within TSN enabled OT Network. Our experiments also demonstrate how shaping of certain traffic types get affected in absence of precise time synchronization and propose possible solutions based on experiment results. Based on our experimental results we provide recommendations on how traffic type assignments should be done and which traffic shaping mechanisms should be used for a particular traffic type.

Sarker, Taposh Kumer [University of Texas at El Pa↗

Enhancing Automotive Intrusion Detection Through Multi-Modal Fusion: A CAN FD-LiDAR Approach

As vehicles become smarter and more autonomous, they increasingly depend on advanced sensors and communication technologies to operate securely. However, such growing dependence on technology—whether it’s CAN (Controller Area Network) for internal communication or LiDAR (Light Detection and Ranging) for sensing the world around them—also expands the attack surface for the types of cyber attacks. Traditional intrusion detection systems (IDS) typically monitor these systems in isolation, limiting their ability to detect sophisticated, crosssystem attacks. To address this, we propose a multi-modal fusion approach that combines real-world CAN FD signals (from the HCRL dataset) with LiDAR features (from the nuScenes dataset) to enhance attack detection. Our method employs a twostage ensemble approach. Calibrated XGBoost and LightGBM models initially process CAN FD (Fuzzing Data) and LiDAR data independently, detecting timing anomalies and space abnormalities. They are subsequently logarithmically combined with a logistic regression meta-model along with 17 engineered features capturing cross-modal behavior, prediction conflicts, and nonlinear interactions. This approach achieves an AUC of 0.87 and an F1-score of 0.82, surpassing single-modality baselines and early fusion methods, at merely 2 ms inference latency. Compared with deep learning competitors, it is 3 times more efficient, providing a lightweight, interpretable, and real time solution to automotive cybersecurity.

97 MATHEMATICS AND COMPUTING↗

A Two-Stage Approach for PV Inverter Engagement in Power Factor Correction and Voltage Regulation

The rapid integration of distributed energy resources, like solar photovoltaics (PVs), can lead to overvolt-age challenges due to reverse power flow and a noticeable decrease in power factor at the substation interface. While existing literature extensively explores utilizing smart inverter capabilities for reactive power flexibility using a volt-var curve (VVC), obtaining time-varying operating points of such curves in real-time is challenging due to computational demands and communication requirements. Similarly, employing optimization-based approaches for reactive power control and active voltage regulation in large-scale distribution feeders is difficult due to the complexity of the problem and the challenges in effectively engaging customer-owned resources. This paper proposes a two-stage strategy to harness smart inverters for reactive power support. The first stage formulates short-term planning by optimally designing VVCs (on a daily or hourly basis) for large-scale solar PVs based on projected system needs and communicating optimal curves to smart inverters in advance. Subsequently, the second stage employs a transactive-based method to involve customer-owned PVs for reactive power support, effectively enhancing overall system performance and addressing real-time demands. In conclusion, the efficacy of this approach will be demonstrated using real-world distribution circuits provided by Vermont Electric Power Company (VELCO) and Vermont Electric Cooperative (VEC).

Poudel, Shiva [Pacific Northwest National Laborato↗

Development and Experimental Validation of a High-Power DC Distribution Testbed for Advanced Charging Infrastructure and Energy Management

This paper presents the development of a hardware testbed for DC-distributed high-power charging (HPC) stations. As DC distributed solutions emerge as a viable solution to optimize HPC site operations, challenges such as interoperability, protection, and seamless integration of distributed energy resources (DER) persist. These issues underscore the need for a robust testing facility to investigate compliance of available commercial off-the-shelf (COTS) market devices. The developed testbed features a DC-distributed charging hub including a charger, emulated energy storage system (ESS), and site level communication and controller implementation. It facilitates the testing of COTS hardware, charger prototypes, standards validation and site energy management system (SEMS) controllers at rated power. This paper details the development of the charging infrastructure platform, implementation of communication system, validation of different SEMS algorithms, and understanding improvements required for future expansion. Using the developed testbed, interoperability gaps for SEMS implementation with multi-vehicle concurrent charging via a multi-port charger are experimentally observed. Aimed at supporting the transition to large-scale EV charging infrastructure deployment and DER integration, this testbed plays a crucial role in conformity testing of COTS device interoperability.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Self-Supervised and Interpretable Anomaly Detection Using Network Transformers

Machine learning and deep neural networks (DNNs) have been proposed as a tool to identify anomalies in computer network communications. However, due the obfuscated nature of off-the-shelf machine learning models, their output often does not provide enough information to isolate the source of the anomaly to take corrective measures. In this article, we introduce the network transformer (NeT), a DNN model for anomaly detection that incorporates the graph structure of the communication network in order to improve interpretability. Further, the presented approach has the following advantages: first, enhanced interpretability by incorporating the graph structure of computer networks; second, provides a hierarchical set of features that enables analysis at different levels of granularity; second, self-supervised training that does not require labeled data. The NeT model was evaluated on a set of anomalous scenarios executed in a real industrial control system. The presented approach successfully identified the anomalies, the devices affected, and the specific connections causing the anomalies, providing a data-driven hierarchical approach to analyze the behavior of a cyber network.

97 MATHEMATICS AND COMPUTING↗

Capacities of Entanglement Distribution From a Central Source

Distribution of entanglement is an essential task in quantum information processing and the realization of quantum networks. In our work, we theoretically investigate the scenario where a central source prepares an N -partite entangled state and transmits each entangled subsystem to one of N receivers through noisy quantum channels. The receivers are then able to perform local operations assisted by unlimited classical communication to distill target entangled states from the noisy channel output. In this operational context, we define the EPR distribution capacity and the GHZ distribution capacity of a quantum channel as the largest rates at which Einstein-Podolsky-Rosen (EPR) states and Greenberger-Horne-Zeilinger (GHZ) states can be faithfully distributed through the channel, respectively. We establish lower and upper bounds on the EPR distribution capacity by connecting it with the task of assisted entanglement distillation. We also construct an explicit protocol consisting of a combination of a quantum communication code and a classical-post-processing-assisted entanglement generation code, which yields a simple achievable lower bound for generic channels. As applications of these results, we give an exact expression for the EPR distribution capacity over two erasure channels and bounds on the EPR distribution capacity over two generalized amplitude damping channels. We also bound the GHZ distribution capacity, which results in an exact characterization of the GHZ distribution capacity when the most noisy channel is a dephasing channel.

42 ENGINEERING↗

Machine Learning for Scalable and Optimal Load Shedding Under Power System Contingency

Prompt and effective corrective actions in response to unexpected contingencies are crucial for improving power system resilience and preventing cascading blackouts. The optimal load shedding (OLS) accounting for network limits has the potential to address the diverse system-wide impacts of contingency scenarios as compared to traditional local schemes. However, due to the fast cascading propagation of initial contingencies, real-time OLS solutions are challenging to attain in large systems with high computation and communication needs. In this paper, we propose a decentralized design that leverages offline training of a neural network (NN) model for individual load centers to autonomously construct the OLS solutions from locally available measurements. Our learning-for-OLS approach can greatly reduce the computation and communication needs during online emergency responses, thus preventing the cascading propagation of contingencies for enhanced power grid resilience. Numerical studies on both the IEEE 118-bus system and a synthetic Texas 2000-bus system have demonstrated the efficiency and effectiveness of our scalable OLS learning design for timely power system emergency operations.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Positioning Accuracy in a Concurrent Robot-CNC Hybrid Manufacturing System

Abstract Additive manufacturing (AM) has gained notoriety for offering advantages over traditional manufacturing methods, such as increased design complexity and flexibility. However, it has not found widespread use beyond rapid prototyping. One hindrance to the acceptance of AM processes in industry is the time and cost of fabrication per component. While metal AM by itself can be inexpensive, extra manufacturing steps in the form of subtractive manufacturing (SM) may need to be performed to reach final part tolerances, leading to hybrid additive-subtractive manufacturing (HASM) of a part, which increases time and cost. A potential area to reduce cost is through increasing the efficiency of the HASM process by conducting additive and subtractive manufacturing simultaneously. Usually, HASM is performed in a process where AM is completed in one machine or cell and transferred to another machine or cell for SM in a sequential assembly line process. This efficiency decreases part cost, but high aspect ratio parts or parts with internal geometry that require interleaved additive deposition and machining cannot be produced. One unexplored solution to simultaneous HASM that allows for interleaved operations is to operate the deposition head and machining spindle concurrently within the same machine envelope, known as concurrent HASM (CHASM). In this type of process, both AM and SM occur simultaneously on a batch of small parts or a single large part, maintaining a high efficiency without sacrificing the full range of complex geometries that AM allows for. A potential approach to the single-machine method could be to combine a robot and mill within the same envelope. A challenge to this approach, however, is control of both systems. Most machine controllers have limited external communication or, if a robot has been integrated, only offer movement of either the robot or mill at any given time. As a result, systems must pause either the AM or SM process to switch between them rather than working simultaneously. The present work investigates the positional accuracy of such a CHASM system comprised of a robotic arm and a 3-axis mill. Open-loop tests with limited communication between machines are performed on the system to verify positional error during concurrent robot-mill movements. Under certain conditions, it is demonstrated that position error can stay within 2 mm for the duration of a single layer; however, these tests show that, generally, the open-loop positioning performance of the system is inadequate for CHASM without part-specific hand-tuning of parameters. Based on these results, a set of requirements for successful robot-CNC CHASM is proposed for future integrations.

Goodwin, Jesse↗

CommBench: Micro-Benchmarking Hierarchical Networks with Multi-GPU, Multi-NIC Nodes

Modern high-performance computing systems have multiple GPUs and network interface cards (NICs) per node. The resulting network architectures have multilevel hierarchies of subnetworks with different interconnect and software technologies. These systems offer multiple vendor-provided communication capabilities and library implementations (IPC, MPI, NCCL, RCCL, OneCCL) with APIs providing varying levels of performance across the different levels. Understanding this performance is currently difficult because of the wide range of architectures and programming models (CUDA, HIP, OneAPI). We present CommBench, a library with cross-system portability and a high-level API that enables developers to easily build microbenchmarks relevant to their use cases and gain insight into the performance (bandwidth & latency) of multiple implementation libraries on different networks. We demonstrate CommBench with three sets of microbenchmarks that profile the performance of six systems. Our experimental results reveal the effect of multiple NICs on optimizing the bandwidth across nodes and also present the performance characteristics of four available communication libraries within and across nodes of NVIDIA, AMD, and Intel GPU networks.

Hidayetoglu, Mert↗

RingX: Scalable Parallel Attention for Long-Context Learning on HPC

The attention mechanism has become foundational for remarkable AI breakthroughs since the introduction of the Transformer, driving the demand for increasingly longer context to power frontier models such as large-scale reasoning language models and high-resolution image/video generators. However, its quadratic computational and memory complexities present substantial challenges. Current state-of-the-art parallel attention methods, such as ring attention, are widely adopted for long-context training but utilize a point-to-point communication strategy that fails to fully exploit the capabilities of modern HPC network architectures. In this work, we propose ringX, a scalable family of parallel attention methods optimized explicitly for HPC systems. By enhancing workload partitioning, refining communication patterns, and improving load balancing, ringX achieves up to 3.4 × speedup compared to conventional ring attention on the Frontier supercomputer. Optimized for both bi-directional and causal attention mechanisms, ringX demonstrates its effectiveness through training benchmarks of a Vision Transformer (ViT) on a climate dataset and a Generative Pre-Trained Transformer (GPT) model, Llama3 8B. Our method attains an end-to-end training speedup of approximately 1.5 × in both scenarios. To our knowledge, the achieved 38% model FLOPs utilization (MFU) for training Llama3 8B with a 1M-token sequence length on 4,096 GPUs represents one of the highest training efficiencies reported for long-context learning on HPC systems. Our code implementation is available at https://github.com/jqyin/ringX-attention.

Yin, Junqi [ORNL] (ORCID:0000000338435520)↗

DoCeph: DPU-Offloaded Messaging in Ceph for Reduced Host CPU Utilization

Ceph is a widely used distributed object store, but its messenger layer imposes substantial CPU overhead on the host. To address this limitation, we propose DoCeph, a DPU-offloaded storage architecture for Ceph that disaggregates the system by offloading the communication-intensive messaging component to the DPU while retaining the storage backend on the host. The DPU efficiently manages communication, using lightweight RPC for metadata operations and DMA for data transfer. Moreover, DoCeph introduces a pipelining technique that overlaps data transmission with buffer preparation, mitigating hardware-imposed transfer size limitations. We implemented DoCeph on a Ceph cluster with NVIDIA BlueField-3 DPUs. Evaluation results indicate that DoCeph cuts host CPU usage by up to 92% while sustaining stable throughput and providing larger performance benefits for object writes over 1 MB.

Park, Kuri [Sogang University]↗