Search NASASearch

SEARCH · Search NASA

Results for “core communications”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

CORE-BFS: Communication-Optimized REctangular-partitioned BFS Achieving 160.845 TeraTEPS on Frontier Supercomputer

Distributed Breadth-First Search (BFS) is fundamental to many large-scale graph applications, but its performance on parallel systems is often limited by high communication overhead. This paper presents CORE-BFS, an extremely scalable GPU-based BFS implementation that introduces a unique rectangular 2D partitioning-based design for Frontier supercomputer. To further improve performance, we propose four key optimizations: (1) Rectangular 2D-partition specific data formats that use two compressed row and one compressed column status array bitmaps combined with a Double Compressed Sparse Row (DCSR) format per partition, reducing memory footprint and inter-rank traffic; (2) Adaptive frontier & communication strategy that unifies top-down and bottom-up traversal on the rectangular layout, uses lazy synchronization in top-down levels, and switches variants based on frontier size to minimize communication overhead; (3) Frontier-split degree-aware update that maps frontier vertices to thread-centric, wavefront-centric, and block-centric kernels based on their degree to improve GPU utilization and memory coalescing; (4) Row-reduction pipeline that overlaps bottom-up adjacency list processing with row-wise bitmap reduction to hide inter-rank latency. Together, these techniques increase parallelism while reducing memory and communication overhead. On the Graph500 benchmark, CORE - BFS scales up to 9,248 Frontier nodes with scale-42 graphs and reaches 160.845 TTEPS, delivering a 5.42 × speedup over our previous Frontier implementation.

Yang, Haoshen [Rutgers University]

Integration of 5G and Time Sensitive Networks in Fossil Energy Generation Systems: A Case Study

Precise timing and data transmission within stringent time constraints are critical for numerous applications, such as robotics, virtual and augmented reality, industrial automation, energy and medical business, and various other sectors. Time-sensitive networks (TSN) and fifth-generation wireless communications (5G) are crucial for industrial communications, enabling convergent communication for various services using a common network core. Applications that are time-sensitive and require deterministic communications with low latency fall into this category, such as generation plant systems. This paper presents a simulated model of 5G-TSN for a fossil power plant based on the wired network parameters implemented in situ. Metrics analysis, comparison and future work are presented. © 2024 IEEE.

01 COAL, LIGNITE, AND PEAT

Implementation of ISO 15118-202 messages within Everest EV Charging Open Source Framework [SWR-25-56]

This software implements the messages defined in the ISO 15118-202 standard within the Everest EV Charging open source framework. The protocol and messages defined in the ISO 15118-202 standard enable the exchange of additional information which is not available for exchange within the currently deployed EV/EVSE communications protocols. This information includes co-identification parameters, error message exchange and more. This fork of the everest-core repository adds a prototype of the Extensible Supply Equipment Communication Controller (SECC) Discovery Protocol (ESDP) implemented based on a draft of the ISO 15118-202 standard. This is achieved through additions and modifications to the EvseV2G module. The implementation provides a demonstration of the ESDP messages, encoding and decoding but does not include a full integration within the Everest framework. Much of the information being sent over ESDP in this implementation is set statically for the sake of demonstrating the protocol itself. This fork of the ext-switchev-iso15118 repository adds a prototype of the Extensible Supply Equipment Communication Controller (SECC) Discovery Protocol (ESDP) implemented based on a draft of the ISO 15118-202 standard. The implementation provides a demonstration of the ESDP messages, encoding and decoding but does not include a full integration within the Everest framework. Much of the information being sent over ESDP in this implementation is set statically for the sake of demonstrating the protocol itself. This fork adds the ESDP features for only the EVCC controller because that is the only portion that is utilized in the everest Software-in-the-Loop.

Watt, Ed [National Renewable Energy Laboratory (NR

Coherent Erbium Spin Defects in Colloidal Nanocrystal Hosts

We demonstrate nearly a microsecond of spin coherence in Er 3+ ions doped in cerium dioxide nanocrystal hosts, despite a large gyromagnetic ratio and nanometric proximity of the spin defect to the nanocrystal surface. The long spin coherence is enabled by reducing the dopant density below the instantaneous diffusion limit in a nuclear spin-free host material, reaching the limit of a single erbium spin defect per nanocrystal. We observe a large Orbach energy in a highly symmetric cubic site, further protecting the coherence in a qubit that would otherwise rapidly decohere. Spatially correlated electron spectroscopy measurements reveal the presence of Ce 3+ at the nanocrystal surface, which likely acts as extraneous paramagnetic spin noise. Even with these factors, defect-embedded nanocrystal hosts show tremendous promise for quantum sensing and quantum communication applications, with multiple avenues, including core-shell fabrication, redox tuning of oxygen vacancies, and organic surfactant modification, available to further enhance their spin coherence and functionality in the future.

77 NANOSCIENCE AND NANOTECHNOLOGY

SAIGE-GPU: accelerating genome- and phenome-wide association studies using GPUs

Genome-wide association studies (GWAS) at biobank scale are computationally intensive, especially for admixed populations requiring robust statistical models. SAIGE is a widely used method for generalized linear mixed-model GWAS but is limited by its CPU-based implementation, making phenome-wide association studies impractical for many research groups. We developed SAIGE-GPU, a GPU-accelerated version of SAIGE that replaces CPU-intensive matrix operations with GPU-optimized kernels. The core innovation is distributing genetic relationship matrix calculations across GPUs and communication layers. Applied to 2068 phenotypes from 635 969 participants in the Million Veteran Program, including diverse and admixed populations, SAIGE-GPU achieved a 5-fold speedup in mixed model fitting on supercomputing infrastructure and cloud platforms. We further optimized the variant association testing step through multi-core and multi-trait parallelization. Deployed on Google Cloud Platform and Azure, the method provided substantial cost and time savings. Source code and binaries are available for download at https://github.com/saigegit/SAIGE/tree/SAIGE-GPU-1.3.3. A code snapshot is archived at Zenodo for reproducibility (DOI: [10.5281/zenodo.17642591]). SAIGE-GPU is available in a containerized format for use across HPC and cloud environments and is implemented in R/C++ and runs on Linux systems.

Rodriguez, Alex [Argonne National Laboratory (ANL)

Bypass heat exchanger configuration to reroute core flow

A turbine engine assembly includes a heat exchanger that is in thermal communication with an exhaust gas flow. A bypass passage is provided that is configured to selectively route at least a portion of the exhaust gas flow around the heat exchanger in response to a predefined engine operating condition.

Terwilliger, Neil

Gaia: segmented germanium detector for high-energy X-ray fluorescence and spectroscopic imaging

We present Gaia, a monolithic array of 96 high-purity germanium pixel detectors integrated with a custom low-noise application-specific integrated circuit (ASIC) and a field-programmable gate array (FPGA)-based data acquisition system. The sensor operates at ∼100 K using a commercial closed-cycle cryocooler, with the in-vacuum electronics thermally isolated from the cold finger to ensure thermal stability. The system demonstrates an average energy resolution of 711 eV at 122 keV, measured using a 57 Co source, and 253 eV at 5.89 keV, measured with 55 Fe across all channels. The readout architecture incorporates a high-performance FPGA paired with a dual-core ARM processor, forming a complete embedded Linux-based computing platform. Communication between the processor and FPGA is handled via memory-mapped I/O, and data are streamed over high-speed gigabit Ethernet. A full-scale 384-pixel Gaia detector, based on this 96-element module, is currently under fabrication.

36 MATERIALS SCIENCE

Effect of heteroatom incorporation on electronic communication in metal chalcogenide nanoclusters

Metal chalcogenide nanoclusters (NC), specifically of type TM 6 E 8 (L) 6 (TM = transition metal, E = chalcogen, L = ligand) have garnered attention in recent years as promising catalysts and biosensors due to their remarkable electronic and magnetic properties, as well as their ability to undergo supramolecular assembly into 2D materials. Furthermore, the undercoordinated metal chalcogenide NCs have shown distinct surface reactivity, which is strongly dependent on the composition of the TM core. The differences in the reactivity of the undercoordinated species have been attributed to differences in ligand binding energies. Although ligand binding energies in homometallic NCs have been extensively studied, little is known about the effect of heteroatoms in the core on the strength of ligand binding in metal chalcogenide NCs. In this work, we provide new insights into this topic by examining the relative stability of [Co 6−x Fe x S 8 (PEt 3 ) 6 ] + (x = 0–6) NCs towards fragmentation using collision energy-resolved collision-induced dissociation (CID) experiments. We observe that the ligand binding energy gradually decreases until four Fe atoms are incorporated into the cluster core and then gradually increases until all the Co atoms are replaced with Fe. This experimental trend was compared with the results of density functional theory (DFT) calculations, which indicate drastic differences in the electronic communication between Co and Fe atoms in the TM core. By understanding the effect of heteroatom incorporation on ligand binding energy to the NC core, our work provides important insights into the effect of atom-by-atom substitution on the functional properties of tunable nanostructures.

Havenridge, Shana [Argonne National Laboratory (AN

QRCODE: Massively parallelized real-time time-dependent density functional theory for periodic systems

We present a new software module, QRCODE (Quantum Research for Calculating Optically Driven Excitations), for massively parallelized real-time time-dependent density functional theory (RT-TDDFT) calculations of periodic systems in the open-source Qbox software package. Our approach utilizes a custom implementation of a fast Fourier transformation scheme that significantly reduces inter-node message passing interface (MPI) communication of the major computational kernel and shows impressive scaling up to 16,344 CPU cores. In addition to improving computational performance, QRCODE contains a suite of various time propagators for accurate RT-TDDFT calculations. As benchmark applications of QRCODE, we calculate the current density and optical absorption spectra of hexagonal boron nitride (h-BN) and photo-driven reaction dynamics of the ozone-oxygen reaction. We also calculate the second and higher harmonic generation of monolayer and multi-layer boron nitride structures as examples of large material systems. Our optimized implementation of RT-TDDFT in QRCODE enables large-scale calculations of real-time electron dynamics of chemical and material systems with enhanced computational performance and impressive scaling across several thousand CPU cores.

97 MATHEMATICS AND COMPUTING

OmniFed: A Modular Framework for Configurable Federated Learning from Edge to HPC

Federated Learning (FL) is critical for edge and High Performance Computing (HPC) where data is not centralized and privacy is crucial. We present OmniFed, a modular framework designed around decoupling and clear separation of concerns for configuration, orchestration, communication, and training logic. Its architecture supports configuration-driven prototyping and code-level override-what-you-need customization. We also support different topologies, mixed communication protocols within a single deployment, and popular training algorithms. It also offers optional privacy mechanisms including Differential Privacy (DP), Homomorphic Encryption (HE), and Secure Aggregation (SA), as well as compression strategies. These capabilities are exposed through well-defined extension points, allowing users to customize topology and orchestration, learning logic, and privacy/compression plugins, all while preserving the integrity of the core system. We evaluate multiple models and algorithms to measure various performance metrics. By unifying topology configuration, mixed-protocol communication, and pluggable modules in one stack, OmniFed streamlines FL deployment across heterogeneous environments. Github repository is available at https://github.com/at-aaims/OmniFed.

Tyagi, Sahil [ORNL] (ORCID:0009000783144745)

SmartFuse: Reconfigurable Smart Switches to Accelerate Fused Collectives in HPC Applications

Communication switches have sometimes been augmented to process collectives (e.g., the IBM BlueGene project and the Mellanox SHArP switch). In this work, we find that there is a great acceleration opportunity through the further augmentation of switches to accelerate more complex functions that combine communication with computation. We consider three types of such functions. The first is fully-fused collectives built by fusing multiple existing collectives like Allreduce with Alltoall. The second is semi-fused collectives built by combining a collective with another computation. The third we refer to as higher-order collectives built by combining multiple computations and communications, such as to perform matrix-matrix multiply (PGEMM). In this work, we propose a framework called SmartFuse to accelerate fused collective functions. The core of SmartFuse is a reconfigurable smart switch to support these operations. The semi/fully fused collectives are implemented with a CGRAlike architecture, while higher-order collectives are implemented with a more specialized computational unit that can also schedule communication. Supporting our framework is software to evaluate and translate relevant parts of the input program, compile them into a control data flow graph, and then map this graph to the switch hardware. The proposed framework, once deployed, has the strong potential to accelerate existing HPC applications transparently by encapsulation within an MPI implementation. Experimental results show that this approach improves the performance of the PGEMM kernel, MINIFE, and AMG by, on average, 94%, 15%, and 13%, respectively.

Haghi, Pouya

Low spontaneous Brillouin scattering in anti-resonant hollow-core fibers in GHz frequency range

Brillouin light scattering (BLS) is a powerful experimental tool that can be used to gain insights into the fundamental and applied properties of matter, like dispersions of quasiparticles in a solid, as well as their spatiotemporal dynamics. Many applications of light scattering favor the use of optical fibers in place of free-space optics. In this study, we compare the performance of anti-resonant hollow core fibers to that of conventional solid core fused silica fibers for BLS experiments in the GHz frequency range. Conventional fibers are barely suitable for low-noise measurements because of the spontaneous scattering of photons on various phononic modes present in the core and cladding. In the case of the hollow-core fiber, we identify a range of discrete phononic modes and associate them with the various acoustic modes of the structure surrounding the hollow core using finite-element numerical simulations. The measured relative intensity of the spontaneous BLS signal from these modes is orders of magnitude smaller than that of a solid-core fiber, making anti-resonant hollow-core fibers one of the best solutions for the single-mode light guidance for BLS and potentially other low-noise photonic experiments.

36 MATERIALS SCIENCE

Board on Earth Sciences and Resources and Its Activities

The National Academies’ Board on Earth Sciences and Resources (BESR) and its standing committees provide an ongoing forum for advancing the understanding and communication of Earth sciences and resource topics, including emerging topics and innovative techniques. BESR activities help to provide evidence-based information to members of the executive and legislative branches of the federal government, the private sector, states and tribes, academia, non-governmental organizations, and the public to support decision-making. BESR and its standing Committee on Solid Earth Geophysics (COSEG), also supported by this award, fulfill this role through development and administration of consensus studies, as well as workshops and other convening activities related to the Earth sciences; overseeing selected activities of the Board’s standing committees, such as disciplinary meetings and webinars; and communicating, sharing information, and providing opportunities for interaction and exchange among technical and non-technical stakeholders. The core support received from DOE helps BESR and COSEG maintain a central body of volunteer experts and National Academies staff who can respond to pressing needs and requests from federal sponsors and other members of the Earth science community, to maintain the health and relevance of the Earth sciences discipline, and to provide the Earth science community with a privileged interface to the government to support scientific and engineering advances and decision making related to Earth sciences and engineering. Initiation and oversight of Earth science activities at the National Academies is an enduring function of the BESR and COSG that helps to ensure development and completion of projects and activities that are responsive to the needs of sponsors and the broader Earth science and research enterprise.

58 GEOSCIENCES

P38 heterogeneous multi-tiled system with support for message queues (MoSAIC) v0.1

The proposed system is written in the hardware description language (HDL) verilog targeting an FPGA board. It is intended as a testbed to explore architecture tradeoffs in multi-tiled heterogeneous architectures. Although we target FPGAs, the system can be implemented as a monolithic SoC or a package comprised of many chiplets that are interconnected in the same package using a NoC. The proposed NoC is lightweight and follows an axi-lite interface. The endpoints of the NoC are a heterogeneous mix of "tiles" as endpoints that are general purpose processors, fixed function accelerators, and programmable accelerators. We assume that the network interfaces for the NoC endpoints are all addressable in a global name-space in that they represent an address range (for memory addresses) or a range of unique identifiers that are associated with each individual tile. This makes the functionality abstract from the standpoint of the NoC design details. Message queues offer a direct inter-processor interface between peer general purpose cores and diverse accelerators that comprise an SoC. Although they share the same NoC infrastructure for inter-tile communication within an SoC or SiP, the hardware message queues bypass the memory hierarchy and thus do not pollute the memory state or invoke the cache coherence mechanism.

Gonzalez, LouisaPatricia

Mechanical Solutions Scan Report

Power lines, poles, and towers are the backbone of the United States (U.S.) electric-power grid. These transmission and distribution networks route electricity from generator to loads. The characteristics of these routes are rapidly changing -- trending towards decentralized renewable generation, electric heating, vehicle charging, and large data-center loads. Coupled with aging infrastructure and the increased frequency of extreme weather events, there is concern about the future reliability and transmission capacity of conductors and adjacent components. This scan report seeks to provide an overview of mechanical solutions to challenges caused by extreme weather events associated with components of transmission and distribution infrastructure, including conductor heat sag, ice accumulation, wind, and wildfire. Many options could increase transmission capacity or reliability, and these are at various stages of technological readiness. Some have only been lab tested, while some have been widely deployed in the U.S. or overseas for decades. The solution categories and providers featured in this report are intended to be comprehensive at the time of publication and to serve as a reference for decision-makers concerned about transmission and distribution reliability. There are two other categories of large, complex solutions, which are not covered in this report: replacing existing conductors with advanced conductors and implementing digital grid enhancing technologies. A separate scan report titled “Advanced Conductor Scan Report,” which discusses advanced carbon-core conductors, was published by the Idaho National Laboratory (INL) in 2023. Information on digital technologies, such as dynamic line ratings, power-flow controllers, and other power electronics and communications-based devices, can be found on the Grid- Enhancing Technologies landing page. Mechanical grid-enhancing technologies, or solutions covered in this report, often do not require full equipment replacement and do not rely on digital components. Mechanical technologies are overlooked because they may be older, simpler, or seemingly “more obvious” than digital or carbon-core technologies. However, it is wise to consider mechanical solutions in a thorough evaluation of grid enhancing technology solutions.

24 - POWER TRANSMISSION AND DISTRIBUTION

Constellation: The autonomous control and data acquisition system for dynamic experimental setups

The operation of instruments and detectors in laboratory or beamline environments presents a complex challenge, requiring stable operation of multiple concurrent devices, often controlled by separate hardware and software solutions. These environments frequently undergo modifications, such as the inclusion of different auxiliary devices depending on the experiment or facility, adding further complexity. The successful management of such dynamic configurations demands a flexible and robust system capable of controlling data acquisition, monitoring experimental setups, enabling seamless reconfiguration, and integrating new devices with limited effort. This paper presents Constellation, a flexible and network-distributed control and data acquisition software framework tailored to laboratory and beamline environments, that addresses the limitations of existing solutions. The framework is designed with a focus on extensibility, providing a streamlined interface for instrument integration. It supports efficient system setup via network discovery mechanisms, promotes stability through autonomous operational features, and provides comprehensive documentation and supporting tools for operators and application developers such as controllers and logging interfaces. At the core of the architectural design is the autonomy of the individual components, called satellites, which can make independent decisions about their operation and communicate these decisions to other components. This paper introduces the design principles and framework architecture of Constellation, presents the available graphical user interfaces, shares insights from initial successful deployments, and provides an outlook on future developments and applications.

Autonomy

PID-Regulated Heating System for PIP-II Reference Line

The Proton Improvement Project-2 centers on building a new superconducting linear particle accelerator (Linac) at Fermilab. At the heart of the accelerator is the reference line, a critical system that defines the ideal path for the particle beam as it passes through magnets, RF cavities, and other beamline elements. Temperature stability is crucial for the reliable operation of RF components, such as mixers and filters. Fluctuations affect key performance parameters like conversion loss, isolation, and linearity. To mitigate any drift caused by ambient temperature changes, a heating plate assembly is utilized to maintain key components at a controlled temperature of 40°C. The system utilizes an aluminum 36”x36”x0.5” heat plate powered by a MOSFET-based control circuit, delivering approximately 460 W of thermal energy through a resistor array. Real-time temperature feedback is provided by a PT100 Resistance Temperature Detector (RTD), which interfaces with a Proportional–Integral–Derivative (PID) control algorithm to maintain closed-loop temperature regulation. The control signal actively modulates the gate voltage of an N channel MOSFET, dynamically adjusting power delivery in response to deviations from the temperature setpoint. Simulations and LTspice models validate the functionality and responsiveness of the circuit under varying conditions. The prototype has successfully demonstrated stable thermal control, paving the way for integration into the PIP-II infrastructure. The final design will feature an expanded resistor array, as well as communication with a PLC for continuous data acquisition and diagnostics. This work directly supports Fermilab’s broader mission by contributing to the stability and reliability of core accelerator systems, enhancing the precision of particle beam delivery for future physics experiments.

Mosher, Alexander [Fermilab]

Microreactor Automated Control System Test Bed Digital Architecture for Real-Time, Hardware-in-the-Loop Simulation

This work describes progress made towards the development of a real-time hardware-in-the-loop (HIL) test bed for non-nuclear testing of microreactor control schemes and failure modes. Non-nuclear testing is a crucial step in developing robust control algorithms for managing microreactor dynamics. The creation of an HIL simulation harnesses the realistic dynamics of physical analogue systems while additionally considering the challenges of variable communication delay. This collaborative effort between Oak Ridge National Laboratory and Idaho National Laboratory has resulted in a LabVIEW-based gRPC communication protocol which couples a TRANSFORM Modelica simulation of nuclear components to the ViBRANT physical hardware for realistic feedback and visual representation of control action in real time. A modular python client structure is developed to manage FMU-based Modelica simulation and real-time gRPC communication. HIL testing suggests that the modeled reactor with natural convection molten salt loop coolant configuration responds well to PID control of drum positioning for modulation of reactor core power, however, future efforts will be made to explore the added thermal inertial delay of system level control and downstream demand changes. Development of this platform with a generalized methodology provides a foundation for exploring a variety of reactor configurations and failure modes in rapid order to provide insight into the most effective avenues of study for further research and development.

McConnell, Jono [ORNL] (ORCID:0000000238984741)