Search NASA⌕ Search

Engineering topics

Bergman, Keren

Publications and source records attributed to Bergman, Keren.

Scaling comb-driven resonator-based DWDM silicon photonic links to multi-Tb/s in the multi-FSR regime

The use of chip-based micro-resonator Kerr frequency combs in conjunction with dense wavelength-division multiplexing (DWDM) enables massively parallel intensity-modulated direct-detection data transmission with low energy consumption. Resonator-based modulators and filters used in such systems can limit the number of usable wavelength channels due to practical constraints on the maximum achievable free spectral range (FSR). In this work, we introduce the design of multi-Tb/s comb-driven resonator-based silicon photonic links by leveraging the multi-FSR regime. We demonstrate the viability of the link architecture with yield estimates that are supported by extensive wafer-scale measurements of 704 micro-resonators fabricated in a commercial complementary metal–oxide–semiconductor foundry. We show that a 2.80 Tb/s link is realizable with a ≥6 σ yield (∼99.999%), and that aggregate bandwidths of 3.76 Tb/s and 4.72 Tb/s are possible if yield targets are relaxed (3 σ and 1 σ , respectively). All designs represent a 1.94−3.28× boost to aggregate link bandwidth while maintaining BER≤10 −10 performance, with a theoretical bandwidth of 10.51 Tb/s being possible for sufficiently robust resonators. We use high-speed BER measurements to inform co-optimization of data rate and aggressor spacing ( λ ag ), limiting any additional loss-based power penalties to off-resonance insertion loss (IL) and routing loss. This work demonstrates that, through the multi-FSR regime, there is a clear path toward Kerr comb-driven ultra-broadband, high bandwidth silicon photonic links that can support next-generation data centers and high-performance computers.

James, Aneek (ORCID:0000000262527807)↗

Fabrication-robust silicon photonic devices in standard sub-micron silicon-on-insulator processes

Perturbations to the effective refractive index from nanometer-scale fabrication variations in waveguide geometry plague high index-contrast photonic platforms; this includes the ubiquitous sub-micron silicon-on-insulator (SOI) process. Such variations are particularly troublesome for phase-sensitive devices, such as interferometers and resonators, which exhibit drastic changes in performance as a result of these fabrication-induced phase errors. In this Letter, we propose and experimentally demonstrate a design methodology for dramatically reducing device sensitivity to silicon width variations. We apply this methodology to a highly phase-sensitive device, the ring-assisted Mach–Zehnder interferometer (RAMZI), and show comparable performance and footprint to state-of-the-art devices, while substantially reducing stochastic phase errors from etch variations. This decrease in sensitivity is directly realized as energy savings by significantly reducing the required corrective thermal tuning power, providing a promising path toward ultra-energy-efficient large-scale silicon photonic circuits.

Rizzo, Anthony (ORCID:000000034752797X)↗

Performance trade-offs in reconfigurable networks for HPC

Designing efficient interconnects to support high-bandwidth and low-latency communication is critical toward realizing high performance computing (HPC) and data center (DC) systems in the exascale era. At extreme computing scales, providing the requisite bandwidth through overprovisioning becomes impractical. These challenges have motivated studies exploring reconfigurable network architectures that can adapt to traffic patterns at runtime using optical circuit switching. Despite the plethora of proposed architectures, surprisingly little is known about the relative performances and trade-offs among different reconfigurable network designs. We aim to bridge this gap by tackling two key issues in reconfigurable network design. First, we study how cost, power consumption, network performance, and scalability vary based on optical circuit switch (OCS) placement in the physical topology. Specifically, we consider two classes of reconfigurable architectures: one that places OCSs between top-of-rack (ToR) switches—ToR-reconfigurable networks (TRNs)—and one that places OCSs between pods of racks—pod-reconfigurable networks (PRNs). Second, we tackle the effects of reconfiguration frequency on network performance. Our results, based on network simulations driven by real HPC and DC workloads, show that while TRNs are optimized for low fan-out communication patterns, they are less suited for carrying high fan-out workloads. PRNs exhibit better overall trade-off, capable of performing comparably to a fully non-blocking fat tree for low fan-out workloads, and significantly outperform TRNs for high fan-out communication patterns.

Teh, Min Yee↗

A Case For Intra-rack Resource Disaggregation in HPC

The expected halt of traditional technology scaling is motivating increased heterogeneity in high-performance computing (HPC) systems with the emergence of numerous specialized accelerators. As heterogeneity increases, so does the risk of underutilizing expensive hardware resources if we preserve today’s rigid node configuration and reservation strategies. This has sparked interest in resource disaggregation to enable finer-grain allocation of hardware resources to applications. However, there is currently no data-driven study of what range of disaggregation is appropriate in HPC. To that end, we perform a detailed analysis of key metrics sampled in NERSC’s Cori, a production HPC system that executes a diverse open-science HPC workload. In addition, we profile a variety of deep-learning applications to represent an emerging workload. We show that for a rack (cabinet) configuration and applications similar to Cori, a central processing unit with intra-rack disaggregation has a 99.5% probability to find all resources it requires inside its rack. In addition, ideal intra-rack resource disaggregation in Cori could reduce memory and NIC resources by 5.36% to 69.01% and still satisfy the worst-case average rack utilization.

97 MATHEMATICS AND COMPUTING↗

Distributed deep learning training using silicon photonic switched architectures

The scaling trends of deep learning models and distributed training workloads are challenging network capacities in today’s datacenters and high-performance computing (HPC) systems. We propose a system architecture that leverages silicon photonic (SiP) switch-enabled server regrouping using bandwidth steering to tackle the challenges and accelerate distributed deep learning training. In addition, our proposed system architecture utilizes a highly integrated operating system-based SiP switch control scheme to reduce implementation complexity. To demonstrate the feasibility of our proposal, we built an experimental testbed with a SiP switch-enabled reconfigurable fat tree topology and evaluated the network performance of distributed ring all-reduce and parameter server workloads. The experimental results show up to 3.6× improvements over the static non-reconfigurable fat tree. Our large-scale simulation results show that server regrouping can deliver up to 2.3× flow throughput improvement for a 2× tapered fat tree and a further 11% improvement when higher-layer bandwidth steering is employed. The collective results show the potential of integrating SiP switches into datacenters and HPC systems to accelerate distributed deep learning training.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Optically connected memory for disaggregated data centers

Recent advances in integrated photonics enable the implementation of reconfigurable, high-bandwidth, and low energy-per-bit interconnects in next-generation data centers. We propose and evaluate an Optically Connected Memory (OCM) architecture that disaggregates the main memory from the computation nodes in data centers. OCM is based on micro-ring resonators (MRRs), and it does not require any modification to the DRAM memory modules. We calculate energy consumption from real photonic devices and integrate them into a system simulator to evaluate performance. Here, our results show that (1) OCM is capable of interconnecting four DDR4 memory channels to a computing node using two fibers with 1.02 pJ energy-per-bit consumption and (2) OCM performs up to 5.5× faster than a disaggregated memory with 40G PCIe NIC connectors to computing nodes.

97 MATHEMATICS AND COMPUTING↗