Search NASASearch

SEARCH · Search NASA

Results for “bandwidth efficiency”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

On-chip probabilistic inference for charged-particle tracking at the sensor edge

Modern scientific instruments operate under increasingly extreme constraints on bandwidth, latency, and power. Inference at the sensor edge determines experimental data collection efficiency by deciding which information to save for further analysis. Particle tracking detectors at the Large Hadron Collider exemplify this challenge: pixelated silicon sensors generate rich spatiotemporal ionization patterns, yet most of this information is discarded due to data-rate limitations. Concurrently, advancements in co-design tools provide rapid turn-around for incorporating machine learning into application-specific integrated circuits, motivating designs for particle detectors with new integrated technologies. We demonstrate that neural networks embedded in the front-end electronics can infer charged-particle kinematic parameters from a single silicon layer. We regress hit positions and incident angles with calibrated uncertainties, while satisfying stringent constraints on numerical precision, latency, and silicon area. Our results establish a path toward probabilistic inference directly at the edge, opening new opportunities for intelligent sensing in high-rate scientific instruments.

Das, Arghya Ranjan [Purdue U.] (ORCID:000000018451

A Parallel Alternative for Energy-Efficient Neural Network Training and Inferencing

Energy efficiency of training and inferencing with large neural network models is a critical challenge facing the future of sustainable large-scale machine learning workloads. This paper introduces an alternative strategy, called phantom parallelism, to minimize the net energy consumption of traditional tensor (model) parallelism, the most energy-inefficient component of large neural network training. The approach is presented in the context of feed-forward network architectures as a preliminary, but comprehensive, proof-of-principle study of the proposed methodology. We derive new forward and backward propagation operators for phantom parallelism, implement them as custom autograd operations within an end-to-end phantom parallel training pipeline and compare its parallel performance and energy-efficiency against those of conventional tensor parallel training pipelines. Formal analyses that predict lower bandwidth and FLOP counts are presented with supporting empirical results on up to 256 GPUs that corroborate these gains. Experiments are shown to deliver ∼50% reduction in the energy consumed to train FFNs using the proposed phantom parallel approach when compared with conventional tensor parallel methods. Additionally, the proposed approach is shown to train smaller phantom models to the same model loss on smaller GPU counts as larger tensor parallel models on larger GPU counts offering the possibility for even greater energy savings.

Seal, Sudip [ORNL] (ORCID:0000000332330656)

Large-scale real-time signal processing in physics experiments: the ALICE TPC FPGA pipeline

For LHC Run 3, the ALICE Time Projection Chamber was upgraded to operate in continuous readout mode. Interaction rates of up to 50 kHz in Pb-Pb collisions require real-time processing of more than 3 TB s -1 of raw detector data. This requirement is met by a custom FPGA-based processing pipeline that performs the complete front-end data treatment fully in-stream, including common-mode correction, pedestal subtraction, ion-tail filtering, zero suppression, and dense data packing. A central element of the design is a highly parallel common-mode correction algorithm operating directly on the streaming data. It robustly identifies signal-free readout channels on a time-bin basis and applies pad-dependent scaling to compensate for local variations in capacitive coupling in the GEM readout. In combination with pedestal subtraction and ion-tail filtering, this enables accurate baseline restoration under extreme high-occupancy conditions, preventing signal loss while efficiently suppressing noise prior to zero suppression. The pipeline operates continuously at the full detector bandwidth and reduces the raw input rate of approximately 3 TB s -1 to about 900 GBps for Pb-Pb collisions at the target interaction rate. Overall, it represents a large-scale FPGA-based real-time signal-processing implementation for high-energy physics detector readout.

Digital signal processing (DSP)

Multiplexed color centers in a silicon photonic cavity array

Entanglement distribution is central to the modular scaling of quantum processors and establishing quantum networks. Color centers with telecom-band transitions and long spin coherence times are suitable candidates for long-distance entanglement distribution. However, high-bandwidth memory-enhanced quantum communication is limited by high-yield, scalable creation of efficient spin-photon interfaces. Here, we develop a silicon photonics platform consisting of arrays of bus-coupled cavities. The coupling to a common bus waveguide enables simultaneous access to individually addressable cavity-enhanced T center arrays. We demonstrate frequency-multiplexed operation of two T centers in separate photonic crystal cavities. In addition, we investigate the cavity enhancement of a T center through hybridized modes formed between physically distant cavities. Our results show that bus-coupled arrays of cavity-enhanced color centers could enable efficient on-chip and long-distance entanglement distribution.

Komza, Lukasz

HPC Campaign Management: Remote data access with user-defined error bound using ADIOS and ZFP

Remote access to large-scale scientific datasets, like those generated by combustion simulations or other high-performance computing (HPC) applications, presents a significant challenge. Downloading entire datasets is often impractical due to their size and the bandwidth limitations of typical networks. To address this challenge, we propose a novel approach that enables efficient remote access to large datasets distributed across multiple facilities. Our method enables technologies to download only the data values of a select variable, in a select region of interest, to a user-defined accuracy. For this purpose, we extended the ADIOS IO library to provide read functions with user-defined accuracy, a remote data server that understands multidimensional selections of specific variables, steps and accuracy from an ADIOS dataset, and which uses lossy compression on the remote site to reduce the data to be transferred back to the client. In addition, our extension of the ADIOS library collects metadata from multiple datasets in small files called Campaign Archives, which can be shared among project participants on any HPC, cloud or laptop, and which can easily facilitate the discovery of content and pointers to the data location as well as remote access to the data by local tools as if data was local. This feature called Campaign Management, enables a group of scientists to manage related datasets stored in multiple files, across multiple facilities as if it was in a single file/database. We demonstrate the effectiveness of our approach using a 1.5 TB dataset from the S3D combustion simulation on Frontier at the Oak Ridge Leadership Facility. Even a single variable from this dataset, at 64 GB, is too large to be processed on a standard laptop. We show two different reading patterns for 2D plots and 3D visualization, with careful settings that a scientist studying combustion data would do and show that running the same Python scripts on Frontier directly takes comparable time than running them on the local laptop with remote access to the data on Frontier.

Podhorszki, Norbert [ORNL] (ORCID:000000019647542X

A compact and portable gamma-ray spectrometer (GRASP) for inertial confinement fusion and basic science experiments

A compact and portable gamma-ray spectrometer has been designed to diagnose different components of the inertial confinement fusion-relevant γ-ray spectrum with energies between ∼3.7–17.9 MeV. The system is designed to be as compact as possible for convenient transportation and fielding in diagnostic ports on the OMEGA laser, the National Ignition Facility, and other photon-source facilities. The system consists of a conversion foil for Compton scattering in front of four magnetic spectrometer “arms,” each covering a different energy range and constructed out of cylindrical permanent magnet Halbach arrays. Monte Carlo simulations have been used to optimize and assess the performance of the conversion foil, and COSY INFINITY ion-optical simulations have been used to optimize the spectrometer magnets. The performance of the design is assessed for a simulated direct-drive γ-ray spectrum. Spanning its total γ-ray energy bandwidth and using a 1.7 mm thick boron conversion foil, the system’s total energy resolution and efficiency are ∼15.8%–4.5% and 5.4 × 10−7–3.7 × 10−7e−/γ, respectively, with room for improvement. Spectral γ-ray measurements will provide guidance to the inertial confinement fusion program toward achieving high-energy gain relevant to inertial fusion energy and enable new measurement capabilities for basic discovery science.

Instruments & Instrumentation

ExtremeMETA: High-speed Lightweight Image Segmentation Model by Remodeling Multi-channel Metamaterial Imagers

Deep neural networks (DNNs) have heavily relied on traditional computational units, such as CPUs and GPUs. However, this conventional approach brings significant computational burden, latency issues, and high power consumption, limiting their effectiveness. This has sparked the need for lightweight networks such as ExtremeC3Net. Meanwhile, there have been notable advancements in optical computational units, particularly with metamaterials, offering the exciting prospect of energy-efficient neural networks operating at the speed of light. Yet, the digital design of metamaterial neural networks (MNNs) faces precision, noise, and bandwidth challenges, limiting their application to intuitive tasks and low-resolution images. In this study, we proposed a large kernel lightweight segmentation model, ExtremeMETA. Based on ExtremeC3Net, our proposed model, ExtremeMETA maximized the ability of the first convolution layer by exploring a larger convolution kernel and multiple processing paths. With the large kernel convolution model, we extended the optic neural network application boundary to the segmentation task. To further lighten the computation burden of the digital processing part, a set of model compression methods was applied to improve model efficiency in the inference stage. The experimental results on three publicly available datasets demonstrated that the optimized efficient design improved segmentation performance from 92.45 to 95.97 on mIoU while reducing computational FLOPs from 461.07 MMacs to 166.03 MMacs. The large kernel lightweight model ExtremeMETA showcased the hybrid design’s ability on complex tasks.

large convolution kernel

Optimal Control Strategy With Efficiency and Reliability Improvement for Offshore DC Microgrids

Offshore microgrids, due to their remote location and lack of external energy support, face significant challenges in wide-range load operation and maintenance. Consequently, efficiency and reliability are critical concerns for converters in offshore dc microgrids. This article presents an optimal control strategy aimed at enhancing both efficiency and reliability. A normalized nonlinear relationship between power loss and thermal stress of a paralleled converter is first established. Based on this, a dual-objective optimization function with an active weight function as well as a system overall performance index is established. The active weight function dynamically adjusts the control priority based on converter efficiency and switching device thermal stress. Then, the optimal power-sharing strategy is derived by the Lagrange multiplier method with the proposed optimal function. Additionally, to accommodate a wide load range, an optimal selection strategy for operating converter combinations is proposed, requiring only low-bandwidth communication. Experiment verification is given to validate the effectiveness of the proposed control strategy. The experiment results demonstrate that the proposed control strategy can improve the overall performance of offshore microgrids by optimizing efficiency and reliability.

24 POWER TRANSMISSION AND DISTRIBUTION

Impact of quantum well thickness on efficiency loss in InGaN/GaN LEDs: Challenges for thin-well designs

We investigate the impact of quantum well (QW) thickness on efficiency loss in c-plane InGaN/GaN LEDs using a small-signal electroluminescence technique. Multiple mechanisms related to efficiency loss are independently examined, including injection efficiency, carrier density vs current density relationship, phase space filling, quantum-confined Stark effect, and Coulomb enhancement. An optimal QW thickness of around 2.7 nm in these InGaN/GaN LEDs was determined for QWs having constant In composition. Despite improved control of deep-level defects and lower carrier density at a given current density, LEDs with thin QWs still suffer from an imbalance in enhancement effects on the radiative and intrinsic Auger–Meitner recombination coefficients. The imbalance in enhancement effects results in a decline in internal quantum efficiency and radiative efficiency with decreasing QW thickness at low current density in LEDs with QW thicknesses below 2.7 nm. Here, we also investigate how LED modulation bandwidth varies with QW thickness, identifying the key trends and their implications for device performance.

36 MATERIALS SCIENCE

In Situ MOF Pyrolysis Construction of Hierarchical Porous Co‐Nanoparticles/Carbon Cloth Composites for Enhanced Electromagnetic Wave Shielding and Absorption

The development of high-performance electromagnetic protection materials integrating broadband absorption and effective shielding capabilities is hindered by challenges in simultaneously optimizing multiple electromagnetic properties through conventional material designs. This study pioneers a hierarchical porous Co nanoparticle/carbon cloth (Co/CC) composite via controlled annealing of a Co-MOF precursor on carbon cloth. The Co-MOF served a dual role as both magnetic source and pore-forming agent, enabling in situ generation of uniformly dispersed Co nanoparticles and creation of abundant pores/interfaces on the CC fibers during pyrolysis. This unique architecture synergistically enhanced dielectric loss (via interfacial/dipolar polarization) and magnetic loss (via natural resonance, exchange interactions, and eddy currents), significantly improving impedance matching. The hierarchical pores further functioned as integrated “absorption–reflection” units for efficient electromagnetic energy attenuation. Consequently, the Co/CC composite annealed at 800°C (Co/CC-800) achieves minimum reflection loss (−40.69 dB) and 120% effective absorption bandwidth extension (6.16 GHz) as a filler, and exhibits superior electromagnetic interference shielding effectiveness (46.66 dB) as an integrated component. Significantly, Co/CC-800 demonstrated robust photothermal and electrothermal conversion capabilities, ensuring operational stability in ice-covered and humid harsh environments. This work pioneers a pore-structure-mediated strategy to harmonize dielectric–magnetic synergy, providing a new paradigm for designing advanced multifunctional electromagnetic protection materials.

dielectric‐magnetic synergy

Fast and Scalable FFT-Based GPU-Accelerated Algorithms for Block-Triangular Toeplitz Matrices with Application to Linear Inverse Problems Governed by Autonomous Dynamical Systems

In this work, we present an efficient and scalable algorithm for performing matrix-vector multiplications (matvecs) for block Toeplitz matrices. Such matrices, which are shift-invariant with respect to their blocks, arise in the context of solving inverse problems governed by autonomous systems, and time-invariant systems in particular. In this article, we consider inverse problems that infer unknown parameters from observational data of a linear time-invariant dynamical system given in the form of partial differential equations (PDEs). Matrix-free Newton-conjugate-gradient methods are often the gold standard for solving these inverse problems, but they require numerous actions of the Hessian on a vector. Matrix-free adjoint-based Hessian matvecs require solution of a pair of linearized forward/adjoint PDE solves per Hessian action, which may be prohibitive for large-scale inverse problems. Time invariance of the forward PDE problem leads to a block Toeplitz structure of the discretized parameter-to-observable (p2o) map defining the mapping from inputs (parameters) to outputs (observables) of the PDEs. This block Toeplitz structure enables us to exploit two key properties: (1) compact storage of the p2o map and its adjoint, and (2) efficient fast Fourier transform–based Hessian matvecs. The proposed algorithm is mapped onto large multi-GPU clusters and achieves more than 80% of peak bandwidth on NVIDIA A100 GPUs. Excellent weak scaling is shown for up to 48 A100 GPUs. For the targeted problems, the implementation executes Hessian matvecs within fractions of a second, which is orders of magnitude faster than can be achieved by conventional matrix-free Hessian matvecs via forward/adjoint PDE solves.

97 MATHEMATICS AND COMPUTING

FedCSpc: A Cross-Silo Federated Learning System With Error-Bounded Lossy Parameter Compression

Cross-Silo federated learning is widely used for scaling deep neural network (DNN) training over data silos from different locations worldwide while guaranteeing data privacy. Communication has been identified as the main bottleneck when training large-scale models due to large-volume model parameters and gradient transmission across public networks with limited bandwidth. Most previous works focus on gradient compression, while limited work tries to compress parameters that can not be ignored and extremely affect communication performance during the training. Here, to bridge this gap, we propose FedCSpc: an efficient cross-silo federated learning system with an XAI-driven adaptive parameter compression strategy for large-scale model training. Our work substantially differs from existing gradient compression techniques due to the distinct data features of gradient and parameter. The key contributions of this paper are fourfold. (1) Our designed FedCSpc proposes to compress the parameter during the training using the state-of-the-art error-bounded lossy compressor – SZ3. (2) We develop an adaptive compression error bound adjustment algorithm to guarantee the model accuracy effectively. (3) We exploit an efficient approach to utilize the idle CPU resources of clients to compress the parameters. (4) We perform a comprehensive evaluation with a wide range of models and benchmarks on a GPU cluster with 65 GPUs. Results show that FedCSpc can achieve the same model accuracy as FedAvg while reducing the data volume of parameters and gradients in communication by up to 7.39× and 288×, respectively. With 32 clients on a 4 Gb size model, FedCSpc significantly outperforms FedAvg in wall-clock time in the emulated WAN environment (at the bandwidth of 1 Gbps or lower without loss of generality).

SZ3

Femtosecond wavelength-tunable laser system using gain managed nonlinear amplifier

Optical parametric amplifiers require a seed energy at the spectral range of interest for amplification. In the case of pulsed laser systems, seed pulse characteristics such as spectral energy density and spectral uniformity influence laser system design and performance. In this letter, we present the design and modeling of a few micro-joules, wavelength-tunable, femtosecond parametric amplifier system. We employ a gain-managed nonlinear fiber amplifier output as the seed pulse. This seed pulse is relatively uniform across its bandwidth from 1010 to 1180 nm. Its spectral energy density averages 600 pJ/nm, 30 times higher than conventional counterparts. The overall amplification gain is 175, and the conversion efficiency is 17.5%. The inherent quasi-linear chirp of the seed pulse is utilized to alter the output wavelength. Our tunable femtosecond laser is an enabling technology for high-demand applications such as time-resolved spectroscopy, nonlinear and advanced material research, and biomedical imaging and microscopy modalities.

47 OTHER INSTRUMENTATION

Pushing the limits of NAND technology scaling with ferroelectrics

Artificial intelligence (AI) continues to drive transformative advancements across various industries. The data-intensive nature of AI training (and inferencing) has resulted in the generation of unprecedented volumes of data with machine-generated content surpassing human-generated data by more than 100-fold in 2025. Efficiently managing this data influx necessitates advanced digital storage technologies. However, traditional NAND flash memory, which is critical for supporting data flows in AI systems—alongside high-bandwidth memory, for AI training—faces fundamental scaling limitations as it approaches the 1000-layer milestone, encompassing more than 40 trillion transistors. This article delves into the potential of hafnia-based ferroelectric materials as a breakthrough solution to these challenges. Recent advancements indicate that the intrinsic limitations of ferroelectric field-effect transistors (FEFETs) can be mitigated through material and device-level engineering. These advancements enable FEFETs to meet the stringent density, reliability, and scalability requirements of future three-dimensional NAND technology. The role of ferroelectrics in addressing NAND scaling challenges and expanding storage capabilities presents a promising avenue for meeting the storage demands of the AI-driven era.

3D NAND

Stacked-mosaic amplifier and diode delivery concept for kJ-class inertial fusion energy laser driver

We report on a laser amplifier architecture designed as a modular element for scaling to multi-megajoule laser facilities intended for inertial fusion energy (IFE) power plants. This kilojoule-class module features a stacked mosaic of gain media integrated with a diode delivery system for efficient optical pumping of the mosaic-structured gain medium, enabling high-repetition-rate operation at high wall plug efficiencies, necessary for an IFE driver. Details of a diode delivery system capable of pumping the mosaic architecture are presented. Comprehensive numerical modeling demonstrates that the stacked-mosaic approach reduces transverse gain by five orders of magnitude compared to a conventional full-aperture Yb:YAG based amplifier, substantially suppressing transverse amplified spontaneous emission (TASE) and enabling enhanced longitudinal energy extraction. Detailed analysis of thermal management and wavefront distortion in a 4 × 4 mosaic array indicates that temperature gradients and thermally induced aberrations are effectively controlled using gas cooling and commercially available phase-plate and deformable mirror technologies. We further discuss the applicability of a stacked-mosaic architecture to direct-drive IFE schemes, where broad spectral bandwidth is critical, and its compatibility with frequency conversion modules for up-conversion to blue wavelengths. Finally, an example point design for a 10 kJ, 10 Hz Yb:YAG module operating at 175 K with wall-plug efficiency exceeding 10 % is presented, underscoring the feasibility of this approach for next-generation high-energy lasers for IFE drivers. The results establish the stacked-mosaic amplifier as a scalable, robust platform not only for IFE but also for a broad range of advanced scientific and industrial laser applications.

Lasers

Passive radiative thermal management using phase-change metasurfaces

Abstract Realizing innovative composite materials with passive thermal management capabilities and minimal ecological footprints is a challenging but much sought-after goal that would have a transformative effect on renewable energy sciences. We demonstrate an environmentally friendly metasurface utilizing vanadium dioxide (VO 2 ) that offers responsiveness to ambient temperature and potentially long-term stability. The metasurface enables passive thermal management by self-adjusting its absorptivity and emissivity response over a broad bandwidth ranging from visible to mid-infrared (IR) wavelengths. Above the VO 2 phase transition the metasurface exhibits increased mid-IR emissivity and reduced visible/near-IR absorptivity, creating an efficient radiative emission channel in the first atmospheric transparency window with reduced absorption of solar radiation. In contrast, below VO 2 ’s transition temperature, the metasurface increasingly absorbs sun light while minimizing mid-IR radiative heat losses. Moreover, a functional silicon layer eliminates the need for an additional capping layer commonly employed to protect VO 2 from environmental degradation. The additional protective layer often impedes the use and performance of VO 2 based devices in terrestrial as well as spacecraft applications. Therefore, the proposed durable and eco-friendly metasurface will be an excellent candidate for essential passive thermal regulation systems across residential and terrestrial applications.

42 ENGINEERING

Low power on-chip data transmission for wafer-scale monolithic active pixel sensors

Here, this paper details the implementation of the digital pulse shaping subsystem within the Backbone Transmission Line Encoding (BTLE) driver, a low-power, long-distance on-chip data transmission solution designed in a 65 nm CMOS process. Digital pulse shaping is critical for minimizing inter-symbol interference (ISI) caused by bandwidth limitations of on-chip interconnects, especially in wafer-scale monolithic active pixel sensors (MAPS). A duobinary encoder coupled with a parallelized polyphase finite impulse response (FIR) filter is used for efficient shaping of the transmitted signal spectrum. This reconfigurable architecture achieves reliable 160 Mb/s data transfer over a 10 cm on-chip link, as validated by simulations demonstrating low power consumption (FoM 37.3 fJ/bit/mm of transmission line length) and effective ISI mitigation.

47 OTHER INSTRUMENTATION

Lossy Compression: An Online Multi-Stage Technology for High-Fidelity Synchro- Waveform Measurements

Effective real-time monitoring and analysis of distributed grids necessitate the use of synchro-waveform measurements, which capture almost all high-frequency disturbances and transient phenomena. However, due to limitations in high-speed measurements and network bandwidth, it is challenging to transfer all high-fidelity synchro-waveforms losslessly and successfully. To cope with these challenges, a hybrid-based online multi-stage compression algorithm is proposed to significantly improve the compression efficiency for synchro-waveform measurements. Initially, the multiple discrete Wavelet transformation is deployed to deconstruct the waveform components. The delta encoding is further developed to decrease the magnitude. In conjunction with the Lempel-Ziv-Markov chain, the hybrid compression algorithm is implemented to achieve real-time compression for the synchro-waveform measurements. Moreover, an innovative error index that synergizes the time and frequency domain error and correlation is formulated to evaluate the waveform distortion. By integrating compression ratio, suitable parameters can be optimally selected. Finally, the simulation, laboratory experiments, as well as field tests across a spectrum of sampling frequencies and time intervals are conducted to substantiate the efficacy of the proposed method. Here, the outcomes demonstrated that a compression ratio of approximately 15.5 and 17.83 can be reached for 0.5 s and 1 s data under both offline and online scenarios, which equates to a substantial 93.5% to 94.39% reduction in data storage requirements.

High-fidelity synchro-waveform measurements