Search NASA⌕ Search

SEARCH · Search NASA

Results for “High performance computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42

A dynamic 2D Borehole Thermal Energy Storage (BTES) model for enhanced computational efficiency

Progressing toward a future increasingly reliant on renewable energy sources, the development of effective, durable energy storage solutions becomes essential to balance supply and demand fluctuations. Borehole Thermal Energy Storage (BTES) is a long-duration thermal energy storage technology that captures excess heat generated from renewable energy sources and stores it underground for later use, enabling the efficient utilization of sustainable energy. This approach is particularly valuable in district energy networks when integrated with Ground Source Heat Pumps (GSHP) to provide stable heating and cooling. However, traditional three-dimensional (3D) numerical models of BTES systems demand extensive computational resources, limiting their practicality for real-time and large-scale applications. This study introduces a novel two-dimensional (2D) modeling approach that reduces computational costs while maintaining high accuracy. By employing a radial ring-based discretization method, the model simulates heat injection, retention, and retrieval dynamics over seasonal cycles. A new thermal-mass weighted-average temperature parameter is introduced to evaluate the performance of BTES systems. Model validation against FEFLOW simulations demonstrates a 17-fold improvement in computational speed compared to traditional Computational Fluid Dynamics (CFD) models while achieving a mean absolute percentage error (MAPE) of 2 % during charging and 4 % during discharging. Additionally, a trade-off analysis between computational efficiency and accuracy is conducted, ensuring the model's applicability for real-world scenarios. The findings of this research contribute to the development of computationally efficient BTES models, facilitating better optimization, control, and integration into renewable energy systems. This work provides a foundation for further studies in techno-economic analysis, multi-year performance evaluation, and real-time operational strategies for BTES applications, supporting a more sustainable energy future.

2D modeling↗

Adaptive Interface-PINNs (AdaI-PINNs) for transient diffusion: Applications to forward and inverse problems in heterogeneous media

We model transient diffusion in heterogeneous materials using a novel physics-informed neural networks framework (PINNs) termed Adaptive interface physics-informed neural networks or AdaI-PINNs (Roy et al. arXiv preprint arXiv:2406.04626, 2024). AdaI-PINNs utilize different activation functions with trainable slopes tailored to each material region within the computational domain, allowing for a fully automated and adaptive PINNs approach to model interface problems with strongly and weakly discontinuous solutions. To enhance its performance in highly heterogeneous transient diffusion systems, we prescribe a suite of robust practices, including appropriate non-dimensionalization of equations, a biased sampling method, Glorot initialization, and the hard enforcement of boundary and initial conditions. Here we evaluate the efficacy of the proposed method on several benchmark forward and inverse problems. Comparative studies on one-dimensional and two-dimensional benchmark problems reveal that the modified AdaI-PINNs outperform its unmodified counterpart, achieving root-mean-square errors that are at least two orders of magnitude better in forward problems. For inverse problems, the maximum errors in the approximated diffusion coefficients by modified AdaI-PINNs are four orders of magnitude better than those of the unmodified version. Additionally, modified AdaI-PINNs demonstrate improved stability in problems with large material mismatches.

42 ENGINEERING↗

SEC ‐ SAXS / MC Ensemble Structural Studies of the Microtubule Binding Protein Cdt1 Show Monomeric, Folded‐Over Conformations

ABSTRACT Cdt1 is a mixed folded protein critical for DNA replication licensing and it also has a “moonlighting” role at the kinetochore via direct binding to microtubules and the Ndc80 complex. However, it is unknown how the structure and conformations of Cdt1 could allow it to participate in these multiple, unique sets of protein complexes. While robust methods exist to study entirely folded or unfolded proteins, structure–function studies of combined, mixed folded/disordered proteins remain challenging. In this work, we employ orthogonal biophysical and computational techniques to provide structural characterization of mitosis‐competent human Cdt1. Thermal stability analyses shows that both folded winged helix domains1 are unstable. CD and NMR show that the N‐terminal and linker regions are intrinsically disordered. DLS shows that Cdt1 is monomeric and polydisperse, while SEC‐MALS confirms that it is monomeric at high concentrations, but without any apparent inter‐molecular self‐association. SEC‐SAXS enabled computational modeling of the protein structures. Using the program SASSIE, we performed rigid body Monte Carlo simulations to generate a conformational ensemble of structures. We observe that neither fully extended nor extremely compact Cdt1 conformations are consistent with SAXS. The best‐fit models have the N‐terminal and linker disordered regions extended into the solution and the two folded domains close to each other in apparent “folded over” conformations. We hypothesize the best‐fit Cdt1 conformations could be consistent with a function as a scaffold protein that may be sterically blocked without binding partners. Our study also provides a template for combining experimental and computational techniques to study mixed‐folded proteins.

Cell Biology↗

One‐at‐a‐Time Parameter Perturbation Ensemble of the Community Land Model, Version 5.1

Comprehensive land models are subject to significant parametric uncertainty, which can be hard to quantify due to the large number of parameters and high model computational costs. We constructed a large parameter perturbation ensemble (PPE) for the Community Land Model version 5.1 with biogeochemistry configuration (CLM5.1-BGC). We performed more than 2,000 simulations perturbing 211 parameters across six forcing scenarios. This provides an expansive data set, which can be used to identify the most influential parameters on a wide range of output variables globally, by biome, or by plant functional type. We found that parameter effects can exceed scenario effects and that a small number of parameters explains a large fraction of variance across our ensemble. The most important parameters can differ regionally and also based on the forcing scenario. The software infrastructure developed for this experiment has greatly reduced the human and computer time needed for CLM PPEs, which can facilitate routine investigation of parameter sensitivity and uncertainty, as well as automated calibration.

Kennedy, Daniel [NSF National Center for Atmospher↗

High-level hadronic tau lepton triggers of the CMS experiment in proton-proton collisions at √(s) = 13.6 TeV

The trigger system of the CMS detector is pivotal in the acquisition of data for physics measurements and searches. Studies of final states characterized by hadronic decays of tau leptons require the reconstruction and the identification of genuine tau leptons against quark- and gluon-initiated jets at the trigger level. This is a difficult task, particularly as improvements to the LHC have resulted in an increased number of interactions per bunch crossing in recent years. To address this challenge, a series of machine-learning algorithms with high identification efficiency and low computational cost have been incorporated into the high-level trigger for hadronically decaying tau leptons. In this paper, these developments and the trigger performance are summarized using data collected by the CMS experiment in proton-proton collisions at √(s) = 13.6 TeV in 2022–2023, corresponding to an integrated luminosity of 62 fb -1 .

Particle identification methods↗

Electron Cloud Simulations in the Fermilab Booster

As part of Fermilab's Proton Improvement Plan-II (PIP-II), the Fermilab Booster synchrotron will operate at a higher intensity, increasing from 4.5×1012 to 6.7×1012 protons per pulse [ppp]. A potential challenge for achieving high-intensity performance arises from rapid transverse instabilities induced by electron cloud (EC). This research presents EC simulations using PyECLOUD, which is an advanced computational tool that incorporates measurements of the secondary electron yield (SEY) from the Booster's combined function magnet material. By systematically varying beam parameters in PyECLOUD, such as bunch structure, bunch length, and intensity, the EC effects on beam stability and overall performance of Booster can be predicted.

43 PARTICLE ACCELERATORS↗

A Data-Agnostic, Continuous Machine Learning Framework for Application in High Energy Physics and Beyond: Phase 1 Final Scientific/Technical Report

This Phase 1 effort has focused on the development of continual learning frameworks for use in machine learning, specifically in the applied context of High Energy Physics (HEP). Machine learning (ML) is a transformative technology by which computers, typically through the use of neural networks, are able to perform tasks with proficiency that rivals or surpasses that of human users. Model Degradation & Catastrophic Forgetting are two undesired phenomena which can occur in ML where the performance of a model degrades when either deployed on novel data streams, or trained on novel data which are sufficiently different than the data the models were initially trained on. A natural example where these sorts of effects can be observed is in the performance of detectors in harsh environments, where the detector signature may change over the lifetime of the detector as it ages and deteriorates — precisely what occurs in the experiments conducted in HEP. Real world HEP data is therefore an excellent test-ground and use-case for Continual Learning paradigms, which are techniques used in ML to counteract these problems. Ensemble learning is one such technique, where multiple smaller models are trained on subsets of the overall data and are ensembled together during inference. The intuition behind this technique is that, although there are shifts in the distributions which govern the incoming data streams, these shifts are not expected to be homogeneous or global. If a sufficient diversity in solutions within the various sub-models has been achieved, then at least one sub-model is expected to retain its performance within the overall ensemble. One further strength of this approach is that the architectures of the various models do not need to be identical, and in fact even different modalities of data can naturally be combined in this way. This work focused on applying ensemble learning techniques to derive results using two main datasets, anomaly detection in HEP data & time-series forecasting in semiconductor manufacturing data. Semiconductor manufacturing involves data with surprising similarity to that of HEP (e.g. wafer maps look very similar to digi-occupancy maps) and Cerium Lab’s prominence within the semiconductor industry makes semiconductor manufacturing a natural opportunity for commercialization of this work. Our efforts have led to two strong results. The first is that we evaluated the proposed ensembling techniques using previously proposed machine learning architectures for use in anomaly detection, namely AutoEncoder based models and their derivatives. We also developed new architectures which have not been evaluated in this context before. In fact, this work marks the first use of Vision Transformers for anomaly detection in HEP. Second, we demonstrated that ensemble learning significantly improves model performance in scenarios prone to degradation, validating its effectiveness across both HEP and semiconductor datasets. These results further support ensemble learning as a powerful strategy for mitigating catastrophic forgetting and maintaining robust performance in evolving data environments.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Distributed Stochastic Optimization of a Neural Representation Network for Time-Space Tomography Reconstruction

4D time-space reconstruction of dynamic events or deforming objects using X-ray computed tomography (CT) is an important inverse problem in non-destructive evaluation. Conventional back-projection based reconstruction methods assume that the object remains static for the duration of several tens or hundreds of X-ray projection measurement images (reconstruction of consecutive limited-angle CT scans). However, this is an unrealistic assumption for many in-situ experiments that causes spurious artifacts and inaccurate morphological reconstructions of the object. To solve this problem, we propose to perform a 4D time-space reconstruction using a distributed implicit neural representation (DINR) network that is trained using a novel distributed stochastic training algorithm. Our DINR network learns to reconstruct the object at its output by iterative optimization of its network parameters such that the measured projection images best match the output of the CT forward measurement model. Here, we use a forward measurement model that is a function of the DINR outputs at a sparsely sampled set of continuous valued 4D object coordinates. Unlike previous neural representation architectures that forward and back propagate through dense voxel grids that sample the object's entire time-space coordinates, we only propagate through the DINR at a small subset of object coordinates in each iteration resulting in an order-of-magnitude reduction in memory and compute for training. DINR leverages distributed computation across several compute nodes and GPUs to produce high-fidelity 4D time-space reconstructions. We use both simulated parallel-beam and experimental cone-beam X-ray CT datasets to demonstrate the superior performance of our approach.

36 MATERIALS SCIENCE↗

Travelling wave‐based fault detection and location in a real low‐voltage DC microgrid

Abstract This paper discusses a device‐level implementation of a travelling wave (TW) protection device (PD) designed for a real low‐voltage DC microgrid. The TWPD fault detection and location algorithm is executed on a commercial digital signal processor (DSP) board, involving signal sampling at 1 MHz via the DSP board's analog‐to‐digital converter (ADC). The analogue input card measures positive pole, negative pole and pole‐to‐pole voltages at the TWPD location. Upon a successful fault detection using a second‐order high‐pass filter, the voltage data is normalised and multi‐resolution analysis (MRA) is performed on a 128‐sample buffer around the TW arrival time. MRA employs the discrete wavelet transform (DWT) to capture high‐frequency voltage patterns, and then the Parseval's energy theorem quantifies these TW characteristics by computing the energy of reconstructed wavelet coefficients. These energy values per decomposed frequency band are the basis for training a random forest classifier that predicts fault location and type. The TWPD is fully implemented and connected to a real DC microgrid in Albuquerque, NM, USA, for validation, and results are shown for field tests verifying the performance under faults.

Paruthiyil, Sajay Krishnan [Department of Electric↗

Capability in Theory, Modeling, and Validation for a Range of Innovative Fusion Concepts using High-Fidelity Moment-Kinetic Models

A computational modeling capability is created and available to the fusion community to understand and design lower-cost and innovative fusion concepts. The approach uses high- fidelity kinetic, moment-kinetic, and moment models and includes sophisticated plasma- boundary interactions. A majority of fusion-relevant simulations are performed with magnetohydrodynamic models and hybrid particle-in-cell codes, with limited-fidelity electron and kinetic physics. However, in fusion configurations like Z-pinches, field-reversed- configurations, plasma jet magneto-inertial fusion, spinning mirrors, and others, kinetic effects (both electron and ions) are critical to understand the physics and design scaling into the highly kinetic regime of a burning fusion plasma. Furthermore, as present fusion machines move towards a burning plasma regime, liquid-metal blankets are needed to handle first-wall heat- flux, reduce erosion, and eventually for energy conversion and fuel breeding. The work performed under this ARPA-E BETHE Capability Team advances the state-of-the-art in modeling and understanding plasma dynamics in fusion devices and its coupling with liquid-metal dynamics. These are critical areas of research for fusion energy to become realizable. To address these complex problems, we have leveraged and extended computational capabilities through the code, Gkeyll (developed jointly with Princeton Plasma Physics Laboratory and academic partners), for kinetic and moment modeling of fusion plasmas. The Concept Teams supported by this Capability Team include the Wisconsin High-field Axisymmetric Mirror (WHAM), Centrifugal Mirror Experiment (CFME), Plasma-Jet Magneto- Inertial Fusion (PJMIF), and solid and liquid wall plasma-material interaction studies relevant to a number of fusion concepts including Zap Energy’s Z-pinch. This software is open-source and available to the fusion community as a high-fidelity tool for the design of lower-cost fusion experiments. 3D gyrokinetic simulations of WHAM are now possible for long enough time scales to understand the evolution of interchange instabilities. 3D multi-fluid simulations of CMFE at higher Mach numbers are now possible for detailed design iterations with the goal of stability. The state-of-the-art in understanding shock formation and shock mitigation regimes in merging liners for PJMIF have been furthered by our kinetic simulations. Our novel models and frameworks studying plasma-material interaction by incorporating wall emission for various solid wall materials of relevance to pulsed and steady fusion concepts have advanced the state-of-the-art in our understanding of particle fluxes, heat fluxes, and other quantities at cathodes and anodes. The results from this work may explain discrepancies between experimental and theoretical predictions of achieved current densities in pulsed concepts such as Z-pinches. Another significant contribution of this Capability Team is the development and deployment of a novel experimental platform, LEX (Liquid Electrode eXperiment), at Virginia Tech to understand liquid metal free-surface response to electromagnetic pulses. The novel experiments along with model validation quantified the effect of different materials and sizes of liquid metal droplets on the radiative power balance of fusion plasmas for pulsed concepts. Furthermore, these experiments provided mitigation strategies for violent liquid metal response for high current pulses as would be expected in fusion regimes.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

2.5D Super-Resolution Approaches for X-Ray Computed Tomography-Based Inspection of Additively Manufactured Parts

X-ray computed tomography (XCT) is a key tool in non-destructive evaluation of additively manufactured (AM) parts, allowing for internal inspection and defect detection. Despite its widespread use, obtaining high-resolution CT scans can be extremely time consuming. This issue can be mitigated by performing scans at lower resolutions; however, reducing the resolution compromises spatial detail, limiting the accuracy of defect detection. Super-resolution algorithms offer a promising solution for overcoming resolution limitations in XCT reconstructions of AM parts, enabling more accurate detection of defects. While 2D super-resolution methods have demonstrated state-of-the-art performance on natural images, they tend to under-perform when directly applied to XCT slices. On the other hand, 3D super-resolution methods are computationally expensive, making them infeasible for large-scale applications. To address these challenges, we propose a 2.5D super-resolution approach tailored for XCT of AM parts. Our method enhances the resolution of individual slices by leveraging multi-slice information from neighboring 2D slices without the significant computational overhead of full 3D methods. Specifically, we use neighboring low-resolution slices to super-resolve the center slice, exploiting inter-slice spatial context while maintaining computational efficiency. This approach bridges the gap between 2D and 3D methods, offering a practical solution for high-throughput defect detection in AM parts.

Sullivan, Haley↗

A computational investigation of high-flux, plate-and-frame membrane modules for industrial carbon capture

In this work, we study the application of membrane-based separation systems for carbon capture, considering plate-and-frame membrane modules. The successful deployment of membrane CO 2 capture system relies on high-performing membranes as well as effective membrane modules that can fully exploit the developed membranes. A plate-and-frame membrane module is especially attractive for CO 2 capture from industrial flue gas due to its lower pressure drop compared to its counterparts such as spiral wound modules and hollow fiber modules. To design better plate-and-frame modules, we investigate their basic unit - a single membrane stack through a combination of computational modeling and experimental investigations. The modeling approach is based on Computational Fluid Dynamics (CFD) to represent a multiphysics problem, including the fluid flow and diffusion processes within a membrane module. We use experimental data collected under different operating conditions to validate the CFD model. Numerical results suggest a good agreement between experiments and model outputs for the CO 2 recovery, CO 2 mole fraction in the retentate and permeate, and stage-cut. The CFD model is able to predict accurately the flow behavior, providing valuable insights on the effects of fluid dynamics on mass transfer of CO 2 . We also carry out a sensitivity analysis to identify the effect of key parameters on the CO 2 recovery and the CO 2 purity of the outlet streams.

CFD simulation↗

High-Temperature Gas-Cooled Reactors Multiphysics Simulation Demonstration and Code Validation

This study presents a comprehensive benchmarking and verification effort of several thermal-hydraulic and multiphysics capabilities for high-temperature gas-cooled reactor applications. The first part of this effort focuses on the running-in verification of Griffin’s multiphysics capabilities, specifically for simulating the evolution of pebble-bed reactor cores from startup to equilibrium. Since Fiscal Year 2024, improvements and enhancements have been implemented in Griffin, including simplifying the process to specify streamlines and developing the online cross-section generation capability. In the absence of validation data, code-to-code comparisons are conducted with kugelpy, showing good agreement for integral quantities like k-eff predictions and predictions for maximum power density. However, accuracy issues are noted for more detailed quantities like the spatial distribution of fission rate densities which will require further work to address. The second part of this report presents an improved System Analysis Module (SAM) core channel model where the effects of cross flow are considered during the pressurized loss of forced cooling transient, resulting in an improved agreement of the predicted pebble temperature with respect to the predictions from the SAM 2D porous media model. Additionally, the wall channeling effect due to variable porosity at the near wall region of the core is also investigated. Furthermore, to demonstrate Griffin’s online cross-section generation capability, a Multiphysics simulation is performed by coupling Griffin to the SAM core channel model. In the third part of the report, as a part of the Organisation for Economic Co-operation and Development/Nuclear Energy Agency (OECD/NEA) thermal-hydraulic code validation benchmark activity for a high-temperature gas-cooled reactor, the High Temperature Test Facility (HTTF) is investigated first using the NekRS computational fluid dynamics (CFD) code to study the flow mixing phenomenon in the lower plenum of the facility. Then, code-to-code and code-to-data comparisons are performed for Test PG27, which is a pressurized conduction cooldown (PCC) test, using five different codes by six organizations from five countries. The different simulations show good agreements in terms of the general trend but there are differences in some results such as the peak temperatures of different regions and heat removal rate.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Impact of anisotropy on TRISO fuel performance

Manufacturing of tristructural isotropic (TRISO) particles involves the deposition of pyrolytic carbon (PyC) and silicon carbide (SiC) layers using the fluidized bed chemical vapor deposition (CVD) process. The CVD process is known to generate polycrystalline layers with crystallographic textures, which imparts anisotropic thermophysical properties to the layers. Past studies have shown the risk for particle failure increases with an increase in anisotropy. The limit beyond which the anisotropy of PyC layers becomes unacceptable due to failure risk has been identified as a high-priority knowledge gap. This work presents a first systematic study on the effects of anisotropic thermal and mechanical properties on TRISO fuel performance. This computational study, performed using the fuel performance code BISON, investigates how the anisotropy in elasticity and thermal properties affect the stresses, temperature, and failure of a TRISO particle. The influence of other factors, such as operating temperature and particle geometry on the anisotropy effects, also has been analyzed. The studies utilize the recently published anisotropic elasticity and thermal behavior models for TRISO PyC and SiC layers implemented using tensors with full anisotropic capability. The spherical TRISO particles with anisotropic properties were found to have greater maximum tensile stress and significantly higher failure probability than the spherical particles with isotropic properties. In conclusion, the fuel performance predicted using these recently developed models was found to be comparable with the performance obtained using the historical models.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

An improved guess for the variational calculation of charge-transfer excitations in large systems

Ab initio quantum-chemical methods that perform well for computing the electronic ground state are not straightforwardly transferable to electronically excited states, particularly in large molecular systems. Wave function theory offers high accuracy, but is often prohibitively expensive. Methods based on time-dependent density functional theory (TD-DFT) are crucially sensitive to the chosen exchange-correlation functional (XCF) parameterization, and system-specific tuning protocols were therefore proposed to address the method's robustness. Methods based on the variational relaxation of the excited-state electron density showcased promising results for the calculation of charge-transfer excitations, but the complex shape of the electronic hypersurface makes convergence to a specific excited state much more difficult than for the ground state when standard variational techniques are applied. We address the latter aspect by providing suitable initial guesses, which we obtain by two separate constrained algorithms. Combined with the squared-gradient minimization algorithm for all-electrons relaxation in a freeze-and-release scheme (FRZ-SGM), we demonstrate that orbital-optimized density functional theory (OO-DFT) calculations can reliably converge to the charge-transfer states of interest even for large molecular systems. We test the FRZ-SGM method on a phenothiazine-anthraquinone CT excitation in a supramolecular Pd(II) coordination cage complex as a function of the cage conformation. This compound has been studied experimentally prior to our work. We compare this freeze-and-release scheme to two XCF reparameterizations, which were recently proposed as low-cost TD-DFT-based alternatives to variational methods. Two dye-semiconductor complexes, which were previously investigated in the context of photovoltaic applications, serve as a second example to investigate the convergence and stability of the FRZ-SGM approach. Our results demonstrate that FRZ-SGM provides reliable convergence for charge-transfer excited states and avoids variational collapse to lower-lying electronic states, whereas time-dependent DFT calculations with an adequate tuning procedure for the range-separation parameter provide a computationally efficient initial estimate of the corresponding energies, with a computational cost comparable to that of configuration-interaction singles (CIS) calculations.

Bogo, Nicola↗

Application of Fuel Depletion Chain Simplification to Experiment Analysis in the Advanced Test Reactor

An irradiation experiment analysis can be informed by high-fidelity reactor engineering depletion results, but this comes at a computational cost. Applying depletion chain simplification to the advanced test reactor driver fuel before performing experiment depletions permits their programmatic parameters to be calculated faster, with a small penalty to accuracy. Here, this work contrasts the results of two irradiation experiments with different neutronic characteristics. Overall, the simplified nuclide library produced using a simple one-group microscopic cross-section library for a pressurized water reactor in the depletion chain simplification process performed comparably in terms of accuracy and runtime to the simplified nuclide library produced using a three-group microscopic cross-section library generated specifically for the advanced test reactor experiments being modeled. This is attributed to the additional nuclides and transmutation pathways preserved in the one-group cross-section library, which has data for 297 nuclides, compared to the three-group cross-section library, which has data for 217 nuclides. This indicates that a cross-section library with more nuclides is better than a cross-section library with fewer nuclides for the depletion chain simplification process, even if the cross-section library with fewer nuclides better represents the flux spectrum of the system being considered.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Real-Time GPU-Accelerated OFDR With an Integrated Auxiliary Interferometer

A GPU-accelerated optical frequency domain reflectometry (OFDR) system with an improved integrated auxiliary interferometer is proposed. Unlike conventional approaches that require separate auxiliary interferometers and multiple detection channels, the proposed OFDR system embeds this functionality directly into the signal via an intentional beat component. This enables self-calibration of laser nonlinearity while maintaining a cost-effective hardware configuration. Building on this simplified configuration, the system leverages GPU acceleration with an NVIDIA RTX 4070 Ti to achieve real-time performance, delivering high-throughput signal processing for continuous OFDR interrogation. The signal processing pipeline comprises signal capture, resampling for nonlinearity compensation, and frequency shift computation, all optimized for parallel execution. Hardware benchmarking demonstrates substantial acceleration over CPU implementations, achieving up to a 45× speedup for resampling and frequency shift computations and enabling processing latencies below 30 ms. Thermal response validation is conducted under two complementary scenarios: localized heating using a water bath and cryogenic-temperature conditions using liquid nitrogen. Under localized heating, the system achieves an accuracy of 0.249 °C with a thermal sensitivity of 5.971 GHz/°C, while cryogenic-temperature validation demonstrates a frequency shift response with a sensitivity of 2.383 GHz/°C and an accuracy of 2.04 °C. The high acceleration of the proposed GPU-accelerated OFDR system and its accuracy are achieved by exploiting CUDA-based stride indexing, enabling efficient parallel segmentation and processing of large datasets without additional memory copies. The benchmarking results confirm the robustness, accuracy, and deployability of the proposed OFDR system across a wide temperature range, establishing it as a practical platform for real-time distributed fiber sensing in structurally dynamic environments.

Harb, Salah [Lawrence Berkeley National Laboratory↗

ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model Training

Large Language Models (LLMs) have demonstrated remarkable performance in various natural language processing tasks. However, the training of these models is computationally intensive and susceptible to faults, particularly in the attention mechanism, which is a critical component of transformer-based LLMs. In this paper, we investigate the impact of faults on LLM training, focusing on INF, NaN, and near-INF values in the computation results with systematic fault injection experiments. We observe the propagation patterns of these errors, which can trigger non-trainable states in the model and disrupt training, forcing the procedure to load from checkpoints. To mitigate the impact of these faults, we propose ATTNChecker, the first Algorithm-Based Fault Tolerance (ABFT) technique tailored for the attention mechanism in LLMs. ATTNChecker is designed based on fault propagation patterns of LLM and incorporates performance optimization to adapt to both system reliability and model vulnerability while providing lightweight protection for fast LLM training. Evaluations on four LLMs show that ATTNChecker on average incurs on average 7% overhead on training while detecting and correcting all extreme errors. Compared with the state-of-the-art checkpoint/restore approach, ATTNChecker reduces recovery overhead by up to 49×.

Liang, Yuhang [University of Alabama - Birmingham]↗