Search NASA⌕ Search

SEARCH · Search NASA

Results for “generalization error”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Toward more-robust, AI-enabled subsurface seismic imaging for geotechnical applications

Non-invasive seismic imaging has the potential to cost-effectively evaluate large volumes of subsurface material to inform geotechnical site investigation. However, seismic imaging using full waveform inversion (FWI) requires significant computational time and is dependent on an initial starting model. As a result, FWI has not yet been widely adopted into geotechnical practice. Previous efforts, on relatively simple two-layered models, indicate that data-driven artificial intelligence (AI) models may be as effective as FWI at predicting 2D images of shear wave velocity (V s ). Furthermore, the AI model predictions can be made almost instantaneously after data acquisition and do not require an initial starting model. We examine the generality of these findings by developing a new AI model for subsurface seismic imaging, whereby we make several notable contributions. First, we architect a multimodal AI model that combines time- and frequency-domain representations of the seismic wavefield to predict a 50 m by 20 m subsurface image of V s . Second, we developed a new diverse dataset of 100,000 images with their corresponding seismic wavefields to train the AI model. Third, we propose four physics-informed data augmentations for data-driven seismic imaging. Fourth, we develop two prediction consistency tests to evaluate the model’s performance when the true subsurface is unknown. Our final model, which has been made publicly available, is capable of predicting a subsurface V s image from a single seismic wavefield with an average, mean absolute percent error (MAPE) of 24 %. The predictive model is applied to a field dataset and shown to be consistent with local geology and shear-wave refraction measurements from the same location.

Artificial intelligence↗

Deployment of Traditional and Hybrid Machine Learning for Critical Heat Flux Prediction in the CTF Thermal-Hydraulics Code

Critical heat flux (CHF) marks the transition from nucleate to film boiling, where heat transfer to the working fluid can rapidly deteriorate. Accurate CHF prediction is essential for efficiency, safety, and preventing equipment damage, particularly in nuclear reactors. Although widely used, empirical correlations frequently exhibit discrepancies when compared to experimental data, limiting their reliability in diverse operational conditions. Traditional machine learning (ML) approaches have demonstrated potential for CHF prediction but often suffer from limited interpretability, data scarcity, and insufficient knowledge of physical principles. Hybrid model approaches, which combine data-driven ML with base models, mitigate these concerns by incorporating prior knowledge of the domain. This study integrates an externally trained purely data-driven ML model and two hybrid models (using the Biasi and Bowring CHF correlations) within the CTF subchannel code via a custom Fortran framework. Performance was evaluated using two validation cases: a subset of the Nuclear Regulatory Commission (NRC) CHF database and the Bennett dryout experiments. In both cases, the hybrid models demonstrated significantly lower error metrics compared to conventional empirical correlations, with the best models often reducing relative error by about 5 percentage points. The pure ML model achieved comparable accuracy, outperforming the hybrid Biasi model in the NRC test case (3.3% versus 5.5% relative error) but exhibiting slightly higher error against the hybrid Bowring model in the Bennett test case (7.7% versus 6.1%). Trend analysis of error parity indicated that ML-based models reduced the tendency for CHF overprediction, improving overall accuracy. These results demonstrate that ML-based CHF models can be effectively integrated into subchannel codes and could potentially increase performance compared to conventional methods.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Experimental Considerations for Estimating Degradation in PV Modules

Carefully controlled laboratory experiments and measurements can enable the determination of acceleration factors suitable for extrapolation to durability and performance of a fielded PV module. Ideally, a single mechanism can be identified with appropriate acceleration factors for extrapolation to the field. However, even with a single mechanism, the inherent uncertainty in these factors leads to uncertainty in the extrapolation which is greater the higher the acceleration factor. This course will explain how because of the wide range of acceleration factors for a given degradation mode, utilizing acceleration factors greater than about 10x will typically lead to unacceptable uncertainty in the results. Therefore, if even just a rank ordering of materials is desired, acceleration factors must be minimized which requires a good general understanding of the scale of the different acceleration factors for the degradation mode of interest. In this tutorial we will discuss what the different purposes are for many of the accelerated stress tests used today. E.g., what is a qualification test, a highly accelerated stress test, a rank ordering test, or a service life prediction test. We will discuss how one can understand the relationship between test results and expected field performance. A single accelerated stress test condition cannot duplicate outdoor exposure for all possible degradation pathways; therefore, one must use targeted evaluation of material properties at different stress levels to determine the relevant acceleration factors and fit it to a model. We will also discuss how to interpret the results of experiments understanding what is relevant/not relevant, or not e valuated in a test. There are many common error people make in their test interpretations because they push the stress levels to be too harsh. This creates biases and can mask the relevant failure modes and mechanisms or will erroneously lead one to over design materials against things that aren't relevant. Several case studies will be presented to illustrate appropriate interpretation of accelerated stress testing results.

degradation↗

Archetype-based Redshift Estimation for the Dark Energy Spectroscopic Instrument Survey

We present a computationally efficient galaxy archetype-based redshift estimation and spectral classification method for the Dark Energy Survey Instrument (DESI) survey. The DESI survey currently relies on a redshift fitter and spectral classifier using a linear combination of principal component analysis–derived templates, which is very efficient in processing large volumes of DESI spectra within a short time frame. However, this method occasionally yields unphysical model fits for galaxies and fails to adequately absorb calibration errors that may still be occasionally visible in the reduced spectra. Our proposed approach improves upon this existing method by refitting the spectra with carefully generated physical galaxy archetypes combined with additional terms designed to absorb data reduction defects and provide more physical models to the DESI spectra. We test our method on an extensive data set derived from the survey validation (SV) and Year 1 (Y1) data of DESI. Our findings indicate that the new method delivers marginally better redshift success for SV tiles while reducing catastrophic redshift failure by 10%–30%. At the same time, results from millions of targets from the main survey show that our model has relatively higher redshift success and purity rates (0.5%–0.8% higher) for galaxy targets while having similar success for QSOs. These improvements also demonstrate that the main DESI redshift pipeline is generally robust. Additionally, it reduces the false-positive redshift estimation by 5%–40% for sky fibers. We also discuss the generic nature of our method and how it can be extended to other large spectroscopic surveys, along with possible future improvements.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Innovating the next generation of commercial smart building software

Nearly 30% of commercial building energy use is wasted due to equipment faults and HVAC controls problems. The result is increased emissions, compromised comfort and productivity, and less reliable coordination of building power needs with a clean grid. The energy impact alone represents $17 billion in potential savings. Today’s smart building software provides a robust solution to address these operational deficiencies. Energy management and information systems (EMIS) are saving up to 9% on average, with two-year paybacks. They are being incorporated into energy management processes, commissioning services, and utility programs. As effective as they are, two barriers prevent even deeper benefits; limited personnel to fix problems once they are identified, and the expense and time to manually implement changes in control systems. In partnership with the research community, the EMIS industry is developing new capabilities to overcome these barriers. Moving beyond siloed products for either fault detection and diagnostics, or optimal control, these new capabilities empower users to not only automatically identify faults, but also to push corrective action, and control improvements to their buildings. In this paper, several areas for enhancements are documented: ‘one-time’ correction of faults such as setpoints, schedules, and economizer lockouts; short-term active testing for automated proportional integral derivative (PID) loop tuning and functional testing; and continuous supervisory control for demand flexibility and year-round efficiency. Results are presented from a pair of partner implementations out of a dozen providers integrating these enhancements into their products, including field tests from across the country, and insights into operator acceptance and integration into operations and maintenance practices.

Casillas, Armando↗

Situational awareness-enhancing community-level load mapping with opportunistic machine learning

Motivated by present and forthcoming challenges in the adoption and integration of distributed renewable energy, we develop a machine learning (ML) approach that builds short-fuse mappings connecting the occasionally-unobservable true load in one target community with information-rich signals collected from relatively more instrumented reference communities. Our setting is inspired by and tailored to target communities with significant unobservable behind-the-meter solar generation, where true load (a relatively well-behaved quantity of interest to grid operators) is hard to discern during daytime due to insufficient instrumentation and/or privacy reasons, but that can be related to reference communities with low unobservable distributed variable generation or with sufficient instrumentation. The developed mapping, herein realized with Support Vector Machine regression, is built using nighttime data from all communities, when their distributed generation is low or zero. Our ML algorithm opportunistically learns to correlate signals of interest and then is operationally used the next day to shed light into target community load evolution. The mapping is subsequently rebuilt, rolling its short-fuse scope perpetually forward in time. Here, we demonstrate the efficacy of our approach on nine synthetically generated topologies and associated timeseries stemming from real-world data, on which we observe cumulative error performance that yields lower than 10% and 15% daily-averaged mean absolute percentage errors in target community load estimation on more than about 75% and 90% of days, respectively, in multiple yearly evaluations that shed light on long-term performance also under seasonal and one-off effects. The proposed ML-powered methodology can offer grid operators much-improved visibility into a previously obscure space and can also serve as an additional source of information in broader, multi-modal solar disaggregation solutions.

14 SOLAR ENERGY↗

Bootstrap current modeling in M3D-C1

Bootstrap current plays a crucial role in the equilibrium of magnetically confined plasmas, particularly in quasi-symmetric stellarators and in tokamaks, where it can represent bulk of the electric current density. Accurate modeling of this current is essential for understanding the magnetohydrodynamic (MHD) equilibrium and stability of these configurations. This study expands the modeling capabilities of M3D-C1, an extended-MHD code, by implementing self-consistent physics models for bootstrap current. It employs two analytical frameworks: a generalized Sauter model (Sauter et al. 1999 Phys. Plasmas vol. 6, no. 7, pp. 2834–2839), and a revised Sauter-like model (Redl et al. 2021 Phys. Plasmas vol. 28, no. 2, pp. 022502). The isomorphism described by Landreman et al. (2022 Phys. Rev. Lett. vol. 128, pp. 035001) is employed to apply these models to quasi-symmetric stellarators. The implementation in M3D-C1 is benchmarked against neoclassical codes, including NEO, XGCa and SFINCS, showing excellent agreement. These improvements allow M3D-C1 to self-consistently calculate the neoclassical contributions to plasma current in axisymmetric and quasi-symmetric configurations, providing a more accurate representation of the plasma behavior in these configurations. A workflow for evaluating the neoclassical transport using SFINCS with arbitrary toroidal equilibria calculated using M3D-C1 is also presented. This workflow enables a quantitative evaluation of the error in the Sauter-like model in cases that deviate from axi- or quasi-symmetry (e.g. through the development of an MHD instability).

fusion plasma↗

Electronic structure theory with molecular point group symmetries on quantum annealers

Quantum computation has the potential to revolutionize quantum chemistry through major speedups in computation times and an exponential reduction in computational resources. Here, we combine the symmetry-adapted Jordan–Wigner encoding based on the full Boolean symmetry group $\mathbb{Z}$$^{k}_{2}$ with our new implementation of the Xia–Bian–Kais (XBK) method for improving the efficiency of electronic structure theory calculations on quantum annealers, particularly by reducing the number of qubits needed to achieve the same accuracy. By providing a more extensive symmetry-adapted encoding (SAE) than previous work, we are able to simulate molecules larger than those previously reported that have been studied using methods developed for quantum annealers and without using an active space. We calculated the potential energy surfaces of H 2 , LiH, He 2 , H 2 O, O 2 , N 2 , Li 2 , F 2 , CO, BH 3 , NH 3 , and CH 4 , with the largest molecule in the STO-6G basis set requiring 16 qubits with our SAE, and compared them with full configuration interaction results. The application of SAE to the XBK method provides an exponential reduction in the size of the Hilbert space and scales well with the size of the problem. It does not introduce significant additional errors for even or large values of a key variational parameter that determines the number of ancilla qubits used in the XBK method’s Hamiltonian embedding, or for certain molecules such as He 2 and H 2 O. Here, we provide an explanation for this behavior and a recommendation on the usage of our method. In addition, we briefly discuss the potential of extracting electronic excited states from our method.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Resilience–runtime tradeoff relations for quantum algorithms

Abstract A leading approach to algorithm design aims to minimize the number of operations in an algorithm’s compilation. One intuitively expects that reducing the number of operations may decrease the chance of errors. This paradigm is particularly prevalent in quantum computing, where gates are hard to implement and noise rapidly decreases a quantum computer’s potential to outperform classical computers. Here, we find that minimizing the number of operations in a quantum algorithm can be counterproductive, leading to a noise sensitivity that induces errors when running the algorithm in non-ideal conditions. To show this, we develop a framework to characterize the resilience of an algorithm to perturbative noises (including coherent errors, dephasing, and depolarizing noise). Some compilations of an algorithm can be resilient against certain noise sources while being unstable against other noises. We condense these results into a tradeoff relation between an algorithm’s number of operations and its noise resilience. We also show how this framework can be leveraged to identify compilations of an algorithm that are better suited to withstand certain noises.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Extracting Topological Orders of Generalized Pauli Stabilizer Codes in Two Dimensions

In this paper, we introduce an algorithm for extracting topological data from translation invariant generalized Pauli stabilizer codes in two-dimensional systems, focusing on the analysis of anyon excitations and string operators. The algorithm applies to Z d qudits, including instances where d is a nonprime number. This capability allows the identification of topological orders that differ from the Z d toric codes. It extends our understanding beyond the established theorem that Pauli stabilizer codes for Z p qudits (with p being a prime) are equivalent to finite copies of Z p toric codes and trivial stabilizers. The algorithm is designed to determine all anyons and their string operators, enabling the computation of their fusion rules, topological spins, and braiding statistics. The method converts the identification of topological orders into computational tasks, including Gaussian elimination, the Hermite normal form, and the Smith normal form of truncated Laurent polynomials. Furthermore, the algorithm provides a systematic approach for studying quantum error-correcting codes. We apply it to various codes, such as self-dual CSS quantum codes modified from the two-dimensional honeycomb color code and non-CSS quantum codes that contain the double semion topological order or the six-semion topological order. Published by the American Physical Society 2024

Physics↗

4th Big Data for Nuclear Power Plants Workshop 2023

The Ohio State University and Idaho National Laboratory organized the 4 th Big Data for Nuclear Power Plants Workshop in November, 2023 in Columbus, Ohio. Workshop topics were chosen to understand the challenges and gaps that need to be addressed to maximize the impact of data on the nuclear industry, as well as the associated applications and risks. Discussions were focused around six specific application areas: Operation and Maintenance; Machine Learning in Nuclear Materials and Advanced Manufacturing; Cybersecurity; High-Performance Computing and Massive Computation; Big Data and Digital Twins; and Nuclear Non-Proliferation. The opportunities, challenges, and risks identified in the six focus areas explored in this workshop are diverse, but some common themes emerge, such as the importance of data integrity, quality, coverage, privacy, and traceability. Big data and AI/ML tools can be leveraged to reduce costs, optimize human tasking, and reduce human error across various application areas. In order for the nuclear industry to benefit from big data and advanced analytic capabilities, it is essential to address challenges and risks, such as data privacy, model reliability, and computational resource availability. Learning from other industries that have successfully implemented big data and AI/ML technologies, like the aerospace industry, can help the nuclear industry successfully integrate these technologies.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

EVALUATION OF HRA METHODOLOGIES FOR APPLICATION IN SDP WORK

This study critically evaluates human reliability analysis (HRA) methodologies applicable to regulatory probabilistic safety assessment (PSA) model, with a particular focus on their role in supporting the significance determination process (SDP) in nuclear safety assessment. Firstly, three widely utilized HRA methods – IDHEAS-ECA, SPAR-H, and ASEP/THERP – were qualitatively and quantitatively assessed. Qualitative assessments were conducted using attributes from the NEA/CSNI/R(2015)1 report, while quantitative evaluations employed regression and correlation analyses to compare predicted human error probabilities (HEPs) against empirical data. Results reveal distinct strengths, for example, IDHEAS-ECA’s robust predictive accuracy and K-HRA’s alignment with operational practices. In addition, dependency analysis and recovery analysis were critically evaluated. For dependency analysis, the methods’ handling of inter-task dependencies and their impact on HEPs were examined, while recovery analysis highlighted strategies for mitigating failure events. Furthermore, strategies were proposed to evaluate performance-shaping factors under conditions of reduced human performance, such as stress, fatigue, or cognitive overload, addressing specific challenges faced in SDP evaluations. Human errors from KINS’s operational performance information system event reports were evaluated as a case study. This study identifies gaps and provides actionable insights to ensure their validity and applicability in SDP HRA applications. This paper is a part of research conducted by KINS, and it should be noted that this result does not represent the regulatory position of KINS.

99 - GENERAL AND MISCELLANEOUS↗

Deployment of neural-network-based neutron microscopic cross sections in the Griffin reactor physics application

The capability to utilize neural networks to predict macroscopic and microscopic cross section parametric spaces has been developed for the Griffin reactor physics application. The LibTorch interface enables Griffin's MOOSE-based materials to interact with LibTorch-trained models, allowing for the evaluation of complex macroscopic or microscopic cross section spaces, which are then used to evaluate the neutronic properties of the Griffin finite element model. This study benchmarks traditional ISOXML-formatted tabulation libraries against neural network-based models for 279 nuclides on 20,160 grid points for zero-dimensional and two-dimensional reactor models. Benchmark metrics include the fundamental mode eigenvalue, fission and absorption rates, and various temperature coefficients of reactivity (isothermal, fuel, and moderator). From the perspective of storage space, the complete set of LibTorch models uses 11 MB on disk, compared to the 10 GB for the ISOXML multigroup library that covers the same grid space. For the two-dimensional performance case considered in Griffin, the Torch model uses 97% less RAM than the reference ISOXML dataset while runtime increases by a factor of 3 when using the LibTorch model compared to the ISOXML dataset with multi-linear interpolation. The LibTorch model consistently yields errors within 0.01% for most analyzed quantities except for the temperature coefficients of reactivity where the maximum discrepancies are up to 0.3 $\frac{pcm}{K}$. Due to the neural network attempting to best predict quantities with no regard for a positive or negative bias for any given quantity, predictions may experience random fluctuations, resulting in both positive and negative errors. Future work will entail both depletion and coupled transient analysis to determine the predictive capabilities of Griffin with neural network-based cross sections.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Directional finite difference method for directly solving 3D gyrokinetic field equations with enhanced accuracy

The gyrokinetic (GK) field equation is a three-dimensional (3D) elliptic equation, but it is often simplified to a set of two-dimensional (2D) equations by assuming that the field does not vary along a specific direction. However, this simplification can introduce inevitable 0th-order numerical errors, as nonlinear mode coupling in toroidal geometry can produce undesirable harmonic modes that violate the assumption. In this work, we propose a novel directional finite difference method (FDM) with a local coordinate transformation to better resolve the target field of interest. The directional FDM can accurately solve 3D GK field equations without simplifications, which can overcome the limitations of conventional methods. The accuracy and efficiency of different FDMs are analyzed in great detail for a variety of geometries, from simple 2D Cartesian coordinates to realistic 3D curvilinear coordinates. The 0th-order numerical errors of simplified 2D GK equations were found to be more problematic for low-harmonic modes and low aspect ratio geometries such as spherical tokamaks. On the other hand, the directional 3D FDM can accurately resolve a much wider range of harmonic modes aligned to the direction of interest, including the low-harmonic modes. In conclusion, we demonstrate that the directional 3D FDM is a highly effective algorithm for solving the 3D GK field equations, achieving accuracy improvements of 10 to 100 times or more, particularly for low-harmonic modes in spherical tokamaks.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Maximizing efficiency of dataset compression for machine learning potentials with information theory

Machine learning interatomic potentials (MLIPs) balance high accuracy and lower costs compared to density functional theory calculations, but their performance often depends on the size and diversity of training datasets. Large datasets improve model accuracy and generalization but are computationally expensive to produce and train on, while smaller datasets risk discarding rare but important atomic environments and compromising MLIP accuracy/reliability. Here, we develop an information-theoretical framework to quantify the efficiency of dataset compression methods and propose an algorithm that maximizes this efficiency. By framing atomistic dataset compression as an instance of the minimum set cover (MSC) problem over atom-centered environments, our method identifies the smallest subset of structures that contains as much information as possible from the original dataset while pruning redundant information. The approach is extensively demonstrated on the GAP-20 and TM23 datasets and validated on 64 varied datasets from the ColabFit repository. Across all cases, MSC consistently retains outliers, preserves dataset diversity, and reproduces the long-tail distributions of forces even at high compression rates, outperforming other subsampling methods. Furthermore, MLIPs trained on MSC-compressed datasets exhibit reduced error for out-of-distribution data even in low-data regimes. We explain these results using an outlier analysis and show that such quantitative conclusions could not be achieved with conventional dimensionality reduction methods. The algorithm is implemented in the open-source QUESTS package and can be used for several tasks in atomistic modeling, from data subsampling, outlier detection, and training improved MLIPs at a lower cost.

36 MATERIALS SCIENCE↗

Digital risk analysis in nuclear engineering projects: Designing for safety, performance, reliability, and security

Cyber-informed engineering and security-by-design frameworks are important in promoting the need to identify cybersecurity concerns early in the systems engineering lifecycle so risks from adversarial cyber-attacks can be eliminated or reduced through engineering design practices. In addition to adversarial risk, risk in operational technology systems also includes non-adversarial and unintentional risk from other factors such as human performance errors, environmental conditions, design flaws, and device degradation or failure. This paper introduces a new concept for characterizing digital risk, both adversarial and non-adversarial, and provides the basis for initial research into a novel digital risk analysis approach focused on incorporating attack difficulty into a multi-attribute analysis technique using robust decision-making. This digital risk characterization is also used to frame a discussion on the challenges of competing objectives and competing stakeholder requirements in an integrated energy system project that incorporates a small modular reactor and industrial facility.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Explainable machine learning reveals that local structural motifs encode the thermodynamic state across the CuZr metallic glass-forming range

Metallic glasses derive their properties from the statistics of local atomic motifs rather than from long-range order, yet a quantitative, chemistry-specific link between motif populations and the underlying glassy state has remained elusive. In this work we combine large-scale molecular dynamics, Voronoi tessellation, deep neural networks, and SHapley Additive exPlanations (SHAP) to identify which local structural motifs define the glassy state of Cu—Zr metallic glasses. A dataset of 17,180 atomistic configurations spanning ten compositions (Cu 20 Zr 80 –Cu 80 Zr 20 ) and four quench rates (10 9 –10 12 K/s) is used to train a feed-forward neural network that regresses temperature across the 50–2000 K liquid–supercooled–glass range, achieving a mean absolute error of 19.89 K and R 2 = 0.9974, confirming that the local structural state is faithfully encoded in motif-level structure. SHAP analysis then reveals that a tightly coupled near-icosahedral family of motifs (coordination numbers (CN) 11–13, including the full icosahedron 001200 and its single-atom-perturbation sibling 10930) collectively encodes the thermodynamic state of the system across the full glass-forming range. The CN = 11–13 ordered members carry negative SHAP values at high populations, tracking the most deeply-quenched configurations, while 10930 shows the reversed signature consistent with its role as a soft-spot host whose population shrinks as the icosahedral network deepens. The analysis demonstrates that explainable machine learning can isolate the minimal motif vocabulary defining the glassy state and recovers the near-icosahedral building blocks previously identified by data-driven analyses of Cu—Zr. The approach provides a general, chemistry-specific route for characterizing the structural state of disordered materials.

36 MATERIALS SCIENCE↗

Understanding GPU Memory Corruption at Extreme Scale: The Summit Case Study

GPU memory corruption and in particular double-bit errors (DBEs) remain one of the least understood aspects of HPC system reliability. Albeit rare, their occurrences always lead to job termination and can potentially cost thousands of node-hours, either from wasted computations or as the overhead from regular checkpointing needed to minimize the losses. As supercomputers and their components simultaneously grow in scale, density, failure rates, and environmental footprint, the efficiency of HPC operations becomes both an imperative and a challenge. We examine DBEs using system telemetry data and logs collected from the Summit supercomputer, equipped with 27,648 Tesla V100 GPUs with 2nd-generation high-bandwidth memory (HBM2). Using exploratory data analysis and statistical learning, we extract several insights about memory reliability in such GPUs. We find that GPUs with prior DBE occurrences are prone to experience them again due to otherwise harmless factors, correlate this phenomenon with GPU placement, and suggest manufacturing variability as a factor. On the general population of GPUs, we link DBEs to short- and long-term high power consumption modes while finding no significant correlation with higher temperatures. We also show that the workload type can be a factor in memory’s propensity to corruption.

Oles, Vlad↗