Search NASASearch

SEARCH · Search NASA

Results for “Statistical error”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

A Markov chain Monte Carlo (MCMC) Bayesian inference approach to analyze apparent activation barriers and reaction orders from microreactor data

Statistical analysis of steady-state catalytic kinetic data is often limited by data sparsity due to the slow pace at which the data is collected. Data sparsity and limitations in statistical analysis make it difficult to differentiate between mechanistic models and catalytic sites. A Bayesian inference tool is reported for catalysis researchers to estimate error in the determination of reaction orders from steady state microreactor data. The benefits of a Bayesian inference approach are discussed, as an alternative to the more common frequentist approach. The approach incorporates prior knowledge of the system and the data collected to form an error estimate on reaction orders. We investigated the effects of three distinct data treatments—individual fitting of trials, pooled analysis, and constrained regression methods—on the precision and uncertainty of reaction order determinations. To assess the robustness of our findings, we conducted sensitivity analyses to evaluate the influence of Bayesian parameters on uncertainty estimation. Additionally, we utilized synthetic data to illustrate how data quality impacts the precision of uncertainty assessments. We show Bayesian analysis can obtain a more precise estimation of error with a sparse data set than a frequentist analysis. Finally, this work provides strong evidence that the adoption of Bayesian analysis of kinetic data may help researchers make more precise arguments as to the strength of their evidence for a particular mechanistic hypothesis, or in comparing across different catalysts.

42 ENGINEERING

Thinking Bayesian for plasma physicists

Bayesian statistics offers a powerful technique for plasma physicists to infer knowledge from the heterogeneous data types encountered. To explain this power, a simple example, Gaussian Process Regression, and the application of Bayesian statistics to inverse problems are explained. The likelihood is the key distribution because it contains the data model, or theoretic predictions, of the desired quantities. By using prior knowledge, the distribution of the inferred quantities of interest based on the data given can be inferred. Because it is a distribution of inferred quantities given the data and not a single prediction, uncertainty quantification is a natural consequence of Bayesian statistics. The benefits of machine learning in developing surrogate models for solving inverse problems are discussed, as well as progress in quantitatively understanding the errors that such a model introduces.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Unbinned extraction of $γ$ from $B\to DK$ with normalizing flows

We introduce an unbinned method for extracting the CKM angle $γ$ from the decay chain $B^\pm \to (D \to K_S π^+ π^-) K^\pm$ using normalizing flows (NFs). The NFs, trained on $D$ decay data, learn a faithful continuous representation of the amplitude and strong phase variation over the $D\to K_Sπ^+π^-$ Dalitz plot whose fidelity improves with increased data sample sizes. With this input, the $B$ decay data can be used to extract the parameters $r_B$, $δ_B$, and $γ$. We test the method on Monte Carlo generated data, where it successfully recovers the injected value of $γ$ within uncertainties. The present implementation propagates statistical uncertainties from finite training data via an ensemble of independently trained flows, and does not attempt to capture the effects of systematic experimental errors. We explore two versions of the method that differ in how the trigonometric constraint on phase variation is encoded, and comment on the possible extension to Bayesian NFs, which would provide direct uncertainty estimates on the learned densities without requiring ensemble training.

Grossman, Yuval [Cornell U., LEPP]

Synthesis of ARM User Facility Surface Rainfall Datasets to Construct a Best Estimate Value Added Product (PrecipBE)

Surface precipitation measurements are essential for Earth system model (ESM) evaluation and understanding cloud processes. An ever-growing need for robust, temporally evolving, and easy-to-use statistical datasets provides motivation for a baseline ground-based precipitation properties data product. The U.S. Department of Energy Atmospheric Radiation Measurement (ARM) user facility operates an extensive suite of precipitation instruments with various sensitivities and operating mechanisms, which render the decision of which instrument to use based on one or more fixed thresholds challenging and prone to errors and bias. Using a long-term instrument inter-comparison from a unique per-precipitation event perspective, rather than instantaneous sample comparison, we demonstrate that ARM rainfall-measuring instruments are generally consistent with each other at the statistical level. Inter-instrument deviations at the single event level can be large, especially for specific rainfall event properties such as maximum precipitation rates. A machine-learning (ML) analysis using a random forest regressor indicates that in some cases, depending on instrument, local site climatology, and/or specific deployment configuration, certain atmospheric state variables influence the measured quantities in an unpredictable manner. Thus, a-priori weighting of different instruments does not necessarily lead to more accurate and less biased synthesis of instrument data. These results motivate the design of the ARM precipitation best-estimate (PrecipBE) value-added product, which incorporates all valid precipitation data while considering data quality and other instrument limitations. PrecipBE consists of time series and tabular statistics datasets in an easy-to-use and insightful per-precipitation event format. It provides a large set of precipitation event properties supplemented with ancillary data from ARM datasets that correspond to the detected precipitation events. We describe the PrecipBE algorithm and demonstrate its use via the examination of a single-day output as well as a long-term trend analysis of precipitation events at the ARM Southern Great Plains (SGP) site, covering more than 30 years of data. The trend analysis tentatively suggests a long-term temporal tendency for mainly shorter and less intense precipitation events at the SGP site, but a long-term increase in annual rainfall by more than 36 mm (5 %) per decade. This rainfall trend is catalyzed primarily by more extreme event properties of relatively rare, intense precipitation events, with event total and 1 min maximum precipitation rate at a 1 year timeframe increasing up to 5 mm and 9 mm h −1 (several percent) per decade, respectively. While the currently available PrecipBE datasets (at https://adc.arm.gov/discovery/, last access: 8 December 2025) cover rainfall from multiple ARM deployments up to March 2025, PrecipBE is planned to be expanded to include solid-phase precipitation and will soon become an operational product with a several-day lag from real-time. We invite the ARM user community to leverage this new product and welcome user feedback to further enhance the dataset.

Silber, Israel [Pacific Northwest National Laborat

The Art of Automation: Translating Electron Microscopy Workflows Into Automated Processes

Acquiring data using a scanning transmission electron microscope (STEM) is a complex, multi-step process. The intricacy of the process depends on the type of sample, composition of the material, desired results of the experiment, resolution requirement and other experimental factors. Each experiment presents unique complications, such as sample drift and contamination, that the microscopist must consider when acquiring data. All these challenges are handled fluidly and expertly by experienced microscopists, but to reach new levels of innovation in material development, including greater reproducibility, throughput, and precision, the automation of these workflows is essential. The initial phase of this work involved translating intuition-based workflows into discrete, programmable steps. Some common key stages in STEM workflows are the initial tuning, scanning the sample for areas of interest, and then acquiring the data. Each stage can be broken further into specific parameter adjustments, such as aberration correction and dwell time optimization, depending on the experiment. When deconstructing various experiments each step was assessed for automation feasibility based on the amount of real time operator decisions. There are steps that lend themselves to automation more readily than others, such as course focusing and sample screening, but there is potential for full automation of all stages with time. As an initial step, an automated montage routine was developed, allowing for the efficient acquisition of large portions of the sample without requiring continuous intervention from the operator. The automation of this small process of the procedure demonstrates the value of this capability. A major challenge in automation arises from discrepancies between commanded, reported and actual stage movements. Using systematic tests, stage movement was quantified. This error can be corrected algorithmically for more accurate workflows in the future. Expanding automation capabilities would result in larger, more efficient data acquisition which allows for more robust statistical analysis. Additionally, this work lays the groundwork for a closed loop system where machine learning algorithms would intake automatically acquired data and make real time decisions. By progressively automating this instrument, this work establishes the foundation for fully automated experimentation in transmission electron microscopy.

97 MATHEMATICS AND COMPUTING

Measurement of Atmospheric Neutrino Oscillation Parameters Using Convolutional Neural Networks with 9.3 Years of Data in IceCube DeepCore

The DeepCore subdetector of the IceCube Neutrino Observatory provides access to neutrinos with energies above approximately 5 GeV. Data taken between 2012 and 2021 (3387 days) are utilized for an atmospheric ν μ disappearance analysis that studied 150 257 neutrino-candidate events with reconstructed energies between 5 and 100 GeV. An advanced reconstruction based on a convolutional neural network is applied, providing increased signal efficiency and background suppression, resulting in a measurement with both significantly increased statistics compared to previous DeepCore oscillation results and high neutrino purity. For the normal neutrino mass ordering, the atmospheric neutrino oscillation parameters and their 1 σ errors are measured to be Δ m 32 2 = 2.40 − 0.04 + 0.05 × 10 − 3 eV 2 and sin 2 θ 23 = 0.54 − 0.03 + 0.04 . The results are the most precise to date using atmospheric neutrinos, and are compatible with measurements from other neutrino detectors including long-baseline accelerator experiments. Published by the American Physical Society 2025

Abbasi, R.

CV4Quantum: Reducing the Sampling Overhead in Probabilistic Error Cancellation Using Control Variates

Quasiprobabilistic decompositions (QPDs) play a key role in maximizing the utility of near-term quantum hardware. For example, Probabilistic Error Cancellation (PEC) (an error mitigation technique) and circuit cutting (which enables large quantum computations to be performed on quantum hardware with a limited number of qubits) both involve QPDs. Computations based on QPDs typically incur large sampling overheads that grow exponentially with the number of error-terms mitigated or number of circuit-cuts employed, limiting their practical feasibility. In this work, we adapt the control variates variance reduction technique from the statistics literature in order to reduce the sampling overhead in QPD-based computations. We demonstrate our method using simulation experiments that mimic a realistic PEC scenario. In our experiments, we observed a more than 50% reduction in the number of samples needed to achieve a given precision, in more than 50% of the PEC-based estimations performed in the study when using our approach. We discuss how future research on constructing good control variates can lead to even stronger sampling overhead reduction.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Surrogate model evaluation and building energy benchmarking for commercial buildings

Building energy consumption benchmarking involves challenges associated with various energy patterns for different building types; heating, ventilating, and air-conditioning (HVAC) system types; and climates. Given significant variation in energy use patterns, accurate prediction of long-term energy use using surrogate models remains challenging. Multiple linear regression (MLR) is commonly used for building energy benchmarking because of its simple structure; however, it lacks accuracy compared to other black-box models. Although many studies have compared surrogate models and offer guidance on model selection based on metrics, they do not provide detailed analysis on improving the surrogate model accuracy. In this paper, we implement a surrogate model using polynomial ridge regression (i.e., MLR with interaction terms combined with ridge regularization) for small office and retail strip mall buildings across six HVAC system types and all climate zones, for electricity and natural gas in baseline and proposed scenarios. A simulation workflow is developed using OpenStudio TM /EnergyPlus TM to generate simulation data using measures over a wide range of efficiency inputs. Enhancements based on statistical insights are used for improving the model accuracy using filters, input transformations, and change points. Surrogate models achieved average coefficient of variation of the root mean squared error (CVRMSE) values of 2.17, 1.06, 2.05, and 3.26 for proposed electricity, proposed natural gas, baseline electricity, and baseline natural gas, respectively, with enhancements reducing CVRMSE by an average of 14.9% across all combinations. We provide model interpretation via Shapley additive explanations to determine which input variables most influence energy consumption and provide supportive arguments for enhancements.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

2025 Review and Revision of Federal Guidance Report 15

Federal Guidance Report No. 15 (FGR 15), External Exposure to Radionuclides in Air, Water and Soil, published in 2019 and referred to below as FGR 2019, provides age-specific effective dose rate coefficients for reference persons externally exposed to each of 1252 radionuclides homogenously distributed in environmental media. Soon after completion of FGR 2019, the International Commission on Radiological Protection (ICRP) published a similar report (ICRP Publication 144, 2020) addressing the same radionuclides and many of the external exposure scenarios addressed in FGR 2019. In 2023 Argonne National Laboratory (ANL) published a review of dose coefficients in FGR 2019, concluding that “The external effective dose from beta radiation is not appropriately accounted for in FGR 15 for both low- and high-energy beta emitters in air, in soil, and on soil surfaces.” That conclusion was based largely on comparisons of dose coefficients in FGR 2019 with values in ICRP Publication 144 but also on consideration of some unexpected patterns of equivalent dose rate coefficients across tissues and of effective dose rate coefficients across different soil depths, for a selected set of low-energy beta emitters. In response to the ANL report, the Center for Radiation Protection Knowledge (CRPK) at Oak Ridge National Laboratory (ORNL) performed an extensive reexamination of the methods and published values of FGR 2019. CRPK found that coding errors had resulted in inaccurate estimates, primarily overestimates, in many of the dose coefficients tabulated in FGR 2019 but that many of the differences between results in FGR 2019 and ICRP Publication 144 could be traced to differences in methodology. CRPK has corrected all errors in FGR 2019 associated with the coding errors and has taken the opportunity to improve a major portion of the remaining dose coefficients in FGR 2019, primarily through additional Monte Carlo calculations resulting in improved statistics for tissue equivalent dose rate coefficients for exposure to monoenergetic sources. This has eliminated the need for extrapolation of results from high and medium energies to low energies, as was done in FGR 2019. In addition, inconsistencies in the methodology identified in the review, such as computational representation of the adult female, were eliminated. The revised effective dose rate coefficients are consistent with values in ICRP Publication 144 except for differences in values clearly arising from differences in methodology; and differences in values for a relatively small set of very low-energy radionuclides with highly uncertain dose coefficients regardless of methodology.

54 ENVIRONMENTAL SCIENCES

Power and Limitations of Linear Programming Decoder for Quantum LDPC Codes

Decoding quantum error-correcting codes is a key challenge in enabling fault-tolerant quantum computation. In the classical setting, linear programming (LP) decoders offer provable performance guarantees and can leverage fast practical optimization algorithms. Although LP decoders have been proposed for quantum codes, their performance and limitations remain relatively underexplored. In this work, we uncover a key limitation of LP decoding for quantum low-density parity-check (LDPC) codes: certain constant-weight error patterns lead to ambiguous fractional solutions that cannot be resolved through independent rounding. To address this issue, we incorporate a post-processing technique known as ordered statistics decoding (OSD), which significantly enhances LP decoding performance in practice. Our results show that LP decoding, when augmented with OSD, can outperform belief propagation with the same post-processing for intermediate code sizes of up to hundreds of qubits. These findings suggest that LP-based decoders, equipped with effective post-processing, offer a promising approach for decoding near-term quantum LDPC codes.

Gu, Shouzhen [Yale U.]

Power and Limitations of Linear Programming Decoder for Quantum LDPC Codes

Decoding quantum error-correcting codes is a key challenge in enabling fault-tolerant quantum computation. In the classical setting, linear programming (LP) decoders offer provable performance guarantees and can leverage fast practical optimization algorithms. Although LP decoders have been proposed for quantum codes, their performance and limitations remain relatively underexplored. In this work, we uncover a key limitation of LP decoding for quantum low-density parity-check (LDPC) codes: certain constant-weight error patterns lead to ambiguous fractional solutions that cannot be resolved through independent rounding. To address this issue, we incorporate a post-processing technique known as ordered statistics decoding (OSD), which significantly enhances LP decoding performance in practice. Our results show that LP decoding, when augmented with OSD, can outperform belief propagation with the same post-processing for intermediate code sizes of up to hundreds of qubits. These findings suggest that LP-based decoders, equipped with effective post-processing, offer a promising approach for decoding near-term quantum LDPC codes.

Gu, Shouzhen [Yale U.]

Power and Limitations of Linear Programming Decoder for Quantum LDPC Codes

Decoding quantum error-correcting codes is a key challenge in enabling fault-tolerant quantum computation. In the classical setting, linear programming (LP) decoders offer provable performance guarantees and can leverage fast practical optimization algorithms. Although LP decoders have been proposed for quantum codes, their performance and limitations remain relatively underexplored. In this work, we uncover a key limitation of LP decoding for quantum low-density parity-check (LDPC) codes: certain constant-weight error patterns lead to ambiguous fractional solutions that cannot be resolved through independent rounding. To address this issue, we incorporate a post-processing technique known as ordered statistics decoding (OSD), which significantly enhances LP decoding performance in practice. Our results show that LP decoding, when augmented with OSD, can outperform belief propagation with the same post-processing for intermediate code sizes of up to hundreds of qubits. These findings suggest that LP-based decoders, equipped with effective post-processing, offer a promising approach for decoding near-term quantum LDPC codes.

Gu, Shouzhen [Yale U.]

DESI DR2 results. I. Baryon acoustic oscillations from the Lyman alpha forest

We present the baryon acoustic oscillation (BAO) measurements with the Lyman-𝛼 (Ly⁢𝛼) forest from the second data release (DR2) of the Dark Energy Spectroscopic Instrument (DESI) survey. Our BAO measurements include both the autocorrelation of the Ly⁢𝛼 forest absorption observed in the spectra of high-redshift quasars and the cross-correlation of the absorption with the quasar positions. The total sample size is approximately a factor of 2 larger than the DR1 dataset, with forest measurements in over 820,000 quasar spectra and the positions of over 1.2 million quasars. We describe several significant improvements to our analysis in this paper, and two supporting papers describe improvements to the synthetic datasets that we use for validation and how we identify damped Ly⁢𝛼 absorbers. Our main result is that we have measured the BAO scale with a statistical precision of 1.1% along and 1.3% transverse to the line of sight, for a combined precision of 0.65% on the isotropic BAO scale at 𝑧 eff =2.33. This excellent precision, combined with recent theoretical studies of the BAO shift due to nonlinear growth, motivated us to include a systematic error term in Ly⁢𝛼 BAO analysis for the first time. We measure the ratios 𝐷 𝐻 ⁡(𝑧 eff )/𝑟 𝑑 = 8.632 ± 0.098 ± 0.026 and 𝐷 𝑀 ⁡(𝑧 eff )/𝑟 𝑑 = 38.99 ± 0.52 ± 0.12, where 𝐷 𝐻 = 𝑐/𝐻⁡(𝑧) is the Hubble distance, 𝐷 𝑀 is the transverse comoving distance, 𝑟 𝑑 is the sound horizon at the drag epoch, and we quote both the statistical and the theoretical systematic uncertainty. The companion paper presents the BAO measurements at lower redshifts from the same dataset and the cosmological interpretation.

baryon acoustic oscillations

DESI DR2 Results I: Baryon Acoustic Oscillations from the Lyman Alpha Forest

We present the Baryon Acoustic Oscillation (BAO) measurements with the Lyman-alpha (LyA) forest from the second data release (DR2) of the Dark Energy Spectroscopic Instrument (DESI) survey. Our BAO measurements include both the auto-correlation of the LyA forest absorption observed in the spectra of high-redshift quasars and the cross-correlation of the absorption with the quasar positions. The total sample size is approximately a factor of two larger than the DR1 dataset, with forest measurements in over 820,000 quasar spectra and the positions of over 1.2 million quasars. We describe several significant improvements to our analysis in this paper, and two supporting papers describe improvements to the synthetic datasets that we use for validation and how we identify damped LyA absorbers. Our main result is that we have measured the BAO scale with a statistical precision of 1.1% along and 1.3% transverse to the line of sight, for a combined precision of 0.65% on the isotropic BAO scale at $z_{eff} = 2.33$. This excellent precision, combined with recent theoretical studies of the BAO shift due to nonlinear growth, motivated us to include a systematic error term in LyA BAO analysis for the first time. We measure the ratios $D_H(z_{eff})/r_d = 8.632 \pm 0.098 \pm 0.026$ and $D_M(z_{eff})/r_d = 38.99 \pm 0.52 \pm 0.12$, where $D_H = c/H(z)$ is the Hubble distance, $D_M$ is the transverse comoving distance, $r_d$ is the sound horizon at the drag epoch, and we quote both the statistical and the theoretical systematic uncertainty. The companion paper presents the BAO measurements at lower redshifts from the same dataset and the cosmological interpretation.

79 ASTRONOMY AND ASTROPHYSICS

Rapid neutron and gamma-ray source localization using machine learning

Rapid localization of radiation sources is critical for applications including nuclear emergency response, safeguards, and security. However, conventional imaging systems such as neutron scatter cameras and Compton cameras depend on rare coincidence events, which often result in long acquisition times. In this work, we address the challenge of rapid source localization by developing a machine learning approach to predict the direction of a single radiation source using only count rates from an array of neutron and gamma-ray detectors. The proposed model is a fully connected neural network (FCNN) trained using Monte Carlo simulation data from a 252 Cf source. The model hyperparameters are optimized with a small set of routine 252 Cf measurements. We benchmarked the performance of the trained and optimized machine learning model using additional 252 Cf , 137 Cs , and PuBe measurements under laboratory conditions with varying source-detector configurations. For these measurements, the machine learning model achieved a mean localization error smaller than 30° with 3 x 10 3 system counts, corresponding to 8 s measurement time for the imaging system used in this work. In this low-statistics regime, the method outperformed traditional scatter-based imaging by more than 75% in localization accuracy for the evaluated measurement configurations. These results demonstrate that a machine learning-based approach can significantly reduce the time required for accurate single-source localization, providing a robust and computationally efficient alternative to traditional imaging systems in time-critical nuclear security and emergency response scenarios.

Gamma-ray imaging

High-Fidelity Velocity and Concentration Measurements of Turbulent Buoyant Jets

Accurate models of turbulent buoyant flows are essential for the design of the cooling circuit of nuclear reactors and passive safety systems. However, available models fail to fully capture the physics of turbulent mixing when buoyancy becomes predominant with respect to momentum. Therefore, high-fidelity experiments of well-controlled fundamental flows are needed to develop and validate more accurate models. We analyze experiments of positive and negative turbulent buoyant jets, both in uniform and stratified environments, with the aim of understanding the thermal hydraulics of turbulent mixing with variable density and providing high-fidelity data for the development and validation of turbulence models. Non-intrusive, simultaneous particle image velocimetry and laser-induced fluorescence measurements were carried out to acquire instantaneous velocity and concentration fields on a vertical section parallel to the axis of a jet in the self-similar region. The refractive index matching method was applied to measure high-resolution buoyant jets with up to 8.6% density difference. These data are free of the typical errors that characterize optical measurements of buoyancy-driven flows (e.g. natural and mixed convection) where the refractive index of the fluid is inhomogeneous throughout the measurement domain. Turbulent statistics and entrainment of buoyant jets in uniform and stratified environments are presented. These data are compared with non-buoyant jets in a uniform environment, as a reference to investigate the effects of buoyancy and stratification on turbulent mixing. The results will be used for the assessment of current turbulence models and as a basis for the development of a new model that captures turbulent mixing.

laser induced fluorescence

Composite-dimensional topological codes with boundaries and defects

We introduce new algorithms and provide example constructions of stabilizer models for the gapped boundaries, domain walls, and 0D defects of Abelian composite-dimensional twisted quantum doubles. Using the physically intuitive concept of condensation, our algorithm explicitly describes how to construct the boundary and domain-wall stabilizers starting from the bulk model. This extends the utility of Pauli stabilizer models in describing nontranslationally invariant topological orders with gapped boundaries. To highlight this utility, we provide a series of examples, including a new family of quantum error-correcting codes where the double of ℤ4 is coupled to instances of the double semion (DS) phase. We discuss the codes' utility in the burgeoning area of quantum error correction with an emphasis on the interplay between deconfined anyons, logical operators, error rates, and decoding. We also augment our construction, built using algorithmic tools to describe the properties of explicit stabilizer layouts at the microscopic lattice level, with dimensional counting arguments and macroscopic-level constructions building on pants decompositions. The latter outlines how such codes' representation and design can be automated. Our results are validated by a series of error-correcting threshold calculations comparing our codes' performance with that of standard surface codes. To do so, we introduce a composite-dimensional belief-propagation decoder with ordered statistics that utilizes combination sweeps. Going beyond our worked-out examples, we expect our explicit step-by-step algorithms to pave the path for higher-dimensional codes to be discovered and implemented in near-future architectures that take advantage of various hardware platforms.

Mousa, Mohamad [Purdue University]

Revealing EDL-driven reduction mechanisms in binary, ternary, and quaternary fluorinated electrolytes via an integrated MD–DFT–ML framework

Accurately predicting solid electrolyte interphase (SEI) formation requires explicitly resolving the electric double layer (EDL) structure, which deviates significantly from that of the bulk electrolyte. Although an established molecular dynamics (MD) and Density Functional Theory (DFT) framework can model SEI formation by evaluating reduction reactions of local clusters in the EDL, it suffers from a combinatorial computational bottleneck. To overcome this limitation, we introduce a machine-learning-accelerated simulation workflow (MD–DFT–ML), integrating a gradient-boosted regression model trained on EDL composition data to efficiently predict reduction potentials. We apply this framework to seven fluorinated electrolytes comprising fluorinated anions, a fluorinated ester solvent, two types of diluent (ion-solvating ester vs. non-solvating ether), and an FEC additive. The analysis shows that the EDL selectively accumulates cation-binding species; consequently, the non–cation-binding ether diluent rarely enters the EDL and makes minimal contributions to SEI formation. DFT calculations on statistically representative EDL clusters provide reduction potentials and fluorine-release pathways, while the ML model, which substantially reduces the DFT workload, predicts cluster reduction energies with a mean absolute error of 0.1 eV. The combined MD–DFT–ML approach also quantifies contributions from different sources to LiF formation in the SEI. This methodology establishes a generalizable route for multiscale modeling electrolyte and interphase design for next-generation electrochemical energy-storage systems.

DFT-MD-ML workflow