Search NASASearch

SEARCH · Search NASA

Results for “Statistical error”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Perspectives on Systematic Cloud Microphysics Scheme Development With Machine Learning

Cloud microphysics—the collection of processes that govern the small‐scale formation, evolution, and interactions of liquid droplets and ice crystals in clouds and precipitation—remains a major source of uncertainty in weather and climate models. Although too small in scale to be explicitly resolved in any large‐eddy simulation, weather, or climate model, the representation of cloud microphysical processes has significant impact at the climate scale. Current microphysical schemes are limited by both parametric uncertainty, linked to uncertainty in physical parameter values, and structural uncertainty, arising from incomplete physical understanding of the processes at play or approximations made for computational efficiency. Recent advances in the application of machine learning (ML) to the physical sciences show significant potential for minimizing these limitations by leveraging high‐fidelity simulations and observations. Here we outline the challenges that must be addressed to apply ML toward cloud microphysics scheme development. This perspectives paper synthesizes recent progress in using data‐driven methods, including ML, to improve cloud microphysics parameterizations and highlights opportunities to address key uncertainties. We discuss the roles of aleatoric (irreducible, or statistical) and epistemic (reducible, or systematic) errors in contributing to microphysics parameterization uncertainty. ML can leverage observations to improve microphysical schemes via bottom‐up and top‐down constraints. Methods such as differentiable programming and ML‐enhanced sampling strategies and the creation of large scale benchmark data sets promise to bridge the gap between observations and models and to improve the consistency of cloud microphysical representation across temporal and spatial scales.

Lamb, Kara D. [Columbia Univ., New York, NY (Unite

Misclassification in Workers’ Telecommuting Frequency Choices Using a Generalized Extreme Value Model

Telecommuting frequency is a response variable collected in travel surveys and is, therefore, prone to errors leading to mismeasurements or misclassification. Misclassification of explanatory variables is a common risk when using statistical modeling techniques. We define “misclassification” as a response reported or recorded in the wrong category; for example, a variable is recorded as a 1 when it should be 0. Here, in this context, this study aims to develop a statistical model to analyze telecommuting data which accounts for potential misclassification errors by building on existing literature in econometrics. The empirical analysis was undertaken using the 2017 National Household Travel Survey (NHTS) and the general extreme value (GEV) models available in the literature. Specifically, the frequency of telecommuting days was analyzed using the negative binomial (NB) model recast as the multinomial logit (MNL) model. By nature—and consistent with other studies—NHTS data are prone to errors that can be classified as intentional or unintentional misinformation provided by the person being interviewed. Ignoring these errors while modeling telecommuting frequencies using standard discrete count models can result in biased parameter estimates. The misclassification parameter was calculated for both over-reporting and under-reporting scenarios. The misclassification errors can be as high as 14% over-reported and 10% under-reported, particularly for the neighboring values. Statistical fit comparison between the models shows that models that ignore misclassification have worse data fit and biased parameter estimates with significant policy implications.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Generalization error guaranteed auto-encoder-based nonlinear model reduction for operator learning

Many physical processes in science and engineering are naturally represented by operators between infinite-dimensional function spaces. The problem of operator learning, in this context, seeks to extract these physical processes from empirical data, which is challenging due to the infinite or high dimensionality of data. An integral component in addressing this challenge is model reduction, which reduces both the data dimensionality and problem size. In this paper, we utilize low-dimensional nonlinear structures in model reduction by investigating Auto-Encoder-based Neural Network (AENet). AENet first learns the latent variables of the input data and then learns the transformation from these latent variables to corresponding output data. Our numerical experiments validate the ability of AENet to accurately learn the solution operator of nonlinear partial differential equations. Furthermore, we establish a mathematical and statistical estimation theory that analyzes the generalization error of AENet. Finally, our theoretical framework shows that the sample complexity of training AENet is intricately tied to the intrinsic dimension of the modeled process, while also demonstrating the robustness of AENet to noise.

Auto-encoder

ESM data downscaling: a comparison of super-resolution deep learning models

Abstract Climate projections at fine spatial resolutions are required to conduct accurate risk assessment for critical infrastructure and design adaptation planning. Generating these projections using advanced Earth system models (ESM) requires significant computational resources. To address this issue, various statistical downscaling techniques have been introduced to generate fine-resolution data from coarse-resolution simulations. In this study, we evaluate and compare five deep learning-based downscaling techniques, namely, super-resolution convolutional neural networks, fast super-resolution convolutional neural network ESM, efficient sub-pixel convolutional neural network, enhanced deep residual network (EDRN), and super-resolution generative adversarial network (SRGAN). These techniques are applied to a dataset generated by the Energy Exascale Earth System Model (E3SM), focusing on key surface variables such as surface temperature, shortwave heat flux, and longwave heat flux. Models are trained and validated using paired fine-resolution (0.25 $$^{\circ }$$ ∘ ) and coarse-resolution (1 $$^{\circ }$$ ∘ ) monthly data obtained from a 9-year simulation. Next, blind testing is performed using monthly data obtained from two different years outside of the training and validation set. To evaluate the efficiency of each technique, different statistical metrics are used, including mean squared error (MSE), peak signal-to-noise ratio (PSNR), structural similarity index measure (SSIM), and learned perceptual image patch similarity (LPIPS). The results show that EDRN outperforms other algorithms in terms of PSNR, SSIM, and MSE, but struggles to capture fine-scale features in the data. In contrast, SRGAN, a generative model that uses perceptual loss, excels in capturing fine details at boundaries and internal structures, resulting in lower LPIPS than other methods.

Pawar, Nikhil M. (ORCID:0000000211613289)

A novel conditional generative model for efficient ensemble forecasts of state variables in large-scale geological carbon storage

Integrating monitoring data to efficiently update reservoir pressure and CO 2 plume distribution forecasts presents a significant challenge in geological carbon storage (GCS) applications. Inverse modeling techniques are commonly used to fuse observational data and refine reservoir model parameters, thereby improving state variable forecasts. However, these techniques often rely on linear or Gaussian assumptions, which can limit their effectiveness in accurately predicting state variables. Moreover, simulating large-scale three-dimensional (3D) GCS problems is computationally expensive, making iterative runs in inverse problems prohibitive. To address these challenges, we propose a conditional generative model utilizing the score-based diffusion method for real-time 3D pressure and saturation field distribution predictions. Our approach involves solving the score function with a mini-batch-based Monte Carlo estimator to generate labeled data. This data is subsequently employed to train a fully connected neural network, enabling it to learn the conditional sample generator within a supervised learning framework. This method enables the rapid generation of a large ensemble of predictions, facilitating comprehensive uncertainty quantification of state variables. Here we applied our method to forecast the dynamic 3D distributions of pressure and saturation fields over a 30-year injection period. The statistical assessment with low root mean square error (RMSE) values demonstrates that our method can accurately predict the spatiotemporal distributions of both pressure and saturation fields. Moreover, the developed conditional generative model shows high computational efficiency by generating 100 ensemble forecasts of 3D state variables in less than 10 min. The consistency between ensemble averages and ground truth values further illustrates the model’s capability to capture state variable dynamics during the CO 2 plume injection process. Notably, the ground truth values fall within the ensemble forecasts, indicating that our uncertainty quantification effectively captures variability and potential noise in the observations. Thus, the developed conditional generative model proves to be a more efficient, accurate, and practical tool for GCS applications, facilitating timely risk analysis and informed decision-making.

58 GEOSCIENCES

Development and implementation of high-throughput proteomic and metabolomics assays by using advanced chromatographic and mass spectrometric systems (CRADA Final Report)

The mission of this CRADA with Agilent was to couple powerful MS platforms (QQQ, IM-QTOFMS) with Agilent’s novel Ultra-High-Performance Liquid Chromatography (UHPLC) fast metabolomic workflows and perform ABF Machine Learning (ML) to generated datasets. Agilent transferred UHPLC methods to PNNL and LBNL and methods were implemented and demonstrated in both labs, achieving total acquisition times of < 10 min. Metabolites analyzed using Agilent’s shared methods included metabolites from central carbon metabolism, common across hosts, and metabolites unique to engineered strains. Standards were acquired in an UHPLC-Drift Tube Ion Mobility Mass Spectrometer (DTIMS) system for the first time within the context of ABF and methods were optimized based on Agilent’s protocols. Samples from ABF hosts Pseudomonas putida, Aspergillus pseudoterreus, Aspergillus niger and Rhodosporidium toruloides were analyzed using the UHPLC-DTIMS platform for a total of 276 runs. A data analysis workflow compatible with the Experimental Data Depot (EDD) and completely shareable was developed for the acquired UHPLC-DTIMS data. Samples were analyzed using a Data Independent Acquisition Approach (DIA), which for most of the standards provided more transitions therefore increasing detection confidence. Using the data acquired by PNNL, LBNL, and Agilent’s specifications from previous ML projects, SNL applied an ensemble ML strategy to pick the best performing model for automated LC-method selection. Finally, with the contribution of the participant labs and Agilent, SNL developed an Automated Method Selection (AMS) software tool to predict the best liquid chromatography method for analysis of any new molecules of interest. Samples with novel pathways and new metabolite targets of interest are generated at a high pace in the ABF. Overall, the project advanced rapid metabolomics by combining liquid chromatography, ion mobility spectrometry, and data-independent mass spectrometry with machine learning. This multidimensional approach uses retention time, collision cross-section, precursor mass, and fragment-ion information to distinguish chemically similar metabolites that can be difficult to resolve using conventional liquid- or gas-chromatography methods. The resulting workflow also provided automated metabolite-identification error estimates, addressing a recognized need for statistical confidence measures in metabolomics.

Petzold, Christopher [Lawrence Berkeley National L

Xanthos-Lake Dataset

The Xanthos-Lake v1.0 dataset provides the input data, trained machine-learning models, and simulation outputs needed to characterize lake water balance, snow and ice conditions, and mixing-layer temperature within the Xanthos global hydrological modeling framework. The dataset supports lake representation across a wide range of lake sizes and hydroclimatic conditions by combining xLSIM, a basin-specific machine-learning emulator of lake snow, ice, ice-cover fraction, and mixing-layer temperature, with the Xanthos-Lake water-balance model. The archive contains NetCDF datasets used to train and evaluate xLSIM, trained model weights, processed meteorological and lake-property inputs, and basin- and lake-category-specific simulation outputs. These materials are organized into four primary data groups, described below. Snowice_model_inputs: Contains the NetCDF input data used to train xLSIM. The xLSIM machine-learning framework uses three lake-based datasets. The meteorological forcing dataset provides monthly relative humidity, specific humidity, surface wind speed, maximum and minimum air temperature, downward longwave and shortwave radiation, snowfall, surface air pressure, and total precipitation. Lake surface area is included as an additional static predictor. The target-state dataset provides lake ice thickness, snow depth, snow cover, and lake mixing-layer temperature, while a companion lake-surface dataset provides the lake ice-cover fraction. Before training, ice thickness and snow depth are converted from meters to centimeters, mixing-layer temperature is converted from kelvin to degrees Celsius and constrained to nonnegative values, and ice-cover fraction is converted from a fraction to a percentage. The predictor variables are normalized using statistics calculated across the selected lakes and time steps. Snowice_model_outputs: Contains the NetCDF outputs generated by xLSIM. For each basin, xLSIM produces a file containing observed and predicted lake-state variables for the training, validation, and testing periods. The modeled variables include lake ice thickness, snow depth, snow cover, mixing-layer temperature, and lake ice-cover fraction. For basins without a sufficiently persistent snow-and-ice signal, the emulator predicts only mixing-layer temperature. The outputs also include training and validation loss histories, the selected model configuration, identifiers of the lakes used in training, and SHAP-based feature-importance information at the global, lake, and seasonal-regime levels. The trained machine-learning model weights are provided separately within the dataset archive. Together, these files support model evaluation and subsequent coupling with the Xanthos-Lake water-balance framework. XanthosLAKES: Contains the NetCDF input data used by the Xanthos-Lake framework. Monthly meteorological inputs include relative and specific humidity, downward shortwave and longwave radiation, mean, maximum, and minimum air temperature, wind speed, precipitation, snowfall, and surface air pressure. Static lake-property datasets provide lake identifiers, geographic locations, surface area, volume, mean depth, elevation, drainage area, fetch, outlet-routing information, and associated Xanthos grid-cell attributes. Separate bathymetric datasets provide the coefficients of the area–depth and volume–depth relationships for each aggregated lake unit. GLEV-based records provide observed lake surface area and evaporation data used to initialize lake states, define reference conditions, and calibrate and evaluate the model. Xanthos-Lake Outputs: Contains the basin- and lake-category-specific NetCDF outputs generated by Xanthos-Lake. Monthly variables include lake surface area, storage volume, outlet discharge, evaporation rate, evaporation volume, lake–groundwater exchange, lake inflow, ice thickness, snow depth, snow-cover fraction, ice-cover fraction, and mixing-layer temperature. The files also contain lake-specific calibration and validation statistics, including normalized root-mean-square error, mean absolute error, Nash–Sutcliffe efficiency, Kling–Gupta efficiency, and percent bias. Stored calibrated and derived parameters include the weir discharge coefficient, fractional freeboard, groundwater exchange coefficient, reference water level, corresponding reference surface area and storage volume, weir-width adjustment factor, and the fraction of routed inflow entering the lake. Basin identifiers, lake category, simulation period, calibration and validation periods, and parameter-schema information are retained as NetCDF metadata.

Abeshu, Guta [Pacific Northwest National Laborator

Comparing Compressed and Full-Modeling analyses with FOLPS: implications for DESI 2024 and beyond

The Dark Energy Spectroscopic Instrument (DESI) will provide unprecedented information about the large-scale structure of our Universe. In this work, we study the robustness of the theoretical modelling of the power spectrum of F OLPS , a novel effective field theory-based package for evaluating the redshift space power spectrum in the presence of massive neutrinos. We perform this validation by fitting the AbacusSummit high-accuracy N -body simulations for Luminous Red Galaxies, Emission Line Galaxies and Quasar tracers, calibrated to describe DESI observations. We quantify the potential systematic error budget of F OLPS finding that the modelling errors are fully sub-dominant for the DESI statistical precision within the studied range of scales. Additionally, we study two complementary approaches to fit and analyse the power spectrum data, one based on direct Full-Modelling fits and the other on the ShapeFit compression variables, both resulting in very good agreement in precision and accuracy. In each of these approaches, we study a set of potential systematic errors induced by several assumptions, such as the choice of template cosmology, the effect of prior choice in the nuisance parameters of the model, or the range of scales used in the analysis. Furthermore, we show how opening up the parameter space beyond the vanilla ΛCDM model affects the DESI observables. These studies include the addition of massive neutrinos, spatial curvature, and dark energy equation of state. We also examine how relaxing the usual Cosmic Microwave Background and Big Bang Nucleosynthesis priors on the primordial spectral index and the baryonic matter abundance, respectively, impacts the inference on the rest of the parameters of interest. This paper pathways towards performing a robust and reliable analysis of the shape of the power spectrum of DESI galaxy and quasar clustering using F OLPS .

79 ASTRONOMY AND ASTROPHYSICS

FRAM Version 7.1’s Bias

The Fixed-Energy Response-Function Analysis with Multiple Efficiency (FRAM) code was developed at Los Alamos National Laboratory to measure the gamma-ray spectrometry of the isotopic composition of plutonium, uranium, and other actinides. For FRAM versions 4 and earlier, the reported uncertainties of the results come from the propagation of the statistics in the peak areas only. No systematic error components are included in the reported uncertainties. For FRAM versions 5 and 6, we examined the FRAM analytical results of both the archival plutonium data and the data specifically acquired for the isotopic uncertainty analysis project and found the relationship between the bias and other parameters. We worked out the equations representing the biases of the measured isotopes from each measurement using internal spectral parameters, such as peak resolution and shape, region of analysis, and burnup (for plutonium) or enrichment (for uranium). The resulting biases were included in the reported uncertainties of FRAM v.5 and v.6.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Perturbative Stability and Error-Correction Thresholds of Quantum Codes

Topologically ordered phases are stable to local perturbations, and topological quantum error-correcting codes enjoy thresholds to local errors. We connect the two notions of stability by constructing classical statistical mechanics models for decoding general Calderbank-Shor-Steane codes and classical linear codes. Our construction encodes correction success probabilities under uncorrelated bit-flip and phase-flip errors, and simultaneously describes a generalized ℤ 2 lattice-gauge theory with quenched disorder. We observe that the clean limit of the latter is precisely the discretized imaginary-time path integral of the corresponding quantum code Hamiltonian when the errors are turned into a perturbative 𝑋 or 𝑍 magnetic field. Motivated by error-correction considerations, we define general order parameters for all such generalized ℤ 2 lattice-gauge theories, and show that they are generally lower bounded by success probabilities of error correction. For CSS codes satisfying the low-density parity-check condition and with a sufficiently large code distance, we prove the existence of a low-temperature ordered phase of the corresponding lattice-gauge theories, particularly for those lacking Euclidean spatial locality and/or when there is a nonzero code rate. We further argue that these results provide evidence for stable phases in the corresponding perturbed quantum Hamiltonians, obtained in the limit of continuous imaginary time. To do so, we distinguish space- and timelike defects in the lattice-gauge theory. A high free-energy cost of spacelike defects corresponds to a successful “memory experiment” and suppresses the energy splitting among the ground states, while a high free-energy cost of timelike defects corresponds to a successful “stability experiment” and points to a nonzero gap to local excitations.

quantum error correction

The National Climate Data Base (NCDB): A Bias-Corrected High-Resolution Climate Dataset

Assessing renewable energy resources under future climate scenarios has been highlighted in recent years to analyze and understand potential impacts of future change in renewable generation on the power sector. Solar energy is well-known as the most plentiful among various renewable resources and usually converted to electricity using photovoltaics (PV) technologies, and the global deployment of PV technology has increased rapidly in recent decades. In this study, we develop a statistical technique to downscale the future projection of solar irradiance for PV energy-related applications. A set of Regional Climate Model (RCM)-based projections obtained from the North American Coordinated Regional Climate Downscaling Experiment (NA-CORDEX) are used as inputs to statistical methods to generate high-resolution global horizontal irradiance (GHI) over the contiguous United States (CONUS). The main steps of the statistical downscaling method include (1) regridding RCM output (0.22 degree and daily resolutions) to handle the modeled-observed data sets on a common grid, (2) correcting bias of RCM GHI using satellite-derived observation, and (3) implementing temporal and spatial downscaling to generate GHI at 8-km and hourly resolution. Basically, complex physical processes and interactions between solar radiation and various atmospheric constituents lead solar irradiance to be highly variable and uncertain. Underrepresentation of clouds from the RCM parameterizations is the main source of error and uncertainty in modeling solar irradiance. Thus, we adapt and use the high-quality satellite-derived data from the National Solar Radiation Database (NSRDB) to analyze the bias and error of RCM GHI as well as estimate the statistical parameters for spatial and temporal downscaling. This presentation will summarize the comprehensive analysis conducted to produce and assess the results under two climate scenarios (RCP4.5 and RCP8.5). We will also present a detailed validation demonstrating the strengths of the proposed downscaling method and future extension of this research.

climate data

The National Climate Database (NCDB): An Unbiased 100-Year Dataset for PV Modeling

In this study, we develop a statistical technique to downscale the future projection of solar irradiance for photovoltaics (PV) energy-related applications. A set of Regional Climate Model (RCM)-based projections obtained from the North American Coordinated Regional Climate Downscaling Experiment (NA-CORDEX) are used as inputs to statistical methods to generate high-resolution global horizontal irradiance (GHI) over the contiguous United States (CONUS). The main steps of the statistical downscaling method include (1) regridding RCM output (0.22 degree and daily resolutions) to handle the modeled-observed data sets on a common grid, (2) correcting bias of RCM GHI using satellite-derived observation, and (3) implementing temporal and spatial downscaling to generate GHI at 8-km and hourly resolution. Basically, complex physical processes and interactions between solar radiation and various atmospheric constituents lead solar irradiance to be highly variable and uncertain. Underrepresentation of clouds from the RCM parameterizations is the main source of error and uncertainty in modeling solar irradiance. Thus, we adapt and use the high-quality satellite-derived data from the National Solar Radiation Database (NSRDB) to analyze the bias and error of RCM GHI as well as estimate the statistical parameters for spatial and temporal downscaling. This presentation will summarize the comprehensive analysis conducted to produce and assess the results under two climate scenarios (RCP4.5 and RCP8.5). We will also present a detailed validation demonstrating the strengths of the downscaling method, a summary of the 100-year dataset from 2001-2100, and future extension of this research.

bias correction

Optimal binning of correlated measurements

Experimental measurements are commonly represented on a discrete grid, requiring a balance between granularity and statistical noise. Two strategies have traditionally been used to improve such representations: selecting an appropriate bin width to control discretization error and applying kernel-based smoothing to suppress fluctuations. Despite their shared goal, these approaches have largely developed independently, without a unified statistical description of how discretization and correlation jointly determine measurement precision. Here, we extend the discussion of optimal interval averaging to a correlation-aware setting by Gaussian process regression, which explicitly accounts for correlations among neighboring bins. Starting from first principles, we derive the mean-squared error of discretized measurements and obtain closed-form asymptotic expressions for the optimal bin width and correlation length. When recast in reduced variables, the theory reveals distinct universal scaling laws governing the error in the correlation-free and correlation-controlled regimes. Characterized by intrinsically smooth intensity profiles and counting-based statistics, neutron scattering measurements are well suited for demonstrating the enhanced error contraction enabled by inter-bin correlations. We show that such improvement is achievable over the experimentally accessible Q-range and across multiple instruments and material systems. These results show that explicitly accounting for correlations systematically reshapes the limits of precision in discretized, noise-limited measurements. More broadly, the framework provides a transferable statistical foundation for optimizing data representation, inference, and experimental design across the physical and data sciences.

Tung, Chi-Huan [ORNL] (ORCID:0000000221972074)

Leveraging interpolation models and error bounds for verifiable scientific machine learning

Effective verification and validation techniques for modern scientific machine learning workflows are challenging to devise. Statistical methods are abundant and easily deployed, but often rely on speculative assumptions about the data and methods involved. Error bounds for classical interpolation techniques can provide mathematically rigorous estimates of accuracy, but often are difficult or impractical to determine computationally. Here, in this work, we present a best-of-both-worlds approach to verifiable scientific machine learning by demonstrating that (1) multiple standard interpolation techniques have informative error bounds that can be computed or estimated efficiently; (2) comparative performance among distinct interpolants can aid in validation goals; (3) deploying interpolation methods on latent spaces generated by deep learning techniques enables some interpretability for black-box models. We present a detailed case study of our approach for predicting lift-drag ratios from airfoil images. Code developed for this work is available in a public Github repository.

97 MATHEMATICS AND COMPUTING

Examining the Impact of Local Constraint Violations on Energy Computations in DFT

ABSTRACT This work examines the impact of locally imposed constraints in Density Functional Theory (DFT). Using a metric referred to as the extent of violation index (EVI), we quantify how well exchange‐correlation functionals adhere to local constraints. Applying EVIs to a diverse set of molecules for GGA functionals reveals constraint violations, particularly for semi‐empirical functionals. We leverage EVIs to explore potential connections between these violations and errors in chemical properties. While no correlation is observed for atomization energies, a significant statistical correlation emerges between EVIs and total energies. Similarly, the analysis of reaction energies suggests weak positive correlations for specific constraints. However, definitive conclusions about error cancellation mechanisms cannot be made at this time. These observations revealed by EVIs may be useful for consideration when designing future generations of semilocal functionals.

Khanna, Vaibhav [Department of Chemistry Universit

Hierarchical Bayesian Inverse Problems: A High-Dimensional Statistics Viewpoint

This paper analyzes hierarchical Bayesian inverse problems using techniques from highdimensional statistics. Furthermore, our analysis leverages a property of hierarchical Bayesian regularizers that we call approximate decomposability to obtain non-asymptotic bounds on the reconstruction error attained by maximum a posteriori estimators. The new theory explains how hierarchical Bayesian models that exploit sparsity, group sparsity, and sparse representations of the unknown parameter can achieve accurate reconstructions in high-dimensional settings.

MAP estimation

Robust error calibration for serial crystallography

Serial crystallography is an important technique with unique abilities to resolve enzymatic transition states, minimize radiation damage to sensitive metalloenzymes and perform de novo structure determination from micrometre-sized crystals. This technique requires the merging of data from thousands of crystals, making manual identification of errant crystals unfeasible. cctbx.xfel.merge uses filtering to remove problematic data. However, this process is imperfect, and data reduction must be robust to outliers. We add robustness to cctbx.xfel.merge at the step of uncertainty determination for reflection intensities. This step is a critical point for robustness because it is the first step where the data sets are considered as a whole, as opposed to individual lattices. Robustness is conferred by reformulating the error-calibration procedure to have fewer and less stringent statistical assumptions and incorporating the ability to down-weight low-quality lattices. We then apply this method to five macromolecular XFEL data sets and observe the improvements to each. The appropriateness of the intensity uncertainties is demonstrated through internal consistency. This is performed through theoretical CC 1/2 and I /σ relationships and by weighted second moments, which use Wilson's prior to connect intensity uncertainties with their expected distribution. This work presents new mathematical tools to analyze intensity statistics and demonstrates their effectiveness through the often underappreciated process of uncertainty analysis.

Mittan-Moreau, David W.