Search NASA⌕ Search

SEARCH · Search NASA

Results for “error estimation and analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Dynamic data-driven multiscale modeling for predicting the degradation of a 316L stainless steel nuclear cladding material

Here, we have developed a long short-term memory stacked ensemble (LSTM-SE) surrogate modeling approach that can provide rapid predictions of microstructural evolution and the resultant mechanical properties of American Iron and Steel Institute (AISI) 316L series stainless steel (316LSS) fuel cladding under conditions of varying temperature and radiation dose rate. To acquire training data, we developed and implemented a kinetic Monte Carlo (KMC) model to simulate precipitation kinetics of M 23 C 6 , γ', and G phases within SS316L cladding. Experimentally reported precipitation kinetics of SS316L in literature were linked to the kinetic parameters of the simulated precipitation in our KMC model. The model was then used to simulate microstructure evolution under synthetically generated treatments of varying temperature and radiation dose rate, for periods of up to 3000 hours. Changes in volume fraction, number density, and particle size of precipitates were recorded, and particle area fractions were correlated using statistical methods to develop the surrogate model. Simultaneously, the mechanical properties of the simulated microstructures were evaluated using microstructure-based finite element method (FEM) analysis to determine the elastic modulus, yield stress, ultimate tensile strength, and elongation to failure of the aged microstructures. Using this approach, our surrogate model can predict precipitation behavior within 0.25% volume fraction and mechanical properties within 6% relative error from the values predicted by the KMC and FEM models using 50 training simulations as input. The trained recurrent neural network-based model can return estimations of precipitation kinetics and mechanical properties ~1000 times faster than the physics-based codes. This work demonstrates, as a proof of concept, that reactor material service lifetimes under variable service conditions can be predicted for a statistics-based model from a practicably obtainable dataset.

36 MATERIALS SCIENCE↗

Baryon fraction from the BAO amplitude: a consistent approach to parameterizing perturbation growth

Galaxy clustering constrains the baryon fraction Omega_b/Omega_m through the amplitude of baryon acoustic oscillations and the suppression of perturbations entering the horizon before recombination. This produces a different pre-recombination distribution of baryons and dark matter. After recombination, the gravitational potential responds to both components in proportion to their mass, allowing robust measurement of the baryon fraction. This is independent of new-physics scenarios altering the recombination background (e.g. Early Dark Energy). The accuracy of such measurements does, however, depend on how baryons and CDM are modeled in the power spectrum. Previous template-based splitting relied on approximate transfer functions that neglected part of information. We present a new method that embeds an extra parameter controlling the balance between baryons and dark matter in the growth terms of the perturbation equations in the CAMB Boltzmann solver. This approach captures the baryonic suppression of CDM prior to recombination, avoids inconsistencies, and yields a clean parametrization of the baryon fraction in the linear power spectrum, separating out the simple physics of growth due to the combined matter potential. We implement this framework in an analysis pipeline using Effective Field Theory of Large-Scale Structure with HOD-informed priors and validate it against noiseless LCDM and EDE cosmologies with DESI-like errors. The new scheme achieves comparable precision to previous splitting while reducing systematic biases, providing a more robust way to baryon-fraction measurements. In combination with BBN constraints on the baryon density and Alcock-Paczynski estimates of the matter density, these results strengthen the use of baryon fraction measurements to derive a Hubble constant from energy densities, with future DESI and Euclid data expected to deliver competitive constraints.

Crespi, Andrea [U. Waterloo (main); Waterloo U., I↗

Bayesian calibration and uncertainty quantification of a rate-dependent cohesive zone model for polymer interfaces

In this work we present a rate-dependent cohesive zone model for the fracture of polymeric interfaces and performs a Bayesian calibration, an uncertainty quantification, and a sensitivity analysis for the model. The proposed cohesive zone model accounts for both reversible elastic and irreversible rate-dependent separation sliding deformation at the interface. The viscous dissipation due to the irreversible opening at the interface is modeled using elastic-viscoplastic kinematics that incorporates the effects of strain rate. Inverse calibration of parameters for such complex models through trial and error is challenging due to the large number of parameters of the model. Moreover, the calibrated parameter values are often non-unique and uncertain when the available experimental data is limited. To tackle this challenge, we employ a Bayesian calibration approach to identify parameters from experimental data, the resulting parameters significantly enhance the accuracy of the model. To quantify the uncertainty associated with the inverse parameter estimation, a modular Bayesian approach is employed to calibrate the unknown model parameters, accounting for the parameter uncertainty of the cohesive zone model. The advantages of the Bayesian calibration over a deterministic parameter fit are demonstrated. Further, to quantify the model uncertainties, such as incorrect assumptions or missing physics, a discrepancy function is introduced, which significantly improves the model’s prediction. Finally, the total uncertainty of the model is quantified in a predictive setting. A sensitivity analysis is performed to assess how changes in the input variables of the model affect the peak load, facilitating the identification of a concise set of highly influential parameters. The present approach can be used for calibration and uncertainty quantification for other complex computational mechanics models. It should also facilitate the designing of interface materials under uncertainty.

42 ENGINEERING↗

An Approach to Dynamic Human Reliability Analysis and Its Data Collection Framework

Human reliability analysis (HRA) is a method for evaluating human errors in a variety of complex systems such as nuclear power plants, military systems, aircraft, and chemical plants. Most HRA methods currently used by regulatory institutes or utilities are called static HRA and are carried out by simple worksheets or simple calculators. To date, there are many unsolved or intrinsic challenges in static HRA. For example, existing static HRA does not realistically model and evaluate human actions as they would be performed at actual systems. There is no method with HRA to objectively estimate the time required for human actions despite being essential to HRA processes. In addition, many HRA methods still rely on a dataset generated prior to the 1980s, from unrelated industry experience or simply from expert judgment. Accordingly, this study attempted to research how to overcome the challenges of existing HRA via dynamic risk assessment (a.k.a., simulation-based or computation-based risk assessment) techniques. First, this study developed a dynamic HRA method, named as PRocedure-based Investigation Method of EMRALD Risk Assessment – HRA (PRIMERA-HRA). The PRIMERA-HRA mainly concentrates on providing HRA analysts with specific guidelines on how to reasonably model human actions, assign human reliability data and evaluate output of simulation within a dynamic probabilistic risk assessment tool, called as Event Modeling Risk Assessment using Linked Diagrams (EMRALD). Second, this study also developed a module for performance shaping factors (i.e., the key concept in HRA quantification) applicable to dynamic HRA, then implemented it based on PRIMERA-HRA within the EMRALD tool. Third, this study developed an HRA data collection framework to support dynamic HRA, called as Simplified Human Error Experimental Program (SHEEP). Originally, the SHEEP study aimed to support static HRA and its data collection, but recently extended the scope to the new technologies such as dynamic HRA or HRA for advanced reactors. SHEEP focuses on the use of data collected from simplified simulators to complement—but not replace—data collection studies using full-scope simulators and actual operators. To date, many experiments were conducted under the SHEEP framework. Multiple analyses, such as human performance analysis, human error analysis, task complexity analysis, learning effect analysis and time distribution analysis, were also carried out using the collected data. Then, based on the major insights, an approach to inferring full-scope data based on simplified simulator data was proposed. The PRIMERA-HRA and SHEEP research are expected to evaluate human actions more realistically than existing static HRA, provide an opportunity to collect more HRA data with reasonable cost and labor, then contribute to enhance the quality of HRA.

99 - GENERAL AND MISCELLANEOUS↗

CFD modeling of natural circulation in LiCl-KCl molten salt closed loop

Characterizing flow within a molten salt closed-loop system is crucial for assessing system requirements, evaluating performance, and identifying potential flaws. Direct flow measurement using instrumentation is challenging due to extreme environmental conditions and the limitations associated with measuring molten salt flow under natural convection. Here, this study aims to provide comprehensive insights into the thermal-hydraulic behavior of a closed loop, with a particular focus on temperature distribution and velocity prediction. The Computational Fluid Dynamics (CFD) model demonstrated the capability to effectively simulate and predict both temperature distributions and flow velocities within the molten salt loop. The CFD model's predictive capability was validated by its ability to replicate temperature measurements under varying boundary conditions. The analysis revealed that the CFD model tends to underpredict temperatures in the cold leg and overpredict them in the hot leg, highlighting the need for continuous model refinement and acknowledging the limitations of using a steady-state approach. Furthermore, the potential of using external temperature measurements to estimate internal molten salt temperatures and predict flow velocity was explored, revealing that this approach could introduce up to a 5.5% error in flow velocity calculations. Line probes mapping temperature distributions across the tube's cross-section and molten salt provided valuable insights into temperature gradients, emphasizing the need for a thermal conductivity equation for molten salt with lower uncertainty to achieve more accurate temperature predictions of the system.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Optimality of Gradient-MUSIC for Spectral Estimation

We introduce the Gradient-MUSIC algorithm for estimating the unknown frequencies and amplitudes of a nonharmonic signal from noisy time samples. While the classical MUSIC algorithm performs a computationally expensive search over a fine grid, Gradient-MUSIC is significantly more efficient and eliminates the need for discretization over a fine grid by using optimization techniques. It coarsely scans the 1D landscape to find initialization simultaneously for all frequencies followed by parallelizable local refinement via gradient descent. We also analyze its performance when the noise level is sufficiently small and the signal frequencies are separated by at least 8π/m, where π/m is the standard resolution of this problem. Even though the 1D landscape is nonconvex, we prove a global convergence result for Gradient-MUSIC: coarse scanning provably finds suitable initialization and gradient descent converges at a linear rate. In addition to convergence results, we also upper bound the error between the true signal frequencies and amplitudes with those found by Gradient-MUSIC. For example, if the noise has $\ell^\infty$ norm at most ϵ, then the frequencies and amplitudes are recovered up to error at most Cϵ/m and Cϵ respectively, which are minimax optimal in m and ϵ. Our theory can also handle stochastic noise with performance guarantees under nonstationary independent Gaussian noise. Our main approach is a comprehensive geometric analysis of the landscape, a perspective that has not been explored before.

97 MATHEMATICS AND COMPUTING↗

Hyperspectral imaging for real-time waste materials characterization and recovery using endmember extraction and abundance detection

Hyperspectral imaging, combined with advanced spectral unmixing techniques and artificial intelligence, offers a powerful solution for improving material identification and classification. Here, this study evaluates the effectiveness of the pixel purity index and the sequential maximum angle convex cone algorithms in extracting and validating spectral signatures from pure samples of paper components (cellulose and lignin) and plastic (polypropylene). Principal-component analysis showed that both algorithms captured nearly all relevant variance for the tested materials. Spectral signatures were compared using the spectral angle mapper, revealing high similarity in the short-wave infrared region and greater variability in the visible near-infrared range. The methodology was then applied to a disposable coffee cup to detect and quantify mixed materials, accurately estimating material abundance and object area with less than 1% error. This approach enhances material classification, supporting product verification, quality control, and automated sorting for sustainable waste management and resource recovery.

36 MATERIALS SCIENCE↗

Association Kinetics for Perfluorinated n -Alkyl Radicals

Radical-radical reaction channels are important in the pyrolysis and oxidation chemistry of perfluoroalkyl substances (PFAS). In particular, unimolecular dissociation reactions within unbranched n-perfluoroalkyl chains, and their corresponding reverse barrierless association reactions, are expected to be significant contributors to the gas-phase thermal decomposition of families of species such as perfluorinated carboxylic acids and perfluorinated sulfonic acids. Unfortunately, experimental data for these reactions are scarce and uncertain. Furthermore, obtaining reliable theoretical predictions for such reactions is a laborious and computationally intensive task. Here, in this work, the chemical kinetics of the various association/decomposition reactions producing/decomposing the C 2 -C 4 series of unbranched n-perfluoroalkanes (C 2 F 6 , C 3 F 8 , and C 4 F 10 ) are examined using state-of-the-art ab initio transition-state-theory-based master-equation calculations. The variable-reaction-coordinate transition-state theory (VRC-TST) formalism is employed in computing the microcanonical and canonical rates for the association reactions. Reaction thermochemistry is obtained via composite quantum chemistry calculations and the laddering of error-canceling reaction schemes via a connectivity-based hierarchy approach employing ANL1/ANL0-style reference energies. Lennard-Jones collision model parameters for the considered systems were estimated by a direct dynamics approach, and collisional energy transfer parameters were obtained from analogies to systems of similar size and heavy-atom connectivity. A one-dimensional master equation approach was used to convert the microcanonical rate coefficients from the VRC-TST analysis into temperature- and pressure-dependent rate constants for the association reactions and the reverse dissociation reactions. The data are reported in standardized formats for usage in comprehensive chemical kinetic models for PFAS thermal destruction.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Diabatic error and propagation of Majorana zero modes in interacting quantum dots systems

Motivated by recent experimental progress in realizing Majorana zero modes (MZMs) using quantum dot systems, we investigate the diabatic errors associated with the movement of those MZMs. The movement is achieved by tuning time-dependent gate potentials applied to individual quantum dots, effectively creating a moving potential wall. To probe the optimized movement of MZMs, we calculate the experimentally accessible time-dependent fidelity and local density-of-states using many-body time-dependent numerical methods. Furthermore, our analysis reveals that an optimal potential wall height is crucial to preserve the well-localized nature of the MZM during its movement. Moreover, we analyze diabatic errors in realistic quantum-dot systems, incorporating the effects of repulsive Coulomb interactions and disorder in both hopping and pairing terms. Additionally, we provide a comparative study of diabatic errors arising from the simultaneous versus sequential tuning of multiple gates during the MZMs movement. Finally, we estimate the timescale required for MZM transfer in a six-quantum-dot system, demonstrating that MZM movement is feasible and can be completed well within the qubit's operational lifetime in practical quantum-dot setups.

Density of states↗

Quantifying Error in Photovoltaic Installation Metadata: Preprint

In this research, we quantify the level of metadata error for a fleet of 2860 photovoltaic (PV) systems, using metadata values provided by fleet owners. Using satellite imagery and time series analysis techniques available in open-source Python packages Panel-Segmentation and PVAnalytics, respectively, we evaluate the accuracy of PV system metadata such as location, azimuth, tilt, and mounting configuration (fixed tilt vs. tracking). We find that approximately 75% of provided latitude-longitude coordinates are within 190 meters of the actual solar installation. We were unable to link 7.8% of latitude-longitude coordinates to any solar installation via satellite imagery analysis. We evaluate the level of error in owner-provided mounting configuration (fixed tilt vs. single-axis tracking), finding only 8 systems with an incorrect mounting configuration. When evaluating azimuth and tilt parameters, we find that approximately 64% of the data is correct, with data for 860 systems (approximately 30%) not provided by system owners. To illustrate the importance of having correct solar metadata, we evaluate how incorrect metadata affects solar performance estimates by modeling system AC energy output at ground-truth vs. incorrect latitude-longitude coordinates, mounting configurations, and azimuth-tilt configurations. Energy output estimates can vary significantly if incorrect metadata parameters are used, with incorrect mounting configuration leading to the largest discrepancy with over 20% variation in expected energy output.

azimuth↗

A new method for measuring refractory corrosion of ceramics in glass

Abstract Nuclear waste glass vitrification furnaces are lined with refractory ceramic blocks to contain the molten glass. The refractory liner is susceptible to corrosion and has a finite service lifetime. For this reason, predicting the refractory corrosion in contact with molten glass is integral to estimating melter service lifetime. Standardized laboratory tests varying time and temperature are commonly performed to estimate refractory material loss as a function of glass composition. These data are time and resource‐intensive to collect and are susceptible to considerable measurement error. In order to accelerate glass formulation and design for nuclear waste vitrification, methods are needed to increase laboratory‐scale throughput while maintaining data quality. In this work, a method to remove the residual glass from a corroded coupon using hydrofluoric acid is presented that accelerates the throughput of sample analysis while simultaneously facilitating more accurate measurements.

Amoroso, Jake W. [Savannah River National Laborato↗

Propagating information content: an example with advection

The mathematical algorithm to derive geophysical information from remote sensing observations is called a retrieval. The mathematics of many retrieval problems are ill-posed, and thus a priori information is used to help constrain the derived geophysical variable to realistic values. One quantity of interest, therefore, is the information content of the observation. Perfect information content in the observation would be achieved if the retrieval were able to capture any perturbation in the desired geophysical variable with the proper magnitude. Many new data products can be derived by combining geophysical variables retrieved from multiple different remote sensors. This paper explores, for the first time, how to derive the information content of these derived products. The approach uses traditional error propagation techniques to derive the uncertainty of the derived field twice, both when the observations are used in the retrieval and also when only the a priori information from each remote sensor is propagated. These two uncertainties are then used to provide an estimate of the information content of the derived geophysical variable. This study demonstrates how to propagate the uncertainties from six different instruments to provide the information content for water vapor and temperature advection. A multi-month analysis demonstrates that, in a mean sense, the information content for temperature advection is nearly unity for all heights below 700 m while, the information content for water vapor advection is somewhat more variable but still larger than 0.6 in the convective boundary layer.

Turner, David D. [National Oceanic and Atmospheric↗

Harmonizing direct and indirect anthropogenic land carbon fluxes indicates a substantial missing sink in the global carbon budget since the early 20th century

Inconsistencies in the calculation of the two anthropogenic land flux terms of the global carbon cycle are investigated. The two terms—the direct anthropogenic flux (caused by direct human disturbance in anthromes, currently a carbon source to the atmosphere) and the indirect anthropogenic flux (caused indirectly by human activities that lead to global change and affecting all biomes, currently an atmospheric carbon sink)—are typically calculated independently, resulting in inconsistent underlying assumptions. We harmonize the estimation of the two anthropogenic land flux terms by incorporating previous estimates of these inconsistencies. We recalculate the global carbon budget (GCB) and apply change-point analysis to the cumulative budget imbalance. Cumulative over 1850–2018 (1959–2018), harmonization results in a 13% lesser (4% greater) land use source from anthromes and a 20% (23%) lesser land sink. This recalculation yields a greater non-closure of the GCB, indicating a missing carbon sink averaging 0.65 Pg C year -1 since the early 20th century. The imbalance likely results from a combination of method discontinuity and structural errors in the assessment of the direct anthropogenic land use flux, greater ocean carbon uptake, structural errors in land models, and in how these land terms are quantified for the budget. We caution against overconfidence in considering the GCB a solved problem and recommend further study of methodological discontinuities in budget terms. We strongly recommend studies that quantify the direct and indirect anthropogenic land fluxes simultaneously to ensure consistency, with a deeper understanding of human disturbance and legacy effects in anthromes.

54 ENVIRONMENTAL SCIENCES↗

Misclassification in Workers’ Telecommuting Frequency Choices Using a Generalized Extreme Value Model

Telecommuting frequency is a response variable collected in travel surveys and is, therefore, prone to errors leading to mismeasurements or misclassification. Misclassification of explanatory variables is a common risk when using statistical modeling techniques. We define “misclassification” as a response reported or recorded in the wrong category; for example, a variable is recorded as a 1 when it should be 0. Here, in this context, this study aims to develop a statistical model to analyze telecommuting data which accounts for potential misclassification errors by building on existing literature in econometrics. The empirical analysis was undertaken using the 2017 National Household Travel Survey (NHTS) and the general extreme value (GEV) models available in the literature. Specifically, the frequency of telecommuting days was analyzed using the negative binomial (NB) model recast as the multinomial logit (MNL) model. By nature—and consistent with other studies—NHTS data are prone to errors that can be classified as intentional or unintentional misinformation provided by the person being interviewed. Ignoring these errors while modeling telecommuting frequencies using standard discrete count models can result in biased parameter estimates. The misclassification parameter was calculated for both over-reporting and under-reporting scenarios. The misclassification errors can be as high as 14% over-reported and 10% under-reported, particularly for the neighboring values. Statistical fit comparison between the models shows that models that ignore misclassification have worse data fit and biased parameter estimates with significant policy implications.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Physics-based hybrid machine learning for critical heat flux prediction with uncertainty quantification

Critical heat flux (CHF) is a key quantity in nuclear system modeling due to its impact on heat transfer, safety margins, and reactor performance. This study develops and validates an uncertainty-aware hybrid modeling approach that combines machine learning with physics-based models to predict CHF in cases of dryout. The Biasi and Bowring empirical correlations were paired with three ML uncertainty quantification (UQ) techniques: deep neural network (DNN) ensembles, Bayesian neural networks (BNNs), and deep Gaussian processes (DGPs). A pure ML model without a base model was evaluated for comparison. Model performance was assessed under plentiful (7,350 points) and limited (9 points) training data scenarios using parity, uncertainty distributions, and calibration curves. Results show that the Biasi hybrid DNN ensemble achieved the best overall performance, with a mean absolute relative error of 1.846%, and well-calibrated uncertainty estimates. The BNN-based hybrids showed slightly higher error (2.14%) but superior uncertainty calibration. DGP models underperformed, with over 6% error and poor uncertainty calibration. All hybrid models outperformed pure machine learning configurations, demonstrating resistance against data scarcity. These findings indicate that hybrid modeling significantly improves predictive accuracy, interpretability, and resilience to data scarcity. The integration of uncertainty awareness provides actionable confidence in CHF predictions, which is vital for safety-critical decisions in nuclear applications. This hybrid approach offers a viable pathway for deploying ML models in reactor analysis tools while preserving domain knowledge and physical consistency.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Constraining gravity with a new precision 𝐸 𝐺 estimator using Planck + SDSS BOSS data

The 𝐸 𝐺 statistic is a discriminating probe of gravity developed to test the prediction of general relativity (GR) for the relation between gravitational potential and clustering on the largest scales in the observable Universe. We present a novel high-precision estimator for the 𝐸 𝐺 statistic using CMB lensing and galaxy clustering correlations that carefully matches the effective redshifts across the different measurement components to minimize corrections. A suite of detailed tests is performed to characterize the estimator’s accuracy, its sensitivity to assumptions and analysis choices, and the non-Gaussianity of the estimator’s uncertainty is characterized. After finalization of the estimator, it is applied to Planck CMB lensing and SDSS CMASS and LOWZ galaxy data. We report the first harmonic space measurement of 𝐸 𝐺 using the LOWZ sample and CMB lensing and also updated constraints using the final CMASS sample and the latest Planck CMB lensing map. We find $\hat{𝐸}$$^{Planck+CMASS}_{𝐺}$ = 0.3⁢6$^{+0.06}_{−0.05}$⁢(68.27%) and $\hat{𝐸}$$^{Planck+LOWZ}_{𝐺}$ = 0.4⁢0$^{+0.11}_{−0.09}$⁢(68.27%), with additional subdominant systematic error budget estimates of 2% and 3%, respectively. Using Ω m,0 constraints from Planck and SDSS BAO observations, Λ⁢CDM-GR predicts 𝐸$^{GR}_ {𝐺}$⁡(𝑧 =0.555) = 0.401 ± 0.005 and 𝐸$^{GR}_{𝐺}$⁡(𝑧 =0.316) = 0.452 ± 0.005 at the effective redshifts of the CMASS and LOWZ based measurements. We report the measurement to be in good statistical agreement with the Λ⁢CDM-GR prediction and report that the measurement is also consistent with the more general GR prediction of scale independence for 𝐸 𝐺 . Furthermore, this work provides a carefully constructed and calibrated statistic with which 𝐸 𝐺 measurements can be confidently and accurately obtained with upcoming survey data.

79 ASTRONOMY AND ASTROPHYSICS↗

On the Training and Generalization of Deep Operator Networks

Here, we present a novel training method for deep operator networks (DeepONets), one of the most popular neural network models for operators. DeepONets are constructed by two subnetworks, namely the branch and trunk networks. Typically, the two subnetworks are trained simultaneously, which amounts to solving a complex optimization problem in a high dimensional space. In addition, the nonconvex and nonlinear nature makes training very challenging. To tackle such a challenge, we propose a two-step training method that trains the trunk network first and then sequentially trains the branch network. The core mechanism is motivated by the divide-and-conquer paradigm and is the decomposition of the entire complex training task into two subtasks with reduced complexity. Therein the Gram–Schmidt orthonormalization process is introduced which significantly improves stability and generalization ability. On the theoretical side, we establish a generalization error estimate in terms of the number of training data, the width of DeepONets, and the number of input and output sensors. Numerical examples are presented to demonstrate the effectiveness of the two-step training method, including Darcy flow in heterogeneous porous media.

deep operator networks↗

Development and implementation of high-throughput proteomic and metabolomics assays by using advanced chromatographic and mass spectrometric systems (CRADA Final Report)

The mission of this CRADA with Agilent was to couple powerful MS platforms (QQQ, IM-QTOFMS) with Agilent’s novel Ultra-High-Performance Liquid Chromatography (UHPLC) fast metabolomic workflows and perform ABF Machine Learning (ML) to generated datasets. Agilent transferred UHPLC methods to PNNL and LBNL and methods were implemented and demonstrated in both labs, achieving total acquisition times of < 10 min. Metabolites analyzed using Agilent’s shared methods included metabolites from central carbon metabolism, common across hosts, and metabolites unique to engineered strains. Standards were acquired in an UHPLC-Drift Tube Ion Mobility Mass Spectrometer (DTIMS) system for the first time within the context of ABF and methods were optimized based on Agilent’s protocols. Samples from ABF hosts Pseudomonas putida, Aspergillus pseudoterreus, Aspergillus niger and Rhodosporidium toruloides were analyzed using the UHPLC-DTIMS platform for a total of 276 runs. A data analysis workflow compatible with the Experimental Data Depot (EDD) and completely shareable was developed for the acquired UHPLC-DTIMS data. Samples were analyzed using a Data Independent Acquisition Approach (DIA), which for most of the standards provided more transitions therefore increasing detection confidence. Using the data acquired by PNNL, LBNL, and Agilent’s specifications from previous ML projects, SNL applied an ensemble ML strategy to pick the best performing model for automated LC-method selection. Finally, with the contribution of the participant labs and Agilent, SNL developed an Automated Method Selection (AMS) software tool to predict the best liquid chromatography method for analysis of any new molecules of interest. Samples with novel pathways and new metabolite targets of interest are generated at a high pace in the ABF. Overall, the project advanced rapid metabolomics by combining liquid chromatography, ion mobility spectrometry, and data-independent mass spectrometry with machine learning. This multidimensional approach uses retention time, collision cross-section, precursor mass, and fragment-ion information to distinguish chemically similar metabolites that can be difficult to resolve using conventional liquid- or gas-chromatography methods. The resulting workflow also provided automated metabolite-identification error estimates, addressing a recognized need for statistical confidence measures in metabolomics.

Petzold, Christopher [Lawrence Berkeley National L↗