Search NASA⌕ Search

SEARCH · Search NASA

Results for “Statistical Methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Constraining Galaxy-Halo connection using machine learning

We investigate the potential of machine learning (ML) methods to model small-scale galaxy clustering for constraining Halo Occupation Distribution (HOD) parameters. Our analysis reveals that while many ML algorithms report good statistical fits, they often yield likelihood contours that are significantly biased in both mean values and variances relative to the true model parameters. This highlights the importance of careful data processing and algorithm selection in ML applications for galaxy clustering, as even seemingly robust methods can lead to biased results if not applied correctly. ML tools offer a promising approach to exploring the HOD parameter space with significantly reduced computational costs compared to traditional brute-force methods if their robustness is established. Using our ANN-based pipeline, we successfully recreate some standard results from recent literature. Properly restricting the HOD parameter space, transforming the training data, and carefully selecting ML algorithms are essential for achieving unbiased and robust predictions. Among the methods tested, artificial neural networks (ANNs) outperform random forests (RF) and ridge regression in predicting clustering statistics, when the HOD prior space is appropriately restricted. We demonstrate these findings using the projected two-point correlation function (w p (r p )), angular multipoles of the correlation function (ξ ℓ (r)), and the void probability function (VPF) of Luminous Red Galaxies from Dark Energy Spectroscopic Instrument mocks. Our results show that while combining w p (r p ) and VPF improves parameter constraints, adding the multipoles ξ 0 , ξ 2 , and ξ 4 to w p (r p ) does not significantly improve the constraints.

cosmology↗

Hawai‘i Supernova Flows: a peculiar velocity survey using over a Thousand Supernovae in the near-infrared

ABSTRACT We introduce the Hawai‘i Supernova Flows project and present summary statistics of the first 1217 astronomical transients observed, 668 of which are spectroscopically classified Type Ia Supernovae (SNe Ia). Our project is designed to obtain systematics-limited distances to SNe Ia while consuming minimal dedicated observational resources. To date, we have performed almost 5000 near-infrared (NIR) observations of astronomical transients and have obtained spectra for over 200 host galaxies lacking published spectroscopic redshifts. In this survey paper, we describe the methodology used to select targets, collect/reduce data, calculate distances, and perform quality cuts. We compare our methods to those used in similar studies, finding general agreement or mild improvement. Our summary statistics include various parametrizations of dispersion in the Hubble diagrams produced using fits to several commonly used SN Ia models. We find the lowest dispersions using the SNooPy package’s EBV_model2, with a root mean square deviation of 0.165 mag and a normalized median absolute deviation of 0.123 mag. The full utility of the Hawai‘i Supernova Flows data set far exceeds the analyses presented in this paper. Our photometry will provide a valuable test bed for models of SN Ia incorporating NIR data. Differential cosmological studies comparing optical samples and combined optical and NIR samples will have increased leverage for constraining chromatic effects like dust extinction. We invite the community to explore our data by making the light curves, fits, and host galaxy redshifts publicly accessible.

Do, Aaron (ORCID:0000000334297845)↗

Machine Learning Approaches to Predicting Induced Seismicity and Imaging Geothermal Reservoir Properties

This project developed machine learning (ML) methods, lab data sets, and field data to advance geothermal exploration and geothermal energy production. The work had three focus areas. One involved the development of ML methods to use microearthquakes (MEQs) for imaging geothermal reservoir properties and improving subsurface characterization – most importantly the evolution of permeability within the evolving reservoir. This part of the work included development of ML approaches for automated MEQ location, focal mechanism determination and identification of earthquake precursors. The second area focused on using MEQ signals generated by geothermal exploration and production to predict the relationship between fluid injection and seismicity. Here, we extended to reservoir scale our success in using ML to predict laboratory earthquakes and fault zone stress state. The third focus area was on lab experiments. Here, we developed new ML models for lab earthquake prediction and identification of precursors to failure to improve earthquake forecasting and early warning in geothermal settings. Major outcomes of our work include ML models that learn from MEQ signals during geothermal exploration and production to predict induced seismicity. MEQs occur naturally in connection with drilling and energy production. We developed ML methods to use the seismic waves from these events to characterize the elastic, hydraulic and poromechanical properties of reservoirs. Our work illuminated fracture geometry and the evolution of fracture permeability by incorporating seismic coda wave analysis and ML methods to relate fluid injection and seismicity. We significantly expanded laboratory earthquake prediction to include methods that use both passive measurements of microearthquakes within the lab fault zones and also active source acoustic measurements of fault zone elastic properties. These methods can now predict fault zone stress state, time to failure and the magnitude of lab earthquakes. Our work showed that repetitive stick- slip failure events during frictional sliding (the lab equivalent of earthquakes) are preceded by a cascade of micro-failure events that radiate energy in a manner that foretells unstable failure – manifest as laboratory MEQs. We documented a mapping between fracture properties and statistical attributes of elastic radiation. We extended existing works to geothermal reservoir scale and developed ML methods to determine reservoir permeability, fracture properties, and their evolution during geothermal energy production. An attractive feature of ML algorithms is their ability to handle big datasets and reveal patterns and correlations that may remain invisible to conventional analyses. Our work connected data from field, laboratory and intermediate scales to study permeability, stress, strength, fracture stiffness and geometry. At the field scale we used data from the Newberry Volcano field site, UtahFORGE, EGS Collab, and also the Bedretto underground research lab in Switzerland. These data sets are bridging the gap between the lab scale, theory, and reservoir scale. Our work produced plain language summaries to improve public understanding of DOE research. We also developed openly distributed ML and seismicity datasets for use by all researchers and we published connections between induced seismicity in geothermal areas and reservoir properties including permeability, fracture properties, and stress state. Our models are designed for the large data sets of induced seismicity typically associated with geothermal sites. We produced labeled event catalogs and used them on geothermal data to assess how ML can facilitate geothermal production and exploration. All datasets are available on the GDR Productivity: The project produced 32 publications in peer reviewed journals (two are in review). It supported the work of 6 PhD students, 40 conference presentations, 6 keynote talks at national meetings, and mentoring and professional development for 4 postdoctoral fellows.

15 GEOTHERMAL ENERGY↗

Determining reference standard strength for neutron-irradiated reduced activation ferritic/martensitic steel F82H by Bayesian method

The deterministic approach widely adopted in the design of structural components relies on systematically defined design limits using empirically determined safety factors. However, this approach is not always appropriate because structures are subjected to a variety of loads in the practical environment, which may result in excessively conservative design limits. In recent years, a more rigorous probabilistic approach that incorporates material strength distributions has become an important solution. In the probabilistic approach, the probability density functions of material strength properties underpin the design criteria. Here, the objective of this study is to identify the density distribution functions that best describe tensile properties of irradiated F82H to define a reference strength for DEMO design. Due to the limited number of existing data, this study specifically employs a Bayesian prediction method based on Monte Carlo simulations to determine a material reference value with statistical reliability and to investigate its effectiveness. For example, the dependence of tensile properties of 300 °C irradiated materials on irradiation damage and the range predicted by 95% Bayesian estimation was evaluated. As a statistical model for the dose dependence of statistical parameters, the normal distribution exhibited a better fit for 0.2% proof strength and tensile strength, whereas the distribution of total elongation data gave comparable reference values for both the normal and Weibull distribution models. Both models gave comparable criteria for the distribution of total elongation data. The Weibull model also gave better results for uniform elongation. The function best describing the model was a logarithmic law for both 0.2% proof strength and tensile strength, while a power law for both total and uniform elongation, which allowed for more comprehensive data prediction of irradiation data with statistical accuracy for DEMO reactor design.

36 MATERIALS SCIENCE↗

How are Heterogeneous Nucleation Rate Observations Influenced by Instrument Resolution?

Experimental measurements of the heterogeneous nucleation rate rely on counting the number of nuclei with time. However, the size of a thermodynamically stable nucleus is often a few nanometers in diameter and is below the resolution of most (in situ) measurement techniques that provide a statistically valid sample. Due to the finite resolution of the instruments and analysis methods, it is challenging to capture the incipient nuclei and the subsequent evolution of nuclei density over time. In this work, we demonstrate the impact of instrument resolution on observed nuclei densities by comparing numerical modeling with experimental results. Further, to achieve this, we implemented heterogeneous nucleation within the pore-scale reactive transport modeling framework using classical nucleation theory (CNT). We compared the modeling results with nucleation rates measured using X-ray nanotomography (XnT) and evaluated how these impact the apparent values of the prefactor and interfacial energy based on CNT and the crystal growth rate. Specifically, we applied a resolution threshold (artificial resolution limit) in the model during nuclei counting to resemble an experimental resolution, ranging from 15 to 500 nm. The findings reveal that the instrument resolution significantly impacts the apparent prefactor and interfacial energy. Both apparent prefactor and interfacial energy decrease with a decrease in the instrument resolution. While deviation in the prefactor due to resolution is anticipated, those in the interfacial energy are unexpected. The approach described here allows one to correct apparent nucleation rates that depend on the instrument’s resolution to derive “intrinsic” CNT parameters for the prefactor and interfacial energy.

47 OTHER INSTRUMENTATION↗

Transport coefficient approach for characterizing nonequilibrium dynamics in soft matter

Nonequilibrium states in soft condensed matter require a systematic approach to characterize and model materials, enhancing predictability and applications. Among the tools, X-ray photon correlation spectroscopy (XPCS) provides exceptional temporal and spatial resolution to extract dynamic insight into the properties of the material. However, existing models might overlook intricate details. We introduce an approach for extracting the transport coefficient, denoted as $J(t)$, from the XPCS studies. This coefficient is a fundamental parameter in nonequilibrium statistical mechanics and is crucial for characterizing transport processes within a system. Our method unifies the Green–Kubo formulas associated with various transport coefficients, including gradient flows, particle–particle interactions, friction matrices, and continuous noise. We achieve this by integrating the collective influence of random and systematic forces acting on the particles within the framework of a Markov chain. We initially validated this method using molecular dynamics simulations of a system subjected to changes in temperatures over time. Subsequently, we conducted further verification using experimental systems reported in the literature and known for their complex nonequilibrium characteristics. The results, including the derived $J(t)$ and other relevant physical parameters, align with the previous observations and reveal detailed dynamical information in nonequilibrium states. This approach represents an advancement in XPCS analysis, addressing the growing demand to extract intricate nonequilibrium dynamics. Further, the methods presented are agnostic to the nature of the material system and can be potentially expanded to hard condensed matter systems.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Seeing is Believing: Autonomous Microscopy and the Data Revolution in Materials Science

This presentation explores the transformative potential of autonomous electron microscopy and artificial intelligence (AI) in accelerating materials science discovery, particularly for energy applications and materials operating in extreme environments. We discuss pioneering self-driving laboratories at NREL designed to intelligently probe material synthesis and degradation across multiple scales, aiming to rapidly bridge the gap between atomic-level understanding and the development of high-performance, reliable materials. Utilizing advanced machine learning techniques, such as few-shot learning and multimodal analysis integrating imaging and spectroscopy, we demonstrate methods to extract actionable descriptors for material behavior, quantify complex microstructural evolution, and statistically link synthesis parameters to defect populations. This AI-driven approach promises to accelerate the creation of predictive materials tailored for specific missions, enabling faster development cycles and enhanced material assurance.

36 MATERIALS SCIENCE↗

Seeing is Believing: Autonomous Microscopy and the Data Revolution in Materials Science

This presentation explores the transformative potential of autonomous electron microscopy and artificial intelligence (AI) in accelerating materials science discovery, particularly for energy applications and materials operating in extreme environments. We discuss pioneering self-driving laboratories at NREL designed to intelligently probe material synthesis and degradation across multiple scales, aiming to rapidly bridge the gap between atomic-level understanding and the development of high-performance, reliable materials. Utilizing advanced machine learning techniques, such as few-shot learning and multimodal analysis integrating imaging and spectroscopy, we demonstrate methods to extract actionable descriptors for material behavior, quantify complex microstructural evolution, and statistically link synthesis parameters to defect populations. This AI-driven approach promises to accelerate the creation of predictive materials tailored for specific missions, enabling faster development cycles and enhanced material assurance.

97 MATHEMATICS AND COMPUTING↗

Raptor

Raptor is an efficient Python-based tool for predicting the formation and morphology of stochastic lack of fusion defects in metal AM processes. A major obstacle for the qualification and certification of additively manufactured parts in critical applications continues to be performance variability caused in part by porosity-related defects. High-fidelity process models that could predict these defect features are currently too computationally expensive for component-level analysis. To address this, Raptor employs a high-performance geometric method to model the dynamic melt pool rather than relying on computationally intensive thermal fluid dynamics. This allows Raptor to rapidly identify regions of unmelted material that correspond to lack of fusion pores. The efficiency of this approach significantly reduces the time and resources needed for generating 3D defect predictions, which enables users to conduct large-scale parameter studies and evaluate how process variations affect part quality. The framework offers operational flexibility; users can execute simulations through a simple command line interface or integrate core functions as a library within larger computational workflows. Simulation outputs include 3D porosity maps for visualization and tools for quantitative morphological analysis. These results are suitable for direct comparison with experimental characterization data from methods such as X-ray computed tomography and can be used for statistical process optimization.

Subraveti, Vamsi [Vanderbilt Univ., Nashville, TN ↗

Nucleon Structure Studies at Jefferson Lab and the Electron Ion Collider

The research programs of Thomas Jefferson Laboratory (JLab) and the future Electron- Ion Collider (EIC) focus on one of the main goals of strong interaction studies : understanding the structure of nucleons in terms of the quarks and gluons composing them. Their structure is encoded in functions such as Generalized Parton Distributions (GPDs), which describe how quarks and gluons’ transverse position and longitudinal momentum are distributed inside nucleons. GPDs allow to obtain three-dimensional pictures of nucleons and to understand some of their fundamental properties, such as their internal pressure or the emergence of their spin from the dynamics of the partons composing them. At JLab and the EIC, electron beams are used to probe nucleons. Measurement of reactions such as Deeply Virtual Compton Scattering (DVCS) allows access to GPDs. The first longitudinally polarized-target experiment of the CLAS12 program at Jlab took place in 2022-2023. Combining polarized electron beams and nucleon targets, this experiment offers unique access to observables that allow the measurement of different types of GPDs. In particular, the DVCS beam- and target-spin asymmetries for protons and neutrons in deuterium will be measured for the first time. They give access to kinds of GPDs that are still poorly known, and the comparison between proton and neutron data will allow the extraction of the flavor dependence of the structure of nucleons. Specific analysis methods have been implemented to work with a polarized nuclear target and are presented in this thesis. These methods allow to obtain preliminary results for the asymmetries, waiting for the complete statistics to be available. In the long term, the experimental program for the EIC has been established with a strong emphasis on the measurement of the structure of nucleons at high energy. Measurements of reactions such as DVCS impose strict requirements on the electromagnetic calorimeter that will allow to measure the energy of the scattered electrons and photons. This calorimeter, which is under development, will be based on scintillating crystals read by Silicon Photomultipliers (SiPMs). A new type of glass-based scintillating material was tested, evaluating the possibilities to meet the technical requirements concerning their light yield and resistance to radiation damage in particular. Several models of SiPMs have been characterized, demonstrating they can operate over the vast energy range necessary to address the physics case at the EIC and providing guidelines for developing their readout electronics.

Pilleux, Noemie↗

Bayesian Framework for Bioburden Density Estimation in Planetary Protection

To comply with the international planetary protection policy set forth by the Committee on Space Research and NASA Agency level requirements, spacecraft destined to biologically sensitive planetary bodies have to minimize terrestrial biological contamination. Analysis, testing and inspection are the standard forward verification activities that are used to demonstrate compliance with the biological contamination requirements. For testing of spacecraft surface areas, a swab or wipe sample is collected from surfaces prior to last access and subsequently processed in the lab using NASA Approved Planetary Protection Methods for Culture Based Assays. Raw data resulting from this assay is then statistically treated employing a mathematical paradigm stemming from the 1970’s Viking Lander Project to generate the bioburden density and total microbial bioburden present. This standard approach arbitrarily accounts for error and provides an upper conservative bound as it reports the maximum number of spores estimated to be present on flight hardware surfaces. A bioburden density estimate factors in the following variables: the observed bioburden count, representative volume processed, sampling efficiencies. Notably, to account for error in the approach, a 0 observed count is arbitrarily changed to a count of 1 for each hardware grouping. The data generated by spacecraft bioburden verification campaigns in the past have resulted in <80% of wipes and <90% of swabs containing a bioburden count of 0. As such, having a robust and well documented statistical approach for dealing with the probability of low incident rates is necessary to be able to estimate spacecraft bioburden. Being able to statistically describe the bioburden distribution and associated confidence level is a gamechanger for the development of bioburden allocations during mission design and will allow for tighter management of risk throughout spacecraft build. Thus, Empirical Bayes statistical approach was evaluated to estimate the microbial bioburden on spacecraft to mitigate the aforementioned mathematical concerns and provide a probabilistic bioburden distribution of the flight hardware surface. For application of this approach to performing bioburden calculations, a range of non-informative prior assumptions on hardware surfaces are explored for Bayesian analyses while informative priors using posterior distributions from prior assays are utilized for Empirical Bayes analyses. Several non-informative priors are currently under investigation to assess fitness including use of these priors to serve as a foundation to build off of NASA specification values or a basis of risk to account for unknowns during the integration and testing process. Informative priors under consideration are generated using sampled bioburden values from hardware originating within like processing environments (e.g. vendor cleaning process or similar assembly process), temporal spacecraft status events as a prediction for hardware cleanliness of future samples, and heritage system bioburden actuals to predict allocation for subsequent missions. Informative priors and probabilistic bioburden distributions are then validated using data sets from the Mars Exploration Rover, Mars Science Laboratory, and InSight missions. Using Empirical Bayes approach to generate a probabilistic bioburden distribution as demonstrated through mission use cases provides a valid approach for use in the end-to-end requirements verification process.

97 - MATHEMATICS AND COMPUTING↗

Probabilistic neural networks for improved analyses with phenomenological R -matrix

Here we present a method for measurement analyses based on probabilistic deep neural networks that provide several advantages over conventional analyses with phenomenological models. These include predicting physical quantities directly from data, the rapid generation of statistically robust uncertainties, and the ability to bypass some parameters that may induce ambiguities and complications in data analysis. As deep learning methods make predictions through “black boxes,” the uncertainty quantification is typically challenging. We use a probabilistic framework that provides thorough uncertainty quantification and is straightforward to follow in practice. With the network architecture based on the Transformer, we demonstrate the current method for predicting nuclear resonance parameters from scattering data using the phenomenological R-matrix model.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Thermonuclear 28 P(p, γ ) 29 S reaction rate and astrophysical implication in ONe nova explosion

An accurate 28 P(p, γ) 29 S reaction rate is crucial to defining the nucleosynthesis products of explosive hydrogen burning in ONe novae. Using the recently released nuclear mass of 29 S, together with a shell model and a direct capture calculation, we reanalyzed the 28 P(p, γ) 29 S thermonuclear reaction rate and its astrophysical implication. We focus on improving the astrophysical rate for 28 P(p, γ) 29 S based on the newest nuclear mass data. Our goal is to explore the impact of the new rate and associated uncertainties on the nova nucleosynthesis. We evaluated this reaction rate via the sum of the isolated resonance contribution instead of the previously used Hauser-Feshbach statistical model. The corresponding rate uncertainty at different energies was derived using a Monte Carlo method. Nova nucleosynthesis is computed with the 1D hydrodynamic code SHIVA. The contribution from the capture on the first excited state at 105.64 keV in 28 P is taken into account for the first time. We find that the capture rate on the first excited state in 28 P is up to more than 12 times larger than the ground-state capture rate in the temperature region of 2.5 × 10 7 K to 4 × 10 8 K, resulting in the total 28 P(p, γ) 29 S reaction rate being enhanced by a factor of up to 1.4 at ~1 × 10 9 K. In addition, the rate uncertainty has been quantified for the first time. It is found that the new rate is smaller than the previous statistical model rates, but it still agrees with them within uncertainties for nova temperatures. The statistical model appears to be roughly valid for the rate estimation of this reaction in the nova nucleosynthesis scenario. Using the 1D hydrodynamic code SHIVA, we performed the nucleosynthesis calculations in a nova explosion to investigate the impact of the new rates of 28 P(p, γ) 29 S. Our calculations show that the nova abundance pattern is only marginally affected if we use our new rates with respect to the same simulations but statistical model rates. Finally, the isotopes whose abundance is most influenced by the present 28 P(p, γ) 29 S uncertainty are 28 Si, 33,34 S, 35,37 Cl, and 36 Ar, with relative abundance changes at the level of only 3% to 4%.

Astronomy & Astrophysics↗

Resolving local ordering and structure in Mn x Ge 1- x Te alloys through thermodynamic ensembles of pair distribution functions

Characterizing local bonding environments in complex materials is essential for understanding and optimizing their properties. Equally as important is the ability to predict local motifs as a function of synthesis conditions, enhancing chemists’ ability to design properties into materials. In this study, we present an approach to leverage statistical mechanics to generate temperature- and energy-informed ensemble averaged pair distribution functions (PDFs). This method, which we have named Thermodynamic Ensemble Averages of PDFs for Ordering and Transformations (TEAPOT), utilizes density functional theory (DFT) to relax supercells while incorporating energetic penalties for local order, enabling accurate and computationally efficient analysis of local structure. We apply this method to the neutron PDF measurements of the pseudobinary MnTe–GeTe (MGT) alloy, demonstrating its capability to resolve complex local distortions and chemical ordering. Our results reveal detailed insights into phase transformations and local distortions driven by Mn substitution. For compositions that globally present as rock salt, our analysis reveals that Ge coordination geometry is heavily impacted by synthesis temperature. We propose that high temperature synthesis conditions promote a lowered Ge polyhedra distortion, promoting high charge carrier mobility due to the alignment of local and global structure. Incorporating statistical mechanics and computation into experimental analysis thus guides synthesis of tailored local structure.

36 MATERIALS SCIENCE↗

Generating mock galaxy catalogues for flux-limited samples like the DESI Bright Galaxy Survey

ABSTRACT Accurate mock galaxy catalogues are crucial to validate analysis pipelines used to constrain dark energy models. We present a fast HOD-fitting method which we apply to the AbacusSummit simulations to create a set of mock catalogues for the DESI Bright Galaxy Survey, which contain r-band magnitudes and $(g-r)$ colours. The halo tabulation method fits HODs for different absolute magnitude threshold samples simultaneously, preventing unphysical HOD crossing between samples. We validate the HOD fitting procedure by fitting to real-space clustering measurements and galaxy number densities from the MXXL BGS mock, which was tuned to the SDSS and GAMA surveys. The best-fitting clustering measurements and number densities are mostly within the assumed errors, but the clustering for the faint samples is low on large scales. The best-fitting HOD parameters are robust when fitting to simulations with different realizations of the initial conditions. When varying the cosmology, trends are seen as a function of each cosmological parameter. We use the best-fitting HOD parameters to create cubic box and cut sky mocks from the AbacusSummit simulations, in a range of cosmologies. As an illustration, we compare the ${}^{0.1}M_r\lt -20$ sample of galaxies in the mock with BGS measurements from the DESI one-percent survey. We find good agreement in the number densities, and the projected correlation function is reasonable, with differences that can be improved in the future by fitting directly to BGS clustering measurements. The cubic box and cut-sky mocks in different cosmologies are made publicly available.

79 ASTRONOMY AND ASTROPHYSICS↗

One-to-one aeroservoelastic validation of operational loads and performance of a 2.8 MW wind turbine model in OpenFAST

Abstract. This article presents a validation study of the popular aeroservoelastic code suite OpenFAST leveraging weeks of measurements obtained during normal operation of a 2.8 MW land-based wind turbine. Measured wind conditions were used to generate one-to-one turbulent flow fields (i.e., comparing simulation to measurement in 10 min increments, or bins) through unconstrained and constrained assimilation methods using the kinematic turbulence generators TurbSim and PyConTurb. A total of 253 bins of 10 min of normal turbine operation were selected for analysis, and a statistical comparison in terms of performance and loads is presented. We show that successful validation of the model was not strongly dependent on the type of inflow assimilation method used for mean quantities of interest, which had median modeling errors per wind-speed interval generally within 5 %–10 % of the measurement. The type of inflow assimilation method did have a larger effect on the fatigue predictions for blade-root flapwise and tower-base fore–aft quantities, which surprisingly saw larger errors from the assumed higher-fidelity assimilation methods. Avenues for further work are discussed and include possible improvements to the aerodynamic, structural, and controller modeling that may offer insight on the origin of the up to ∼ 40 % median overprediction of fatigue for these quantities.

17 WIND ENERGY↗

Statistically-driven Experimental Design to Improve Reference-free Quantification of Small Molecules by Liquid Chromatography-Mass Spectrometry

Non-targeted analysis of small molecules and metabolites in unknown, complex samples using liquid chromatography-tandem mass spectrometry remains challenging. One of the main bottlenecks is the extensive unannotated regions of metabolomics mass spectrometry data, resulting in knowledge gaps. Small molecule annotation in mass spectrometry data has conventionally relied on reference standards and libraries for compound identification and confirmation, which can constrain compound identification to those molecules already known, thus limiting the ability to discover new knowledge and new markers. Retention time prediction can facilitate and expedite unknown compound identification in non-targeted analysis of complex metabolomics samples. Additionally, accurate retention time predictions can also inform sample mixture design for LC-MS/MS analyses. However, current machine learning-based methods for retention time prediction are typically developed for specific chromatographic platforms and are not generalizable across scales. And while technologies and methods to improve reference-free metabolite identification for more comprehensive annotation of unknowns has received much attention, development of the same for quantitation without reference standards has been much more limited, despite its importance in toxicological, environmental, food safety, forensics, and clinical applications. We believe that a reference-free quantitation strategy that exploits mass spectrometry data already collected for reference-free identification can provide much more insight on unknowns, and move the metabolomics field for more complete unknowns characterization. As such, we pursue two efforts to improve upon current state-of-the-art methods in non-targeted analysis: (1) machine learning-based retention time prediction and (2) statistical design of experiments framework for reference-free quantitation. In this work, we develop and demonstrate (1) a generalizable retention time prediction capability across chromatographic conditions and scales, and (2) a statistical design-based framework for response factor contribution elucidation and reference-free quantitation. Evaluation of our retention time prediction model, PrediToR, showed approximately 24% improvement over current models, and we observed approximately 10X improvement in concentration estimation accuracy from our statistical design-based response factor model over a primarily ionization efficiency-based model. We expect that future efforts to improve upon these new capabilities will further advance non-targeted analysis of small molecules towards truly reference-free metabolomics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Robust error calibration for serial crystallography

Serial crystallography is an important technique with unique abilities to resolve enzymatic transition states, minimize radiation damage to sensitive metalloenzymes and perform de novo structure determination from micrometre-sized crystals. This technique requires the merging of data from thousands of crystals, making manual identification of errant crystals unfeasible. cctbx.xfel.merge uses filtering to remove problematic data. However, this process is imperfect, and data reduction must be robust to outliers. We add robustness to cctbx.xfel.merge at the step of uncertainty determination for reflection intensities. This step is a critical point for robustness because it is the first step where the data sets are considered as a whole, as opposed to individual lattices. Robustness is conferred by reformulating the error-calibration procedure to have fewer and less stringent statistical assumptions and incorporating the ability to down-weight low-quality lattices. We then apply this method to five macromolecular XFEL data sets and observe the improvements to each. The appropriateness of the intensity uncertainties is demonstrated through internal consistency. This is performed through theoretical CC 1/2 and I /σ relationships and by weighted second moments, which use Wilson's prior to connect intensity uncertainties with their expected distribution. This work presents new mathematical tools to analyze intensity statistics and demonstrates their effectiveness through the often underappreciated process of uncertainty analysis.

Mittan-Moreau, David W.↗