Search NASA⌕ Search

SEARCH · Search NASA

Results for “physics informed machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

Non-equilibrium entropy production and information dissipation in a non-Markovian quantum dot

This study measures trajectory-level entropy production and information dissipation in a driven, non-Markovian quantum dot using time-resolved optical dynamics and machine-learning-based analysis. Although not a 2D-material system, it is relevant because it demonstrates quantitative extraction of nonequilibrium dynamics from nanoscale optical fluctuations, which is conceptually connected to the proposed studies of transient charge and spin dynamics at interfaces.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A Novel Machine Learning Method for Surface PM2.5 Estimations from Geostationary Satellites

Particulate matter (PM) with a diameter of less or equal to 2.5 μm, known as PM , affects human health as it penetrates the respiratory system. The Environmental Protection Agency (EPA) measures the atmospheric concentration of PM using air quality monitors stationed throughout the Continental United States (CONUS). Such measurements are points on a spatial domain and therefore, might not be representative of the air quality at nearby areas considering that the composition of the atmosphere is highly variable from place to place. Satellite based AOD permits a spatially uniform means of estimating PM and new geostationary satellites provide high temporal and spatial resolution estimation of AOD. However, the concentration of PM is non-linearly dependent on other atmospheric parameters that include relative humidity, temperature, and height of the planetary boundary layer. This information may be estimated at similar spatial and temporal resolutions as AOD from numerical modeling such as from the National Oceanic and Atmospheric Administration’s (NOAA) High Resolution Rapid Refresh (HRRR) model which resolves near real-time atmospheric conditions over the CONUS. The estimation of PM concentration is a multi-parametric problem that considers the effect of temporal dependencies among the different parameters. Deep learning approaches are appropriate for such complex estimation problems as they intrinsically capture relations among multiple non-linear parameters. This study compares deep-learning methods to traditional regression analysis to demonstrate the capabilities of these methods in predicting PM2.5 concentrations. Additionally, a novel ensemble learning approach is employed to identify scientific processes that could further improve the estimation of PM concentration. Utilizing Long Short-Term Memory (LSTM) neural networks, which are suitable for multivariate time series estimation problems as they are capable of learning long-term dependencies, individual models are created for each EPA station and trained on the aforementioned dataset collocated over each station. Individual station models are merged if the model's performance is improved by reducing the root mean squared error (RMSE) metric. This ensemble training method ultimately reduces the RMSE value. Evaluation of these results provide insights into physical processes and related observable parameters that may contribute to PM concentrations. Identified parameters evaluated to be statistically different between the merged and unmerged models are expected to improve overall performance. These new parameters are then utilized for reevaluation of the deep learning methods with an extreme gradient boosting model with an RMSE of 5.5 providing the best results.

George Priftis↗

Machine learning and process-based modeling of spatiotemporal changes in active layer thickness across Alaska

Permafrost degradation poses a growing threat to infrastructure stability and ecosystem resilience in the rapidly warming Arctic. We investigated the spatiotemporal dynamics of active layer thickness (ALT) across Alaska by integrating field observations, environmental datasets, a physically based Stefan model, and machine learning (ML) techniques. Using weather projections from the Coupled Model Intercomparison Project Phase 6 under two Shared Socioeconomic Pathways (SSP 2-4.5 and SSP 5-8.5), we assessed ALT sensitivity to projected future weather conditions. The random forest (RF) model outperformed the Stefan approach in predicting ALT on the training dataset (R² = 0.84 vs. 0.53) but demonstrated lower generalizability on the test dataset (R² = 0.24 vs. 0.54). The root mean square error (RMSE) for the RF model for training and testing ranged from 14 to 22 cm, compared to 17 and 18 cm for the Stefan model. Variable importance analysis revealed that mean annual temperature and slope angle were the strongest predictors of ALT, accounting for 19% and 18% of the variance, respectively, followed by sediment transport index (14%) and stream power index (11%). Comparative analysis of baseline ALT predictions showed the Stefan model tended to project a thicker active layer (mean ± SD: 65 ± 16 cm), compared to the RF model (mean ± SD: 59 ± 8.8) cm). Both models indicated a latitudinal gradient in ALT, with shallower depths at higher latitudes. Projected ALT increases by 2100 were estimated at 3.3 ± 2.2 cm under SSP 2-4.5 and 5.9 ± 4.0 cm under SSP 5-8.5 for the ML model, whereas the Stefan model projected substantially larger increases of 13 ± 2.6 cm (SSP 2-4.5) and 28 ± 4.4 cm (SSP5-8.5). Spatial analysis showed the greatest ALT increases in northern Alaska, with relatively smaller changes in southern regions. These findings highlight the complex, multifactorial nature of ALT dynamics and the value of hybrid modeling approaches. As rising temperatures accelerate permafrost thaw, changes in ALT can disrupt ecosystems, damage infrastructures, and enhance the release of stored soil carbon, highlighting the urgent need for improved predictive capabilities to inform adaptation strategies in the Arctic.

Climate sciences↗

Chemical-Free Lithium Separation from High-Salinity Brines Using Model-Informed and Machine Learning-Optimized Multi-Column Zwitterionic Chromatography

Direct Lithium Extraction (DLE) technologies often struggle to produce high-purity lithium salts from high-salinity brines, as current approaches require chemical-based elution, regeneration, and precipitation steps, resulting in significant environmental footprints. A novel salt fractionation approach using carboxybetaine resin, known as zwitterionic chromatography (ZIC), has demonstrated that lithium ions can be separated from divalent cations under high-salinity conditions using only water as eluent, with no regeneration required. To enable continuous and scalable deployment of this approach, we developed a chemical-free Multi-column Zwitterionic Chromatography (MZC) process and its theoretical and process models. To predict and optimize this nontraditional separation system, we introduced a novel anti-Langmuir isotherm, and the isotherm parameters were estimated through a machine learning-driven optimization based on artificial neural network ensembles with numerical feasibility assessment. Using machine learning-driven optimization, the MZC process achieved 98.0% lithium recovery, 99.5 % Li/(Li + Mg + Ca) purity, a 31.3% productivity increase, and a 33% reduction in water use compared to batch operation. The proposed MZC process enables lithium separation at $0.6-1.2 kg-1 Li, with costs dominated by resin manufacturing, while offering lower separation costs and carbon footprint compared with conventional carbonation. Overall, these findings position the MZC process as an effective polishing step within scalable and sustainable lithium production pipelines.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Autonomous fabrication of tailored defect structures in 2D materials using machine learning-enabled scanning transmission electron microscopy

Materials with tailored quantum properties can be engineered from atomic-scale assembly techniques, but existing methods often lack the agility and accuracy to precisely and intelligently control the manufacturing process. Here, we demonstrate a fully autonomous approach for fabricating atomic-level defects using electron beams in scanning transmission electron microscopy (STEM) that combines advanced machine learning and automated beam control. As a proof of concept, we achieved controlled fabrication of MoS-nanowire (MoS-NW) edge structures by iterative and targeted exposure of MoS 2 monolayer to a focused electron beam to selectively eject sulfur atoms, utilizing high-angle annular dark-field (HAADF) imaging for feedback-controlled monitoring of structural evolution of defects. A machine learning framework combining a random forest model and a convolutional neural network (CNN) was developed to decode the HAADF image and accurately identify atomic positions and species. This atomic-level information was then integrated into an autonomous decision-making platform, which applied predefined fabrication strategies to instruct beam control about atomic sites to be ejected. The selected sites were subsequently exposed to a localized electron beam using an FPGA-controlled scan routine with precise control over beam positioning and duration. While the MoS-NW edge structures produced exhibit promising mechanical and electronic properties, the proposed methods to build the autonomous fabrication framework is material-agnostic and can be extended to other 2D materials for the creation of diverse defect structures and heterostructures beyond Mo S2 .

Engineering↗

Deep learning based event reconstruction for cyclotron radiation emission spectroscopy

The objective of the cyclotron radiation emission spectroscopy (CRES) technology is to build precise particle energy spectra. This is achieved by identifying the start frequencies of charged particle trajectories which, when exposed to an external magnetic field, leave semi-linear profiles (called tracks) in the time–frequency plane. Due to the need for excellent instrumental energy resolution in application, highly efficient and accurate track reconstruction methods are desired. Deep learning convolutional neural networks (CNNs) - particularly suited to deal with information-sparse data and which offer precise foreground localization—may be utilized to extract track properties from measured CRES signals (called events) with relative computational ease. In this work, we develop a novel machine learning based model which operates a CNN and a support vector machine in tandem to perform this reconstruction. A primary application of our method is shown on simulated CRES signals which mimic those of the Project 8 experiment—a novel effort to extract the unknown absolute neutrino mass value from a precise measurement of tritium β - -decay energy spectrum. When compared to a point-clustering based technique used as a baseline, we show a relative gain of 24.1% in event reconstruction efficiency and comparable performance in accuracy of track parameter reconstruction.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Causal discovery from data assisted by large language models

Knowledge-driven discovery of novel materials necessitates the development of causal models for property emergence. While in the classical physical paradigm, the causal relationships are deduced based on physical principles or via experiment, the rapid accumulation of observational data necessitates learning causal relationships between dissimilar aspects of material structure and functionalities based on observations. For this, it is essential to integrate experimental data with prior domain knowledge. Here, we demonstrate this approach by combining high-resolution scanning transmission electron microscopy data with insights derived from large language models (LLMs). By applying ChatGPT to domain-specific literature, such as arXiv papers on ferroelectrics, and combining the obtained information with data-driven causal discovery, we construct adjacency matrices for directed acyclic graphs that map the causal relationships between structural, chemical, and polarization degrees of freedom in Sm-doped BiFeO 3 . This approach enables us to hypothesize how synthesis conditions influence material properties and guides experimental validation. Furthermore, the ultimate objective of this work is to develop a unified framework that integrates LLM-driven literature analysis with data-driven discovery, facilitating the precise engineering of ferroelectric materials by establishing clear connections between synthesis conditions and their resulting material properties.

Causal inference↗

Data Science Shows that Entropy Correlates with Accelerated Zeolite Crystallization in Monte Carlo Simulations

We have performed a data science study of Monte Carlo simulation trajectories to understand factors that can accelerate formation of zeolite nanoporous crystals, a process that can take days or even weeks. In previous work, Monte Carlo simulations predicted and experiments confirmed that using a secondary organic structure-directing agent (OSDA) accelerates crystallization of all-silica LTA zeolite, with experiments finding a three-fold speedup [PCCP 24, 142-148 (2022)]. However, it remains unclear what physical factors cause the speed-up. Here, we apply data science to analyze the simulation trajectories to discover what drives accelerated zeolite crystallization in Monte Carlo going from a one-OSDA synthesis (1OSDA) to a two-OSDA version (2OSDA). We encoded simulation snapshots using the Smooth Overlap of Atomic Positions approach, which represents all 2- and 3-body correlations within a given cutoff distance. Principal component analyses failed to discriminate datasets of structures from 1OSDA and 2OSDA simulations, while the Support Vector Machine (SVM) approach succeeded at classifying such structures with an area-under-curve (AUC) score of 0.99 (where AUC = 1 is a perfect classification) with all 3-body correlations, and as high as 0.94 with only 2-body correlations. SVM decision functions reveal relatively broad / narrow histograms for 1OSDA / 2OSDA datasets, suggesting that the two simulations differ strongly in information heterogeneity. Informed by these results, we performed pair (2-body) entropy calculations during crystallization, resulting in entropy differences that semi-quantitatively account for the speedup observed in the previous Monte Carlo simulations. We conclude that altering synthesis conditions in ways that substantially changes the entropy of labile silica networks may accelerate zeolite crystallization, and we discuss possible approaches for achieving such acceleration.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Scattering-based structural inversion of soft materials via Kolmogorov–Arnold networks

Small-angle scattering techniques are indispensable tools for probing the structure of soft materials. However, traditional analytical models often face limitations in structural inversion for complex systems, primarily due to the absence of closed-form expressions of scattering functions. To address these challenges, we present a machine learning framework based on the Kolmogorov–Arnold Network (KAN) for directly extracting real-space structural information from scattering spectra in reciprocal space. This model-independent, data-driven approach provides a versatile solution for analyzing intricate configurations in soft matter. By applying the KAN to lyotropic lamellar phases and colloidal suspensions—two representative soft matter systems—we demonstrate its ability to accurately and efficiently resolve structural collectivity and complexity. Here, our findings highlight the transformative potential of machine learning in enhancing the quantitative analysis of soft materials, paving the way for robust structural inversion across diverse systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Introducing a multiscale feature integration network for inpainting with applications to enhanced CMB map reconstruction

We introduce a novel neural network, SkyReconNet, which combines the expanded receptive fields of dilated convolutional layers along with standard convolutions, to capture both the global and local features for reconstructing the missing information in an image. We implement our network to inpaint the masked regions in a full-sky cosmic microwave background (CMB) map. Inpainting CMB maps is a particularly formidable challenge when dealing with extensive and irregular masks, such as galactic masks which can obscure substantial fractions of the sky. The hybrid design of SkyReconNet leverages the strengths of standard and dilated convolutions to accurately predict CMB fluctuations in the masked regions by effectively utilizing the information from surrounding unmasked areas. During training, the network optimizes its weights by minimizing a composite loss function that combines the structural similarity index measure (SSIM) and mean squared error (MSE). SSIM preserves the essential structural features of the CMB, ensuring an accurate and coherent reconstruction of the missing CMB fluctuations, while MSE minimizes the pixelwise deviations, thus enhancing the overall accuracy of the predictions. The predicted CMB maps and their corresponding angular power spectra align closely with the targets, achieving the performance limited only by the fundamental uncertainty of cosmic variance. The network’s generic architecture enables application to other physics-based challenges involving data with missing or defective pixels, systematic artifacts, etc. In conclusion, our results demonstrate its effectiveness in addressing the challenges posed by large irregular masks, offering a significant inpainting tool not only for CMB analyses but also for image-based experiments across disciplines where such data imperfections are prevalent.

Cosmic microwave background↗

Data-Driven Kinetic Reaction Networks for Separation Chemistry

Understanding complex, multistep chemical reactions at the molecular level is a major challenge whose solution would greatly benefit the design and optimization of numerous chemical processes. The separation of rare-earth (4f) and actinide (5f) elements is an example where improving our chemical understanding is important for designing and optimizing new chemistries, even with a limited number of observations. Here, in this work, we leverage data-driven artificial intelligence and machine-learning approaches to develop kinetic reaction networks that describe the liquid–liquid extraction mechanism of uranium using N,N-di-2-ethylhexyl-isobutyramide (DEHiBA). Specifically, we compare and contrast the properties of two classes of models: (1) purely data-driven models that are regularized using chemistry-agnostic, L1 regression and (2) chemistry-informed models that are regularized using relative reaction energies provided by quantum mechanical calculations. We observe that purely data-driven models are unbiased, simple, and accurate in their predictions of experimental measurements when provided with sufficient data but are difficult to fully constrain and interpret. In contrast, chemistry-informed models exhibit significantly improved chemical interpretability and consistency, providing a detailed description of the separation process while achieving high accuracy through ensemble averaging. Overall, the dominant species predicted to be extracted into the organic phase is UO 2 (NO 3 ) 2 (DEHiBA) 2 , agreeing with experimental slope analysis, thermodynamic modeling, EXAFS, and crystal structures. This work demonstrates that leveraging the fundamental structure of the problem can lead to efficient learning schemes that provide both accurate predictions and chemical insights at a low computational cost.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Direct inference of nuclear equation-of-state parameters from gravitational-wave observations

The observation of neutron star mergers with gravitational waves (GWs) has provided a new method to constrain the dense-matter equation of state (EOS) and to better understand its nuclear physics. However, inferring nuclear microphysics from GW observations necessitates the sampling of EOS model parameters that serve as input for each EOS used during the GW data analysis. The sampling of the EOS parameters requires solving the Tolman–Oppenheimer–Volkoff (TOV) equations a large number of times—a process that slows down each likelihood evaluation in the analysis on the order of a few seconds. Here, we employ emulators for the TOV equations built using multilayer perceptron neural networks to enable direct inference of nuclear EOS parameters from GW strain data. Our emulators allow us to rapidly solve the TOV equations, taking in EOS parameters and outputting the associated tidal deformability of a neutron star in only a few tens of milliseconds. We implement these emulators in PyCBC to directly infer the EOS parameters using the event GW170817, providing posteriors on these parameters informed solely by GWs. We benchmark these runs against analyses performed using the full TOV solver and find that the emulators achieve speed ups of nearly two orders of magnitude, with negligible differences in the recovered posteriors. Additionally, we constrain the slope and curvature of the symmetry energy at the 90% upper credible interval to be $L$ sym ≲ 106 MeV and $K$ sym ≲ 26 MeV.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Cyber Resilience and Social Equity: Twin Pillars of a Sustainable Energy Future

This paper examines the intersection of security and accessibility within energy systems amidst the rise of grid modernization and digitization, especially considering the regulatory changes and the imperatives of inclusive energy strategies. It addresses the dual need for secure, resilient infrastructure and a commitment to mitigate energy poverty while maintaining equitable access to energy. Amid escalating cybersecurity and physical threats, the paper advocates for sustainable energy delivery systems that ensure robust defenses without compromising the goals of reducing energy poverty and ensuring energy security. This paper identifies the pressing need for Cyber-Informed Engineering (CIE) and Secure-by-Design (SbD) principles, highlighting how these strategies can protect critical infrastructure and democratize access to secure energy, particularly for disadvantaged communities. The analysis underscores the challenges presented by the expansion of attack surfaces, interoperability requirements, and grid-edge analytics, offering innovative solutions that leverage advanced technologies and data-driven insights. Furthermore, this paper addresses the workforce development gap, emphasizing the necessity for public-private partnerships and vendor engagement in creating a skilled cybersecurity workforce. This paper has a dual focus on both the technological aspect of cybersecurity and the social dimension of equity within the context of sustainable energy development. It suggests a comprehensive examination of how these two critical elements interact and support the overarching goal of a sustainable energy future.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Towards Autonomous Lunar Resource Excavation via Reinforcement Learning

To continue on a sustainable and flexible path, NASA needs to address the challenge of collecting and moving large amounts of regolith at the destination. NASA’s Regolith Advanced Surface Systems Operations Robot (RASSOR) is principally designed to mine and deliver regolith for In-Situ Resource Utilization (ISRU) processing. RASSOR’s design enables it to efficiently collect and deposit regolith, return collected material for processing, and myriad related ISRU activities. To reliably perform these operations on the lunar surface, RASSOR software and sensory systems need to be robust and maximize the information extracted from a reduced sensor payload. Herein, we present preliminary findings from the Intelligent Capabilities Enhanced RASSOR project. We created reduced-order simulation environments to develop autonomous trenching controllers via reinforcement learning and prototype state estimation architectures. The goal of reinforcement learning is for an agent to learn a policy (task strategy) through interactions with an environment. When the agent performs an action, a change occurs in environment state and a numerical reward is received which informs the agent whether the action performed was good or not. Since reinforcement learning algorithms learn through trial-and-error, a simulation is a desirable first environment for development and learning. We developed two simulations, the first is a 2D excavation simulation developed to facilitate parameter selection, and a 3D simulation developed using a game physics engine, to simulate simplified soil interactions and increase the fidelity of the dynamic models of the robotic agents. The development of this 3D simulation has enabled the training of additional sensing capabilities and research both at the granular mechanics and operations levels. We experimented with various virtual sensor payloads to identify a combination that enabled efficient excavation operation and learning. Our reward function is based on how much material is excavated per step. A penalty is also received for leaving the dig site and to smooth the acceleration of the drum arms. We implemented pseudo time-of-flight sensors to report distance from each drum to ground and the height above ground which was found to be more efficient than existing solutions. Our findings suggest that reinforcement learning for autonomous operations has learned viable trenching strategies within 3000 training episodes in our simplified 2D environment and helped identify desirable sensing capabilities, arrangements, and considerations such as the positioning of time-of-flight sensors. Future work includes expanding our simulation to more complex environments and scenarios, and transfer learning from simulation to RASSOR 2.0 hardware for deployment in the Regolith Test Bin at NASA's Kennedy Space Center.

rassor↗

Improving ideal MHD equilibrium accuracy with physics-informed neural networks

We present a novel approach to compute three-dimensional magnetohydrodynamic equilibria with isotropic pressure profiles and nested surfaces by parametrizing Fourier modes with artificial neural networks (NNs). The full nonlinear global force residual of single equilibria across the volume in real space is then minimized with first order optimizers and compared to equilibria computed by conventional solvers. Already, we observe competitive computational cost to arrive at the same minimum residuals computable with existing codes. With increased computational cost, lower minima of the residual are computable with the NNs than with any other tested solver, establishing a new lower bound for the force residual. We use minimally complex NNs, and we expect significant improvements for solving not only single equilibria with NNs, but also for creating NN models valid over continuous distributions of equilibria.

ideal magnetohydrodynamics↗

INSPIRED: Inelastic neutron scattering prediction for instantaneous results and experimental design

Inelastic neutron scattering (INS) has unique advantages in probing how atoms vibrate and how the vibrations propagate and interact. Such dynamic information is crucial in understanding various material properties, from heat capacity, thermal conductivity, phase transitions, and chemical reactions to more exotic quantum behavior. The analysis and interpretation of the INS spectra often start from a model structure of the sample, followed by a series of calculations to obtain the simulated spectra to compare with experiments. The conventional way to perform such calculations usually requires significant time, computing resources, and specialized expertise. Here, we present a new program named INSPIRED (Inelastic Neutron Scattering Prediction for Instantaneous Results and Experimental Design), which enables users to perform rapid INS simulations in several different ways on their personal computers in just a few clicks, with the crystal structure as the only input file. Specifically, the users can choose a pre-trained symmetry-aware neural network (coupled with an autoencoder) to predict the phonon density of states (DOS), 1D S(E) and 2D S(|Q|,E) spectra for any given structure. One can also choose an existing density functional theory (DFT) calculation from a database (containing over 12,000 crystals), and quickly obtain the simulated INS spectra for single crystals and powders. It is also possible to use pre-trained universal machine learning force fields to relax a given crystal structure, calculate the phonon dispersion and DOS, and, subsequently, the INS spectra. All these functions are implemented with a PyQt graphic user interface. Finally, we expect these new tools will benefit broad user communities and significantly improve the efficiency of experiment design, execution, and data analysis for INS.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Advancing Multiscale Simulation of Plasma-Surface Interfaces

We report the development of an atomistic-informed, surface-state-dependent predictive model for particle exchange in a carbon-tungsten plasma-surface interface. The predictive model uses machine learning (ML) techniques to learn the energy and angular distributions for particle exchange and rate functions for surface state evolution from molecular dynamics simulations of cumulative bombardment of tungsten by energetic carbon ions. Each predictive component is sensitive to the energy and trajectory of incident plasma species and the surface state. The surface state is represented by a set of surface state descriptors, which were derived from the atomistic surface state for each independent carbon bombardment event. These descriptors are representative of the composition and degree of amorphization of the outermost angstrom of surface material and were chosen to optimize predictive performance for particle exchange at the interface. The distributions for particle exchange (reflection/sputtering) are demonstrated to vary with each surface state descriptor, motivating the development of surface-state-dependent particle exchange models for plasma simulations. The performance of various ML methods was compared, including polynomial quantile regression, artificial neural networks, k-nearest neighbors, and random forest algorithms, with polynomial regression performing the best for interpolation and extrapolation of learned relationships. In addition to the particle exchange model, a neutral network was developed and used to identify data sufficiency throughout surface descriptor space, which will enable real-time feedback during future data production to ensure data is produced where it is most needed, and we provide commentary on improvements to the data production workflow for future endeavors.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Statistically-driven Experimental Design to Improve Reference-free Quantification of Small Molecules by Liquid Chromatography-Mass Spectrometry

Non-targeted analysis of small molecules and metabolites in unknown, complex samples using liquid chromatography-tandem mass spectrometry remains challenging. One of the main bottlenecks is the extensive unannotated regions of metabolomics mass spectrometry data, resulting in knowledge gaps. Small molecule annotation in mass spectrometry data has conventionally relied on reference standards and libraries for compound identification and confirmation, which can constrain compound identification to those molecules already known, thus limiting the ability to discover new knowledge and new markers. Retention time prediction can facilitate and expedite unknown compound identification in non-targeted analysis of complex metabolomics samples. Additionally, accurate retention time predictions can also inform sample mixture design for LC-MS/MS analyses. However, current machine learning-based methods for retention time prediction are typically developed for specific chromatographic platforms and are not generalizable across scales. And while technologies and methods to improve reference-free metabolite identification for more comprehensive annotation of unknowns has received much attention, development of the same for quantitation without reference standards has been much more limited, despite its importance in toxicological, environmental, food safety, forensics, and clinical applications. We believe that a reference-free quantitation strategy that exploits mass spectrometry data already collected for reference-free identification can provide much more insight on unknowns, and move the metabolomics field for more complete unknowns characterization. As such, we pursue two efforts to improve upon current state-of-the-art methods in non-targeted analysis: (1) machine learning-based retention time prediction and (2) statistical design of experiments framework for reference-free quantitation. In this work, we develop and demonstrate (1) a generalizable retention time prediction capability across chromatographic conditions and scales, and (2) a statistical design-based framework for response factor contribution elucidation and reference-free quantitation. Evaluation of our retention time prediction model, PrediToR, showed approximately 24% improvement over current models, and we observed approximately 10X improvement in concentration estimation accuracy from our statistical design-based response factor model over a primarily ionization efficiency-based model. We expect that future efforts to improve upon these new capabilities will further advance non-targeted analysis of small molecules towards truly reference-free metabolomics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗