Search NASA⌕ Search

SEARCH · Search NASA

Results for “model identification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Online Dynamic Mode Decomposition Based System Identification of Multi-Zone Building HVAC Systems

Many works have recently been conducted to reduce the electricity consumption of smart buildings and allow them to support various grid services. Most of these works require accurate system models for the various appliances in the building including heating, ventilation, and air conditioning (HVAC) units. In this paper, we investigate a recursive data-driven system identification strategy to construct the thermal model for a time-varying building with a multi-zone HVAC unit. The online dynamic mode decomposition (DMD)-based strategy is employed to identify the multi-zone thermal building dynamics, where a simple information update (rank-1) is selected to avoid computational complexity. The DMD-based identification strategy is validated using a real gymnasium building equipped with a 4-zone HVAC unit, and its performance is compared with that of the traditional nuclear-norm subspace identification (N2SID) strategy.

Wu, Tumin [University of Tennessee, Knoxville (UTK↗

TPCpp-10M: Simulated proton-proton collisions in a time projection chamber for AI foundation models

Scientific foundation models hold great promise for advancing nuclear and particle physics by improving analysis precision and accelerating discovery. Yet, progress in this field is often limited by the lack of openly available large scale datasets, as well as standardized evaluation tasks and metrics. Furthermore, the specialized knowledge and software typically required to process particle physics data pose significant barriers to interdisciplinary collaboration with the broader machine learning community. This work introduces a large, openly accessible dataset of 10 million simulated proton-proton collisions, designed to support self-supervised training of foundation models. To facilitate ease of use, the dataset is provided in a common NumPy format. In addition, it includes 70,000 labeled examples spanning three well defined downstream tasks: track finding, particle identification, and noise tagging, to enable systematic evaluation of the foundation model's adaptability. The simulated data are generated using the Pythia Monte Carlo event generator at a center of mass energy of $\sqrt{s}$ = 200 GeV and processed with Geant4 to include realistic detector conditions and signal emulation in the sPHENIX Time Projection Chamber at the Relativistic Heavy Ion Collider, located at Brookhaven National Laboratory. This dataset resource establishes a common ground for interdisciplinary research, enabling machine learning scientists and physicists alike to explore scaling behaviors, assess transferability, and accelerate progress toward foundation models in nuclear and high energy physics. The complete simulation and reconstruction chain is reproducible with the sPHENIX software stack. All data and code locations are provided under Data Accessibility.

Data Analysis, Statistics and Probability (physics↗

Digital image correlation and infrared thermography data for seven unique geometries of 304L stainless steel

Material Testing 2.0 (MT2.0) is a paradigm that advocates for the use of rich, full-field data, such as from digital image correlation and infrared thermography, for material identification. By employing heterogeneous, multi-axial data in conjunction with sophisticated inverse calibration techniques such as finite element model updating and the virtual fields method, MT2.0 aims to reduce the number of specimens needed for material identification and to increase confidence in the calibration results. To support continued development, improvement, and validation of such inverse methods—specifically for rate-dependent, temperature-dependent, and anisotropic metal plasticity models—we provide here a thorough experimental data set for 304L stainless steel sheet metal. The data set includes full-field displacement, strain, and temperature data for seven unique specimen geometries tested at different strain rates and in different material orientations. Commensurate extensometer strain data from tensile dog bones is provided as well for comparison. We believe this complete data set will be a valuable contribution to the experimental and computational mechanics communities, supporting continued advances in material identification methods.

36 MATERIALS SCIENCE↗

Responses of summer mesoscale convective systems to irrigation over the North China Plain based on convection-permitting model simulations

Extensive irrigation activities in the North China Plain (NCP) significantly influence regional weather and climate. However, previous studies focusing on the NCP were primarily based on coarse-resolution models, which are unable to explicitly resolve convection systems, causing large uncertainty in precipitation simulations. In this study, a convection-permitting model coupled with a dynamic irrigation scheme is utilized to investigate the impacts of irrigation on summertime mesoscale convective systems (MCSs) over the NCP. Sensitivity experiments with irrigation off and on are conducted for 5 summers and an MCS identification and tracking algorithm is applied to both satellite observations and model simulations. We find that incorporating irrigation in the model increases MCS precipitation, which agrees more with observations. The probability distributions of MCS lifetime, area, propagation speed, and intensity are all better simulated with irrigation. Irrigation increases the occurrence frequency of MCSs throughout the entire day. The nighttime increase is partly because of more frequent local initiation of MCS developed from isolated deep convection, while the daytime increase is mainly attributed to the changes in MCSs initiating elsewhere and then propagating to the NCP. On average, irrigation induces additional moisture that is more thermodynamically favorable for precipitation, but this effect is partially offset by the weakened ascending air motion primarily caused by irrigation surface cooling. Compared to weak MCS precipitation events, strong MCS precipitation events experience greater enhancement in precipitation intensity when including irrigation because the offset effect from the change in large-scale ascending air motion is insignificant. In addition, irrigation makes the variation of MCS precipitation intensity more correlated with the variation in ascending motion but less correlated with that in atmospheric moisture. Our results suggest the pronounced impacts of irrigation on MCSs over the NCP which should be included in numerical models to improve regional precipitation simulation and prediction.

54 ENVIRONMENTAL SCIENCES↗

IBR Short Circuit Modeling in ETAP

Mohammad Zadeh from ETAP presented various modeling improvements that included the following: • Address convergence issues by using complete Norton Equivalent model including the IBR series filter impedance. • Fault ride through identification within iterations while short circuit results have not been converged and how that may impact final SC results. • Options for FRT curve 1- using a timer to lock FRT logic vs a hysteresis in FRT curve to avoid unwanted toggling. • Report recent co-operation between software vendors (Aspen, CAPE and ETAP) for adopting a C++ interface for IBR blackbox SC modeling. There was lot of interest to know about the model improvements that are being considered.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Autoregressive distributed lag-based dynamic uniformity modeling and monitoring approaches for superconductor manufacturing

High-temperature superconductors (HTS), known for their high efficiency and low energy loss, have found profound applications across various fields, driving the demand for long, uniformly performing tapes. However, ensuring uniform performance over extended lengths of HTS tapes, often characterized by the consistency of critical current, remains challenging due to fluctuations in growth conditions during manufacturing. Here, to elucidate the mechanisms underlying variations in tape uniformity and enable real-time monitoring of associated parameters, we propose an Autoregressive Distributed Lag (ADL)-based Dynamic Uniformity Modeling and Monitoring (ADUM2) approach. This method integrates uniformity measurement, the identification of critical process parameters and real-time monitoring within the manufacturing process. The ADUM2 approach is applied to the advanced metal organic chemical vapor deposition (A-MOCVD) process, a pilot-scale method for superconductor manufacturing. Our model demonstrates superior performance compared to benchmark methods, accounting for over 80% of the total variance in the data and identifying 13 key process parameters influencing the uniformity of HTS tapes. This study offers significant insights into the high-temperature superconductor manufacturing process and holds the potential to facilitate the production of cost-effective, uniformly performing long superconducting tapes in the future.

autoregressive distributed lag analysis↗

Physics-Informed Active Learning With Simultaneous Weak-Form Latent Space Dynamics Identification

The parametric greedy latent space dynamics identification (gLaSDI) framework has demonstrated promising potential for accurate and efficient modeling of high-dimensional nonlinear physical systems. However, it remains challenging to handle noisy data. Here, to enhance robustness against noise, we incorporate the weak-form estimation of nonlinear dynamics (WENDy) into gLaSDI. In the proposed weak-form gLaSDI (WgLaSDI) framework, an autoencoder and WENDy are trained simultaneously to discover intrinsic nonlinear latent-space dynamics of high-dimensional data. Compared with the standard sparse identification of nonlinear dynamics (SINDy) employed in gLaSDI, WENDy enables variance reduction and robust latent space discovery, therefore leading to more accurate and efficient reduced-order modeling. Furthermore, the greedy physics-informed active learning in WgLaSDI enables adaptive sampling of optimal training data on the fly for enhanced modeling accuracy. The effectiveness of the proposed framework is demonstrated by modeling various nonlinear dynamical problems, including viscous and inviscid Burgers' equations, time-dependent radial advection, and the Vlasov equation for plasma physics. With data that contains 5%–10% Gaussian white noise, WgLaSDI outperforms gLaSDI by orders of magnitude, achieving 1%–7% relative errors. Compared with the high-fidelity models, WgLaSDI achieves 121 to 1779x speed-up.

97 MATHEMATICS AND COMPUTING↗

Latent space dynamics identification for interface tracking with application to shock-induced pore collapse

Capturing sharp, evolving interfaces remains a central challenge in reduced-order modeling, especially when data is limited and the system exhibits localized nonlinearities or discontinuities. Here, we propose LaSDI-IT (Latent Space Dynamics Identification for Interface Tracking), a data-driven framework that combines low-dimensional latent dynamics learning with explicit interface-aware encoding to enable accurate and efficient modeling of physical systems involving moving material boundaries. At the core of LaSDI-IT is a revised autoencoder architecture that jointly reconstructs the physical field and an indicator function representing material regions or phases, allowing the model to track complex interface evolution without requiring detailed physical models or mesh adaptation. The latent dynamics are learned through linear regression in the encoded space and generalized across parameter regimes using Gaussian process interpolation with greedy sampling. We demonstrate LaSDI-IT on the problem of shock-induced pore collapse in high explosives, a process characterized by sharp temperature gradients and dynamically deforming pore geometries. The method achieves relative prediction errors below 9% across the parameter space, accurately recovers key quantities of interest such as pore area and hot spot formation, and matches the performance of dense training with only half the data. This latent dynamics prediction was 10 6 times faster than the conventional high-fidelity simulation, proving its utility for multi-query applications. These results highlight LaSDI-IT as a general, data-efficient framework for modeling discontinuity-rich systems in computational physics, with potential applications in multiphase flows, fracture mechanics, and phase change problems.

Gaussian process↗

Assessment of uranium nitride interatomic potentials

Uranium mononitride (UN) is a promising nuclear fuel due to its high fissile density, high thermal conductivity, and suitability for reprocessing. In this study, two uranium nitride interatomic potentials are assessed: Tseplyaev and Starikov's angular-dependent potential and Kocevski et al.'s embedded atom model potential. Predictions of the thermophysical and elastic properties of UN, UN 2 , and α- and β-U 2 N 3 computed using both potentials are assessed and compared to available experimental data. Notably, the Tseplyaev potential performs better with the energetic aspects of UN, e.g., specific heat capacity and point defect formation energies, whereas the Kocevski potential performs better with the structural aspects of UN, e.g., thermal expansion as well as with the elastic properties. The reasons why the Kocevski potential underestimates the UN specific heat are explained by examining the UN phonon properties modeled using both potentials. The Kocevski potential shows better identification of the mechanical stability ranges of UN, UN 2 , and α- and β-U 2 N 3 , reasonably predicting the melting point of UN and predicting stable structures for UN 2 and α- and β-U 2 N 3 . On the other hand, the Tseplyaev potential predicts a premature phase change of both UN and UN 2 and cannot stabilize α- nor β-U 2 N 3 . However, the Kocevski potential cannot predict a stable α-U phase and is thus not suitable for the calculation of formation energies for non-stoichiometric point defects.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

First observation of antiproton annihilation at rest on argon in the LArIAT experiment

We report the first observation and measurement of antiproton annihilation at rest on argon track and shower multiplicities and particle identification conducted with the LArIAT experiment. Stopping antiprotons from the Fermilab Test Beam Facility’s charged particle test beam are identified using beamline instrumentation and LArIAT’s liquid argon time projection chamber (LArTPC). The charged particle multiplicity from the annihilation vertex is manually evaluated via hand scanning, yielding a mean of 3.2 ± 0.4 tracks and a standard deviation of 1.3 tracks, consistent with a semiautomated reconstruction resulting in 2.8 ± 0.4 tracks and a standard deviation of 1.2 tracks. Both methods are consistent with Monte Carlo simulations within statistical uncertainty. The shower multiplicities and particle identification for outgoing tracks are also consistent with eant4 model predictions. These results, obtained from a low-statistics sample, provide a foundation for higher-statistics studies in larger LArTPCs, which could refine modeling of intranuclear annihilation on argon and inform scenarios such as neutron-antineutron oscillations. Published by the American Physical Society 2025

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Extraction of Vibration Data with Imaging

To date, the primary sensing technology used to measure the vibration response has been accelerometers and strain gages mounted directly to the structure and using either wired or, more recently, wireless telemetry. Cost issues with these sensors and the associated data acquisition systems typically limit the numbers that are deployed on in situ structures. Although there are a few structures with larger sensing counts that in some cases exceed over 1000 sensors, more typical numbers range from ten to one hundred sensors resulting in low spatial resolution when they are applied to physically large systems. When one considers that nuclear power plant structures usually have complex geometries, material properties, connectivity and boundary conditions, it is clear these current approaches to vibration measurements can only provide limited information about a system’s dynamics response characteristics. As an alternative, many non-contact measurement technologies have emerged, including point wise measurement methods such as Global Positioning System (GPS), microwave interferometry, and laser Doppler vibrometry (LDV), as well as simultaneous full-field measurement methods such as electronic speckle pattern interferometry, holography interferometry, and muon tomography, some of which can provide high spatial resolution measurements. Among these methods, digital video imaging techniques have emerged as a feasible solution for full-field vibration measurements that provide significantly more detailed dynamic response information because every pixel becomes a measurement point. Furthermore, recent advances in image processing and computer vision algorithms have been successfully used to process video data for experimental and operational modal analysis. Such full-field measurements have the potential to significantly improve many current structural assessment procedures including system identification (modal parameter estimation), structural health monitoring, load reconstruction, model validation, and model updating. Furthermore, more recent full-field imaging techniques can be accomplished with relatively low-cost, commercially-available off-the-shelf cameras. However, these measurement procedures have other limitations that must be considered such as the ability to only measure visibly accessible points on a structure and a more limited dynamic range and bandwidth than can be achieved with accelerometers or strain gages.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Leveraging machine learning to enhance aerosol classification using Single-Particle Mass Spectrometry

Advancing automated classification of atmospheric aerosols from Single-Particle Mass Spectrometry (SPMS) data remains challenging due to overlapping ion signatures, compositional diversity, and limited labeled data. This study evaluates supervised and semi-supervised learning frameworks to enhance aerosol identification by jointly leveraging labeled and unlabeled spectra. Four models were compared: a supervised Support Vector Machine (SVM), a self-training SVM, a stacked autoencoder classifier, and a stacked autoencoder trained using a temporal-ensembling Mean Teacher approach. All models achieved high and stable accuracies (90.0 %–91.1 %), surpassing previous results on the same dataset (87 %) and matching the performance of state-of-the-art deep learning methods. Despite small global metric differences (≤ 1 %), semi-supervised variants yielded up to 5 %–10 % improvements for compositionally rare particle types – such as soot (0.77 % of spectra, F1-score: 0.93–0.97) and hazelnut pollen (0.98 % of spectra, F1-score: 0.97–1.00) – equating to roughly ∼ 187 additional correctly classified spectra. These gains are scientifically significant, as such rare particles exert disproportionate influence on radiative absorption and ice nucleation processes; their improved detection reduces modeled uncertainties in aerosol absorption optical depth and mixed-phase cloud ice nucleation rates. The models' residual misclassifications (≈ 9 %) largely arise from true spectral overlap among chemically adjacent species (e.g., Na- vs. K-feldspar, coated vs. uncoated feldspars), reflecting physical compositional continuity rather than algorithmic error. Collectively, these findings demonstrate that leveraging unlabeled data to learn robust spectral representations and refine classification enhances both fidelity and interpretability, bridging data-driven analysis with aerosol–climate process understanding.

54 ENVIRONMENTAL SCIENCES↗

DOE Repository Metadata Profile (DRMP): A Metadata Framework for Advancing Interoperability and AI Readiness Across Scientific Repositories

The Department of Energy (DOE) funds a diverse and distributed ecosystem of repositories that steward scientific data, publications, and software across its research programs, user facilities, and national laboratories. While significant progress has been made in standardizing dataset-level metadata, the metadata describing repositories themselves (their identity, governance, access interfaces, policies, and technical capabilities) remains inconsistent and fragmented across DOE-funded systems. This variability limits discoverability, interoperability, automated validation, and AI-driven analysis, all of which are increasingly essential for modern scientific workflows. To address this gap, the DOE Data Curation Working Group (DCWG) developed the DOE Repository Metadata Profile (DRMP). The DRMP is a practical, community-driven framework that defines how repositories can describe themselves in a consistent, machine-actionable, and scalable manner. The DRMP is not a new metadata schema. Instead, it is a mapping profile and structured element set capturing the essential characteristics of DOE repositories. It harmonizes repository-level metadata across six widely adopted community schemas: RE3Data; DCAT-US v3; Schema.org; Dublin Core; DataCite 4.6; and PREMIS 3.0. This harmonization eliminates reinvention and enables interoperability within DOE and across the broader scientific ecosystem. A core objective of the DRMP is to reduce burden on repositories by allowing them to reuse their existing metadata through a Rosetta-style crosswalk rather than redesigning local implementations. The profile introduces a three-level conformance model that supports incremental adoption: • Level 1 – Minimum Viable Record (MVR): foundational identification elements required for workflows, project registration, and basic repository presence. • Level 2 – Interoperable: structured metadata enabling alignment with national and international discovery systems. • Level 3 – AI-Ready: enhanced provenance, policy transparency, fixity, semantic context, and capabilities that support automated reasoning, model training governance, and machine-assisted curation. To support implementation, the DRMP includes JSON Schema definitions, OpenAPI patterns, and MCP templates that allow repositories to publish machine-readable metadata directly within existing platforms. These resources are modular and lightweight, enabling adoption without major architectural change. Adopting the DRMP enables repositories to: • Enhance discoverability and interoperability by aligning identifiers, classifications, and descriptive elements across widely used schema standards. • Support federated discovery and cross-registration across DOE systems, Data.gov, and international catalogs. • Enable AI agents and workflow orchestration systems to interpret repository-level metadata within the American Science Cloud (AmSC) through Model Context Protocol (MCP)-based context publication. • Demonstrate alignment with DOE’s open science, stewardship, and FAIR data priorities. This guidance represents a community-driven step forward. Through voluntary adoption and continued feedback, the DRMP advances a cohesive, machine-actionable description of DOE repositories that supports FAIR data practices, preparing the infrastructure for AI-enabled research, and strengthening the discoverability and reuse of DOE’s scientific outputs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Heterogeneous Multi-Domain Dataset Synthesis to Facilitate Privacy and Risk Assessments in Smart City IoT

The emergence of the Smart Cities paradigm and the rapid expansion and integration of Internet of Things (IoT) technologies within this context have created unprecedented opportunities for high-resolution behavioral analytics, urban optimization, and context-aware services. However, this same proliferation intensifies privacy risks, particularly those arising from cross-modal data linkage across heterogeneous sensing platforms. To address these challenges, this paper introduces a comprehensive, statistically grounded framework for generating synthetic, multimodal IoT datasets tailored to Smart City research. The framework produces behaviorally plausible synthetic data suitable for preliminary privacy risk assessment and as a benchmark for future re-identification studies, as well as for evaluating algorithms in mobility modeling, urban informatics, and privacy-enhancing technologies. As part of our approach, we formalize probabilistic methods for synthesizing three heterogeneous and operationally relevant data streams—cellular mobility traces, payment terminal transaction logs, and Smart Retail nutrition records—capturing the behaviors of a large number of synthetically generated urban residents over a 12-week period. The framework integrates spatially explicit merchant selection using K-Dimensional (KD)-tree nearest-neighbor algorithms, temporally correlated anchor-based mobility simulation reflective of daily urban rhythms, and dietary-constraint filtering to preserve ecological validity in consumption patterns. In total, the system generates approximately 116 million mobility pings, 5.4 million transactions, and 1.9 million itemized purchases, yielding a reproducible benchmark for evaluating multimodal analytics, privacy-preserving computation, and secure IoT data-sharing protocols. To show the validity of this dataset, the underlying distributions of these residents were successfully validated against reported distributions in published research. We present preliminary uniqueness and cross-modal linkage indicators; comprehensive re-identification benchmarking against specific attack algorithms is planned as future work. This framework can be easily adapted to various scenarios of interest in Smart Cities and other IoT applications. By aligning methodological rigor with the operational needs of Smart City ecosystems, this work fills critical gaps in synthetic data generation for privacy-sensitive domains, including intelligent transportation systems, urban health informatics, and next-generation digital commerce infrastructures.

IoT↗

Accurate data-driven surrogates of dynamical systems for forward propagation of uncertainty

Stochastic collocation (SC) is a well-known non-intrusive method of constructing surrogate models for uncertainty quantification. In dynamical systems, SC is especially suited for full-field uncertainty propagation that characterizes the distributions of the high-dimensional solution fields of a model with stochastic input parameters. However, due to the highly nonlinear nature of the parameter-to-solution map in even the simplest dynamical systems, the constructed SC surrogates are often inaccurate. Here, this work presents an alternative approach, where we apply the SC approximation over the dynamics of the model, rather than the solution. By combining the data-driven sparse identification of nonlinear dynamics framework with SC, we construct dynamics surrogates and integrate them through time to construct the surrogate solutions. We demonstrate that the SC-over-dynamics framework leads to smaller errors, both in terms of the approximated system trajectories as well as the model state distributions, when compared against full-field SC applied to the solutions directly. We present numerical evidence of this improvement using three test problems: a chaotic ordinary differential equation, and two partial differential equations from solid mechanics.

42 ENGINEERING↗

Search for Dark Matter Produced in Association with a Dark Higgs Boson in the b b ¯ Final State Using p p Collisions at s = 13 TeV with the ATLAS Detector

A search is performed for dark matter particles produced in association with a resonantly produced pair of b -quarks with 30 < m b b < 150 GeV using 140 fb − 1 of proton-proton collisions at a center-of-mass energy of 13 TeV recorded by the ATLAS detector at the LHC. This signature is expected in extensions of the standard model predicting the production of dark matter particles, in particular those containing a dark Higgs boson s that decays into b b ¯ . The highly boosted s → b b ¯ topology is reconstructed using jet reclustering and a new identification algorithm. This search places stringent constraints across regions of the dark Higgs model parameter space that satisfy the observed relic density, excluding dark Higgs bosons with masses between 30 and 150 GeV in benchmark scenarios with Z ′ mediator masses up to 4.8 TeV at 95% confidence level. © 2025 CERN, for the ATLAS Collaboration 2025 CERN

Aad, G. (ORCID:0000000266654934)↗

Brittle failure analysis and modeling of high-burnup PWR fuel cladding alloys

The aim of this research is the development of methods for predicting mechanical behavior and identification of limiting conditions to prevent brittle failure of high-burnup (HBU) pressure water reactor (PWR) fuel cladding alloys. A finite element (FE) model of the ring compression test (RCT) was created to analyze the failure behavior of zirconium-based alloys with radial hydrides during the RCT. An elastic-plastic material model describes the zirconium alloy. The stress-strain curve needed for the elastic-plastic material model was derived by inverse finite element analyses. Cohesive zone modeling is used to reproduce sudden load drops during RCT loading. Based on the failure mechanism in non-irradiated ZIRLO (R) claddings, a micro-mechanical model was developed that distinguishes between brittle failure along hydrides and ductile failure of the zirconium matrix. Two different cohesive laws representing these types of failure are present in the same cohesive interface. The key differences between these constitutive laws are the cohesive strength, the stress at which damage initiates, and the cohesive energy, which is the damage energy dissipated by the cohesive zone. Statistically generated matrix-hydride distributions were mapped onto the cohesive elements and simulations with focus on the first load drop were performed. Computational results are in good agreement with the RCT results conducted on high-burnup M5 (R) samples. It could be shown that crack initiation and propagation strongly depend on the specific configuration of hydrides and matrix material in the fracture area.

Simbruner, Kai↗

Molecular Vision - Multimodal, multitask retrieval of molecular structure from measured signatures for reference-free compound identification

We are currently at risk of generating false conclusions based on limited methods to identify small molecules in biological systems and in chemical forensics. By definition, the chemical structures of novel small molecules have not been determined, let alone measured or synthesized. Currently, unambiguous structure determination of small molecules is constrained by the time and effort needed to isolate compounds and perform de novo structure elucidation using laboratory-based methods, significantly extending the time to inform mitigation strategies. To address this gap, we have developed a deep learning approach to directly map molecular structure to experimental signatures. We aim to unify measurement technologies employed in untargeted small molecule identification studies—such as infrared (IR) spectrometry, tandem mass spectrometry (MS/MS), ion mobility spectrometry-derived collision cross section (CCS)—through use of a multimodal, multitask deep learning architecture. Where existing methods require direct generation of information-rich spectra and/or properties, an inherently difficult task, we will simplify molecular signature-based identification by posing the problem as a recognition or retrieval task. The model is thus presented with relevant endpoints – structure and one or more molecular signatures – and need only determine whether they are semantically related. Thus, our approach offers the following advantages over existing techniques: (i) circumvents difficulties associated with direct generation of molecular signatures from structure and structure from signatures; (ii) incorporates multiple molecular signatures simultaneously, as available, to support identification; and (iii) enables rapid computation of structural embeddings toward broad coverage of known chemical space. Taken together, the approach removes the need to explicitly obtain or compute reference spectra, representing a powerful method for compound identification that requires only experimentally observed signatures.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗