Search NASASearch

SEARCH · Search NASA

Results for “Databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

The Natural Products Magnetic Resonance Database (NP-MRD) for 2025

The Natural Products Magnetic Resonance Database or NP-MRD (https://np-mrd.org) is a comprehensive, freely accessible, web-based resource for the deposition, distribution, extraction and retrieval of nuclear magnetic resonance (NMR) data on natural products. The NP-MRD was initially established to support compound de-replication and data dissemination for the natural products community. However, that community has now grown to include many users from the metabolomics, microbiomics, foodomics and nutrition science fields. Indeed, since its launch in 2021, the NP-MRD has expanded enormously in size, scope and popularity. The current version of NP-MRD now contains nearly 7X more compounds (281,859 vs. 40,908) and 7X more NMR spectra (5.1 million vs. 817,000) than the first release. More specifically, an additional 4.6 million predicted spectra and another 11,000 spectra simulated from experimental chemical shifts were deposited into the database. Likewise, the number of NMR raw spectral data depositions has grown from a 165 spectra per year to more than 10,000 per year. As a result of this expansion, the number of monthly webpage views has grown from 55 to 20,000 and the number of monthly visitors has increased from 7 to 2500. To address this growth and to better support the expanding needs of its diverse community of users, many additional improvements to the NP-MRD have been made. These include significant enhancements to the data submission process, important improvements to the visualization and display of NMR spectra, notable updates to the database’s spectral search utilities and useful additions to support better NMR spectral analysis/prediction. Significant efforts have also been undertaken to remediate and update many of NP-MRD’s database entries. This manuscript describes these database improvements and expansion efforts, along with how they have been implemented and what future upgrades to the NP-MRD are planned.

Artifical Intelligence

G2Aero Database of Airfoils - Curated Airfoils

This dataset contains a curated set of 19,164 airfoil shapes from various applications and the data-driven design space of separable shape tensors (PGA space), which can be used as a parameter space for machine-learning applications focused on airfoil shapes. We constructed the airfoil dataset in two main stages. First, we identified 13 baseline airfoils from the NREL 5MW and IEA 15MW reference wind turbines. We reparameterized these shapes using least-squares fits of 8-order CST parametrizations, which involve 18 coefficients. By uniformly perturbing all 18 CST coefficients by +/-20% around each baseline airfoil, we generated 1,000 unique airfoils. Each airfoil was sampled with 1,001 shape landmarks whose x-coordinates followed a cosine distribution along the chord. This process resulted in a total of 13,000 airfoil shapes, each with 1,001 landmarks. In the second phase, we gathered additional airfoils from the extensive BigFoil database, which consolidates data from sources such as the University of Illinois Urbana-Champaign (UIUC) airfoil database, the JavaFoil database, the NACA-TR-824 database, and others. We undertook a thorough pre-processing step to filter out shapes with sparse, noisy, or incomplete data. We also removed airfoils with sharp leading edge and those exceeding our threshold for trailing edge thickness. Additionally, we thinned out the collection of NACA airfoils-- parametric sweeps of NACA airfoils with increasing thickness and camber present in BigFoil database-- by selecting every fourth step in the parameter sweeps. Finally, we regularized the airfoils by reparametrizing them with an 8-order CST parametrization (with 1,001 shape landmarks with x coordinated following cosine distribution along the chord) and removing airfoils with high reconstruction errors. This data pre-processing resulted in a set of 6,164 airfoils. In total, our curated airfoil dataset comprises 19,164 airfoils, each with 1,001 landmarks, and is stored in the curated_airfoils.npz file. Using this curated airfoil dataset, we utilized the separable shape tensors framework to develop a data-driven parameterization of airfoils based on principal geodesic analysis (PGA) of separable shape tensors. This PGA space is provided in PGAspace.npz file.

airfoils

An Overview of the Molten Salt Thermal Properties Database--Thermophysical, Version 3.1 (MSTDB-TP v.3.1)

This report presents the current status of the Molten Salt Thermal Properties Database–Thermophysical (MSTDB-TP). Information regarding version 3.1 is provided herein, which contains 820 individual salt entries (data from 180+ independent studies); the thermophysical properties contained in the database include density, viscosity, thermal conductivity, and heat capacity. The major updates to the database include a significant expansion of pseudobinary and higher-order chloride salt mixtures, many of which bearing actinides, and an incorporation of more recent literature data (i.e., that within the past 5 years). Also, modifications have been made to the pure compound data in the database as a consequence of an external quality assessment of duplicate datasets. The user-facing API for the MSTDB-TP, Saline, has been updated to include viscosity estimation capabilities based on the Redlich–Kister formalism; this is an advancement with respect to the existing density estimation capabilities. The graphical user interface was also updated to include a density estimation capability, backed by Saline. Finally, additional preliminary efforts to include surface tension into the database, as well as an investigation on formalisms that would be appropriate for thermal conductivity estimation, are reported herein.

36 MATERIALS SCIENCE

BRE‐X Emissions Database for End‐of‐Life Scenarios of Selective Building Construction Materials to Enable Circular Economy in Construction

In the United States, construction and demolition debris predominately end up in landfills with minimal end‐of‐life Re‐X (recover, recycle, reuse, etc.) scenarios, resulting in large environmental impacts and lost opportunities for material recovery. Except for concrete and metals, which seem to have a few well‐defined end‐of‐life pathways, there seems to be a lack of well‐documented end‐of‐life scenarios for other construction materials, let alone their emissions data. Hence, there is a need for documented end‐of‐life Re‐X scenarios and end‐of‐life data of more building materials to motivate widespread use of Re‐X strategies in building design. This paper outlines the efforts of the National Renewable Energy Laboratory, Carbon Leadership Forum, Building Transparency, and Skidmore, Owings & Merrill to (a) create an open‐access BRE‐X (Building Re‐X) end‐of‐life emissions database consisting of greenhouse gas emissions data associated with various end‐of‐life scenarios for a select list of high‐impact building construction materials, and (b) integrate the BRE‐X end‐of‐life emissions database with CAD/BIM/LCA tools for evaluating various end‐of‐life scenarios. The paper also presents a few existing life cycle inventory databases that contain sparse amounts of end‐of‐life data for a few construction materials and their limitations in terms of scaling and data consolidation. Finally, a sample of how the collected data can be ingested into whole‐building LCA tools using open data formats and a public access link to the BRE‐X end‐of‐life emissions database is also included.

36 MATERIALS SCIENCE

Large language model-driven database for thermoelectric materials

Thermoelectric materials have the ability to convert waste heat into electricity, offering a valuable solution for energy harvesting. However, their widespread use is hindered by low conversion efficiency, the reliance on expensive rare earth elements, and the environmental and regulatory concerns associated with lead-based materials. A fast and cost-effective way to identify highly efficient thermoelectric materials is through data-driven methods. These approaches rely on robust and comprehensive datasets to train models. Although there are several databases on thermoelectric materials, there is still a need to collect and integrate experimental data from peer-reviewed research articles to capture diverse compositions and properties of materials. Here, in this work, we developed a comprehensive database of 7,123 thermoelectric compounds, containing key information such as chemical composition, structural detail, seebeck coefficient, electrical and thermal conductivity, power factor, and figure of merit (ZT). We used the GPTArticleExtractor workflow, powered by large language models (LLM), to extract and curate data automatically from the scientific literature published in Elsevier journals. This process enabled the creation of a structured database that addresses the challenges of manual data collection. The open access database could stimulate data-driven research and advance thermoelectric material analysis and discovery.

Database

Data from TropiRoot 1.0 database: tropical root characteristics across environments

TropiRoot 1.0 is a new tropical root database with root characteristics across environment gradients. It has data extracted from 104 new sources, resulting in more than 8000 rows of data (either species or community data). Most of the data in TropiRoot 1.0 includes root characteristics such as root biomass, morphology, root dynamics, mass fraction, architecture, anatomy, physiology and root chemistry. This initiative represents an approximately 30% increase in the currently available data for tropical roots in the Fine Root Ecology Database (FRED). TropiRoot 1.0, contains root characteristics from 25 different countries where seven are located in Asia, six in South America, five in Central America and the Caribbean, four in Africa, two in North America, and 1 in Oceania. Due to the volume of data, when ancillary data was available, including soil data, these data was either extracted and included in the database or their availability was recorded in an additional column. Multiple contributors checked the entries for outliers during the collation process to ensure data quality. For text-based observations, we examined all cells to ensure that their content relates to their specific categories. For numerical observations, we ordered each numerical value from least to greatest and plotted the values, checking apparent outliers against the data in their respective sources and correcting or removing incorrect or impossible values. Some data (soil and aboveground) have different columns for the same variable presented in different units, including originally published units, but root characteristics data had units converted to match the ones reported in FRED. By filling a gap from global databases, TropiRoot 1.0 expands our knowledge of otherwise so far underrepresented regions, and our ability to assess global trends. This advancement can be used to improve tropical forest representation in vegetation models.

54 ENVIRONMENTAL SCIENCES

The northeast materials database for magnetic materials

The discovery of magnetic materials with high operating temperature ranges and optimized performance is essential for advanced applications. Current data-driven approaches are limited by the lack of accurate, comprehensive, and feature-rich databases. This study aims to address this challenge by using Large Language Models (LLMs) to create a comprehensive, experiment-based, magnetic materials database named the Northeast Materials Database (NEMAD), which consists of 67,573 magnetic materials entries (www.nemad.org). The database incorporates chemical composition, magnetic phase transition temperatures, structural details, and magnetic properties. Enabled by NEMAD, we trained machine learning models to classify materials and predict transition temperatures. Our classification model achieved an accuracy of 90% in categorizing materials as ferromagnetic (FM), antiferromagnetic (AFM), and non-magnetic (NM). The regression models predict Curie (Néel) temperature with a coefficient of determination (R 2 ) of 0.87 (0.83) and a mean absolute error (MAE) of 56K (38K). These models identified 25 (13) FM (AFM) candidates with a predicted Curie (Néel) temperature above 500K (100K) from the Materials Project. This work shows the feasibility of combining LLMs for automated data extraction and machine learning models to accelerate the discovery of magnetic materials.

Ferromagnetism

A database and meta-analysis on the performance of exploding pusher implosions conducted at OMEGA

A database of 222 exploding pusher implosions conducted at the OMEGA Laser Facility is presented. The dataset consists of glass-shell capsules filled with varying pressures of D 2 , T 2 , and 3 He, which were imploded using square laser pulses with intensities ranging from 1 to 1 × 10 15 W/cm 2 . The database includes measurements of bang times, ion temperatures, and yields from the DD, D 3 He, and DT fusion reactions. A semi-analytic exploding pusher model is introduced, which effectively captures the observed trends in the data. This model predicts that the measurements scale according to a power-law relation based on the initial capsule and laser conditions. A generalized power-law scaling relation is directly fit to each dataset, providing a useful interpolation of the entire database. Overall, the database provides a valuable resource to estimating bang times, temperatures, and yields for the design of future experiments. Additionally, it provides a diverse set of data for validating more advanced implosion physics models.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Empirical scaling of the L–H threshold power for metal wall tokamaks using a multi-device database

The empirical scaling for the H-mode power threshold in tokamaks has been revisited using a database with threshold data from machines with a metallic first wall as part of International Tokamak Physics Activity (ITPA) task TC-26. The database contains discharges from ASDEX Upgrade (AUG) (W), JET (Be/W) and Alcator C-Mod (Mo). This was motivated by reports that in like-for-like discharges the power threshold was reduced by approximately 30% after the change from carbon based to metallic first wall materials on AUG (Ryter et al 2013 Nucl. Fusion 53 113003) and JET (Maggi et al 2014 Nucl. Fusion 54 023007). The database contains L–H transition data for all hydrogen isotopes and mixtures, including T and DT from the recent JET campaigns. Compared to the ITPA 2008 scaling (Martin et al 2008 J. Phys.: Conf. Ser. 123 012033), the metal wall scaling has a smaller magnetic field exponent but a larger density exponent. We present an additional parameter to capture the strong dependence of the L–H power threshold (approx. factor 2) on the magnetic configuration in the divertor on JET. The scaling recovers the approximate inverse isotope mass scaling of the threshold power. Alternative scalings involving the plasma current and poloidal magnetic field are explored. Despite the reduction in threshold observed earlier, the scalings based on the metal wall database do not necessarily extrapolate to a lower threshold for ITER compared to the ITPA 2008 scaling, especially at high density. The divertor configuration effect induces the largest uncertainty in the extrapolation.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

VISTA Enhancer browser: an updated database of tissue-specific developmental enhancers

Regulatory elements (enhancers) are major drivers of gene expression in mammals and harbor many genetic variants associated with human diseases. Here, we present an updated VISTA Enhancer Browser (https://enhancer.lbl.gov), a database of transgenic enhancer assays conducted in developing mouse embryos in vivo. Since the original publication in 2007, the database grew nearly 20-fold from 250 to over 4500 experiments and currently harbors over 23 500 images. The updated database provides structured information on experiments conducted at different stages of embryonic development, including enhancer activities of human pathogenic and synthetic variants and sequences derived from a variety of species. In addition to manually curated results of thousands of individual experiments, the new database also features hundreds of manually curated comparisons between alleles. The VISTA Enhancer Browser provides a crucial resource for study of human genetic variation, gene regulation and developmental biology.

59 BASIC BIOLOGICAL SCIENCES

Enzyme Engineering Database (EnzEngDB): a platform for sharing and interpreting sequence–function relationships across protein engineering campaigns

The discovery and engineering of new enzymes is important across the bioeconomy, with diverse applications from foods to pharmaceuticals, sensors to agriculture. However, enzyme engineering, in particular machine learning-guided engineering, is hampered by a lack of data. Currently there exists no database designed to capture and interpret datasets created in this domain, nor are there easy analysis and visualisation tools. We developed the Enzyme Engineering Database to provide a centralized resource and an online analysis tool to consolidate sequence-function data from enzyme engineering campaigns, thereby making three contributions: (i) a database into which researchers can deposit public data, (ii) visualisation and analysis tools for protein engineers to analyse their own data or compare enzyme variants to other engineering campaigns, and (iii) a gold-standard dataset for benchmarking automated extraction along with the first large language model extraction pipeline specific for enzyme engineering campaigns. The Enzyme Engineering Database is accessible at http://enzengdb.org/.

Long, Yueming [California Institute of Technology

The Pan-Arctic Vegetation Cover (PAVC) database v1.1

The Pan-Arctic Vegetation Cover (PAVC) database contains synthesized field-data observations of vegetation cover from 978 Arctic Alaska plots with observations from 2010 to 2021. The cover datasets contain plot data at both the plant functional type (PFT) and species-level resolution, with standardized PFT definitions and species names. We synthesized publicly available point-intercept and visual estimate plots from the Arctic Vegetation Archive of Alaska, the Alaska Vegetation Plots Database, the North Slope Science Catalog, and the National Ecological Observatory Network; as well as previously unpublished data from the Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic).Users will find four synthesized datasets, 4 associated data descriptor (dd) files, and 1 metadata file in the PAVC database:synthesized_species_fcover.csv contains fractional cover (fcover) for unique accepted species names, where names include vegetation identified at the family, genus, species, subspecies, and variety levels, as well as general functional types across all 5 data sources. The synthesized_species_fcover_dd.csv accompanies this dataset with header information.synthesized_pft_fcover.csv contains fcover for the following PFTs: non-vascular plants with lichen and bryophyte subcategories, trees with deciduous and evergreen subcategories, shrubs with deciduous and evergreen subcategories, graminoids (grasses), and forbs (herbaceous flowering plants) measured as total cover. Litter and “other” cover are also included as total cover. Additional “types” include water and bare ground, which were measured as top cover. The synthesized_pft_fcover_dd.csv accompanies this dataset with header information.species_pft_checklist.csv is a lookup table containing the translation from a dataset species name to an accepted species name and to a PFT. This table can be used to clarify our species to PFT adjudications, and to aid users in assigning their own PFTs. Any issues found in this checklist should be reported in the Issues tab of our github.survey_unit_information.csv contains auxiliary information about the plots synthesized in this database. It contains useful information for filtering plots of interest based on temporal, geospatial, and contextual information about the plot surveys.flmd.csv contains metadata information about each file in the database.This research was performed as a part of the NGEE Arctic project. The NGEE Arctic project was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research.The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska.Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES

Prospective Seal Unit Spatial Extent Database for U.S. Sedimentary Basins

The Prospective Seal Unit Spatial Extent Database for U.S. Sedimentary Basins contains a series of spatial datasets representing spatial extents of publicly available data for caprock and seal rock units within the Appalachian Basin, Denver-Julesburg Basin, Great Valley Basin (Sacramento and San Joaquin Basins), Illinois Basin, Michigan Basin, San Juan Basin, U.S. Gulf Coast Basin, and Williston Basin. The database is designed to support carbon storage feasibility and resources assessment for carbon transport and storage (CTS) projects while displaying the spatial extent of prospective seal units and provide a guide to the original data source. This database leverages publicly available data resources from authoritative sources (e.g. U.S. Geological Survey, State Geologic Surveys, and published reports), and aims to help guide users to understand the seal unit's spatial coverage and data gaps from the regional to sub-basin/field scale. The database is organized by seal unit/formation, including the spatial extent for data found to be available for the seal unit. The various datasets represented include spatial extents of the lithologic formation, depth to top structural contour maps, and thickness/isopach maps. Included in this submission are the following resources: 1. Geodatabase/Dataset: “prospective-seal-unit-extents-2025.gdb” 2. ReadMe: “readme-prospective-seal-unit-spatial-extent-dataset-2025.pdf” 3. Data Catalog: “prospective-seal-unit-spatial-extents-data-catalog-2025.xlsx” 4. Data Sources Key: “data-source.csv” Please see NETL disclaimers here: https://netl.doe.gov/home/disclaimer

Basin

M3SF-24LL010301052-Summary of SUPCRTNE testing and Rev0 database

This progress report (Level 3 Milestone Number M3SF-24LL010301052) summarizes research conducted at Lawrence Livermore National Laboratory (LLNL) within the Argillite Disposal R&D work package SF-24LL01030105. SUPCRTNE is being developed as the primary engine for thermodynamic database development in support of geologic disposal of high-level nuclear waste. As used here, “SUPCRT” refers to both a computer program and its supporting database. Additional letters or numbers refer to specific variants (SUPCRT92, (Johnson et al., 1992); SUPCRTBL, (Zimmer et al., 2016); and SUPCRTNE). A SUPCRT database contains the data required to calculate the thermodynamic properties of solid, gas, and aqueous species over a wide range of temperature and pressure. Normally, SUPCRT is used to create a higher-level database that directly supports modeling and simulation codes such as EQ3/6, GWB, PHREEQC, and PFLOTRAN.

58 GEOSCIENCES

Data Qualification Report: SRNL Glass Composition-Properties (ComPro) Database

The Savannah River National Laboratory Glass Composition-Properties (ComPro) database is an extensive database containing pertinent composition and durability data to support the accelerated clean-up mission at the Defense Waste Processing Facility. The activities described in this data qualification report were performed to support the information contained in the database. There were two objectives of the original data qualification process. The first objective was to review supporting documentation to determine if DOE/RW-0333P Quality Assurance Requirements and Description had been implemented during the original work. If the DOE/RW-0333P Quality Assurance Requirements and Description had not been directly implemented during the original work, the second objective was to determine if the controls that were used were adequate to meet the intent of the DOE/RW-0333P Quality Assurance Requirements and Description. The results of these two objectives and the activities performed to support these decisions are described in this document. An assessment of each dataset was made to determine if the data were RW-0333P Compliant, RW-0333P Equivalent or Non-RW-0333P Compliant. The original data qualification was performed in accordance with E7, Conduct of Engineering Manual, Procedure 3.70, Revision 4, Qualification of Data. The specific method that was used was Equivalent Controls as described in E7, 3.70. Revision 2 of this document adds supporting information for the RW-0333P Compliant datasets added to Revision 3 of the database.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W

L-SCIE and SUPCRTNE model and database development

This progress report (Level 3 Milestone Number M3SF-25LL010301052) summarizes research conducted at Lawrence Livermore National Laboratory (LLNL) within the Argillite Host Rock Properties & Processes Work Package SF-25LL01030105. The focus of this milestone is to create an initial SUPCRTNE database and extensively test the SUPCRTNE code. We prepared a draft manuscript describing our workflow and approach for developing a next-generation thermodynamic database for the Spent Fuel and High Level Waste Disposition campaign. The goal is to ease any future formal software qualification effort as was done on the Yucca Mountain Project for codes including SUPCRT92 and EQ3/6. We are attempting to extend the SUPCRTNE database to include more species, mainly from the NEA volumes. We can also take advantage of thermodynamic database activities that are ongoing at Thermochimie and as part of the EURADII program. However, our effort is limited by the funds available and the changes to program scope that was initiated in mid-FY25.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W

An Overview of the Molten Salt Thermal Properties Database–Thermophysical, Version 4.0 (MSTDB-TP V.4.0)

A central repository of thermophysical and thermochemical properties of molten salt compositions of relevance to molten salt reactors (MSRs) is vital in supporting the broad community of MSR developers, who are at various stages of developing and deploying their reactor designs. In general, these MSR designs differ significantly from developer to developer (e.g., with respect to the hardness of the neutron spectra, level of fissile loading, target multicomponent temperatures and power levels, and moderating capabilities). Therefore, the fuel and coolant salts being considered vary greatly: they may be chlorides or fluorides, they utilize different actinides at different ratios, and the cations in the melt are selected based on perceived advantages and disadvantages. Considering the general need for thermal properties, and the vastness of the array of potential candidate salt mixtures, the Molten Salt Thermal Properties Database (MSTDB) was initiated in 2018 with the goal of providing thermophysical and thermochemical characterization of key molten salt compounds and mixtures across their temperature and compositional domains. The MSTDB is thus divided into the thermophysical arm (MSTDB-TP) and the thermochemical arm (MSTDB-TC). The MSTDB is an effort funded by the Department of Energy, Office of Nuclear Energy (DOE-NE) Nuclear Energy Advanced Modeling and Simulation (NEAMS) program, and the MSR Campaign. This report provides an overview of the MSTDB-TP v4.0 in terms of the data contained within, the state of the tools used to access the data, the availability of predictive models that leverage the raw data in the database, the preliminary status of developmental efforts that are currently underway, and an account of future goals for MSTDB-TP. The primary goal for the update from MSTDB-TP v.3.1 to v4.0 was the incorporation of surface tension data into the database; this property is important for thermal hydraulics modeling and species transport in other tools that have been developed under the NEAMS program. A breakdown of the surface tension data that have been added into MSTDB-TP v4.0 is provided herein, and the manner in which the quality of the data has been assessed is also documented. For MSTDB-TP v4.0, newly published thermophysical property data—primarily from collaborative experimental efforts under the MSR Campaign—have been incorporated into the database, and the resulting expansion is documented here. Because of the size to which MSTDB-TP has grown, the raw data format has now been recast into JavaScript Object Notation (JSON) format for easier connection with the MSTDB-TP application programming interface (API). Saline; the pre-existing comma-separated value (CSV) format has been deprecated but is still maintained, accessible, and up to date. As a final effort in packaging the MSTDB-TP v4.0 update, the graphical user interface (GUI) for MSTDB has been updated to allow full accessibility to the density and viscosity predictive models, which are based on Redlich-Kister expansions of MSTDB-TP raw data. Some other major aspects of this report, in terms of preliminary and future work, include: (1) documentation of the formalism and preliminary testing of a kinetic theory model that may act as a predictive model for thermal conductivity; (2) documentation of the candidate predictive models that may be considered in the future for surface tension, making use of the surface tension data now in MSTDB-TP v4.0; (3) a preliminary account of a data collection process that will enable the filling of additional gaps within MSTDB-TP, namely with data which have been collected computationally (e.g., through ab initio molecular dynamics).

22 GENERAL STUDIES OF NUCLEAR REACTORS

UCB-GLOBES: An open-access mass spectral database of identified and unidentified atmospheric organic compounds

Chemical characterization of atmospheric organic aerosols using gas chromatography with 70 eV electron ionization mass spectrometry (GC/EI-MS) has been used for decades in advancing molecular marker detection and identification, though primarily through suspect screening and/or targeted analyses. To advance non-targeted analyses of environmental samples, we have catalogued approximately 27 000 mass spectra (MS) of the trimethylsilyl derivatives of semi-volatile organic aerosol (OA) analytes in the open-access University of California Berkeley Goldstein Library of Organic Biogenic Environmental Spectra (UCB-GLOBES). Analytes were observed in ambient samples from the U.S. and the Central Amazon and/or laboratory simulations of secondary OA (SOA) formation. These samples are representative of OA under urban and biomass burning influences as well as SOA derived from biogenic precursors (e.g., isoprene, monoterpenes, sesquiterpenes) and biomass burning intermediates. MS are documented in UCB-GLOBES without regard to known chemical identity, annotated with extensive metadata such as sample source/experimental conditions, any structural information gained from MS analyses, and predicted chemical properties such as average carbon oxidation state and carbon number. UCB-GLOBES MS are compatible for importing into the NIST MS Search program, and we have also provided a Jupyter Notebook for MS visualization and comparisons. We demonstrate the utility of UCB-GLOBES through MS reanalyses of prior analytes observed in ambient data, finding a 20 % reduction in the number of analytes assigned to OA source categories reliant solely on time series correlation and an overall 11 % increase in new MS-based OA source categorization for the Southeast U.S. For 1513 analytes observed previously in the Central Amazon, we found 375 MS matches using UCB-GLOBES vs. 136 MS matches during prior analyses, representing a 14 % gain in newly confirmed or newly categorized OA species. While OA from laboratory oxidation experiments in UCB-GLOBES are highly diverse chemically, on average only 29 % of UCB-GLOBES MS have a mass spectral match to another MS entry in UCB-GLOBES and/or in databases of known compounds (i.e. NIST MS Database, Adams Essential Oil, MANE Flavor and Fragrance Company). This indicates that roughly 70 % of UCB-GLOBES MS are unique thus far, not observed more than once among the laboratory oxidation samples and ambient data in UCB-GLOBES MS. Further, only 18 % can be positively identified using these databases or known authentic standards. This points to a large gap between these laboratory simulations and ambient OA. Overall, the UCB-GLOBES database can be utilized for improving confidence in OA source categorization and/or identification, novel chemical marker discovery, tracking chemical diversity, de novo structure and properties prediction, and improving MS search and matching algorithms. This can ultimately inform future research priorities for the chemical characterization of atmospheric organic samples.

Mass spectrometry