Search NASA⌕ Search

Engineering topics

Yoo, Pilsun

Publications and source records attributed to Yoo, Pilsun.

Molecular origin of viscoelasticity and influence of methylation in mesophase pitch

The viscoelastic and thermomechanical properties of pitches are responsible for their melt-spinning behavior, which is a critical step for manufacturing high-performance pitch-based carbon fibers. Here, we systematically explore the impact of methyl group modifications on the viscoelastic and thermal properties of mesophase pitches. We employ a range of atomistic modeling approaches, including Density Functional Theory (DFT), Density Functional Tight Binding (DFTB), and Classical Molecular Mechanics (MM), to provide detailed insights into the molecular interactions and structural changes. Further, our results revealed the molecular mechanisms that promote layered structures leading to the anisotropic nature of the viscoelastic behavior of mesophase pitch. Furthermore, we propose a modified molecular representation of naphthalene-based mesophase pitch based on the analysis of x-ray diffraction measurements. This study provides fundamental insights into the molecular structures of mesophase pitch and the role of methyl groups controlling its viscosity, which offer valuable insights into mesophase-based carbon fiber production.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Scaling Ensembles of Data-Intensive Quantum Chemical Calculations for Millions of Molecules

Deep learning models are efficient computational tools that can accelerate the inverse design of molecules with desired functional properties by generating predictions at a fraction of the time required by traditional quantum chemical approaches. To ensure that a model maintains accuracy and transferability across broad regions of the chemical space explored during the inverse design, it must be trained on massively large volumes of simulation data. This requires running large-scale ensemble quantum chemical calculations on high-performance computing (HPC) systems for data collection. However, the efficient execution of such large ensemble calculations and the management of large volumes of output data require tools that can judiciously utilize computational resources and manage metadata overhead on the file system. Therefore, we present a high-performance, scalable, ensemble management framework for performing data-intensive quantum chemical electronic structure calculations for organic molecules. This framework provides abstractions to plug different ab initio, first principles, and first principles-based semi-empirical methods and executes them efficiently at large scale on HPC systems. It dynamically distributes tasks to resources and uses tiered storage for managing large collections of files. We employed this framework to process over ten million organic molecules and generate open-source datasets that provide UV-vis absorption spectra by running time-dependent density-functional tight-binding calculations. It is the largest database containing molecular optical spectra that were simulated with quantum chemical methods in a consistent manner.

Mehta, Kshitij↗

GDB-9-Ex_TD-DFT-PBE0: Dataset containing Time Dependent Density Functional Theory (TDDFT) calculations for organic molecules of the GDB-9-Ex dataset.

This dataset contains data-intensive quantum chemical electronic structure calculations for 96,766 organic molecules of the GDB-9-Ex dataset. Calculations were performed using the Time Dependent Density Functional Theory (TDDFT) first principles method using the ORCA software. It provides UV-vis spectra calculations of molecules with a high level of accuracy. The optical spectra behavior was collected based on the optimized molecular geometries in the DFTB method with 3ob parameters. All calculations utilized the def2-TZVP basis sets with the auxiliary def2/J and def2-TZVP/C basis sets. The time-dependent density-functional theory (TDDFT) approach with the PBE0 exchange-correlation functional and ORCAs default integration grid was employed. For the excitation energy calculations, the lowest 50 excitation states were calculated.

AI dataset↗

GDB-9-Ex_EOM-CCSD: Dataset containing Equation of Motion Coupled Cluster (EOM-CCSD) calculations for organic molecules of the GDB-9-Ex dataset.

This dataset contains data-intensive quantum chemical electronic structure calculations for 80,593 organic molecules of the GDB-9-Ex dataset. Calculations were performed using the Equation of Motion Coupled Cluster (EOM-CCSD) first principles method using the ORCA software. It provides UV-vis spectra calculations of molecules with a high level of accuracy. The optical spectra behavior was collected based on the optimized molecular geometries in the DFTB method with 3ob parameters. All calculations utilized the def2-TZVP basis sets with the auxiliary def2/J and def2-TZVP/C basis sets. The similarity-transformed EOM-CCSD method that used domain-based local pair natural orbitals (DLPNO) approximation which constitutes the STEOM-DLPNO-CCSD method was used. This method is based on the STEOM approach and was found to make accurate predictions of transition energies for organic molecules. For the excitation energy calculations, the lowest 50 excitation states were calculated.

AI dataset↗

Large-scale atomistic model construction of subbituminous and bituminous coals for solvent extraction simulations with reactive molecular dynamics

Large-scale atomistic models for complex polycyclic aromatic hydrocarbon systems help understand the chemical properties and behaviors of complex feedstocks such as coal or petroleum. However, the development and utilization of large-scale models remain limited due to the difficulty in achieving the varied structural characteristics necessary to capture stochastic nature of these feedstocks. Here we demonstrate a systematic workflow to construct stochastic molecular systems from a broad analytical suite: high-resolution transmission electron microscopy (HRTEM), carbon-13 nuclear magnetic resonance spectroscopy ( 13 C NMR), laser desorption ionization mass spectroscopy (LDI-MS), and elemental analysis. We present a model construction and analysis utility of a new Python-based module. We selected one subbituminous and three high-volatile bituminous coals to construct large-scale models (~40,000 atoms). The constructed models were utilized to examine the affinity for solvent extraction (naphthalene or tetralin) and the effect of structural properties (e.g., aromatic cluster size, functional groups, and cross-linking) in reactive molecular dynamics simulations. Complex chemical reactions were monitored with bond order transitions, intermediates formation, and mass distributions. Reactive molecular dynamics simulations suggest a plausible chemical extraction process and products for the complex fossil feedstocks. The results indicated that radical formations with bond breaking of bridging oxygens and carbons were required at high temperatures to facilitate hydrogeneration and extraction of gas molecules from radical-free molecules. We observed that aliphatic chains of tetralin were easily decomposed and combined with radicals to form small size of molecules with aryl bonding, mainly increasing molecules in the 500–1000 Da, while naphthalene had little impact on chemical extraction process.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Deep learning workflow for the inverse design of molecules with specific optoelectronic properties

The inverse design of novel molecules with a desirable optoelectronic property requires consideration of the vast chemical spaces associated with varying chemical composition and molecular size. First principles-based property predictions have become increasingly helpful for assisting the selection of promising candidate chemical species for subsequent experimental validation. However, a brute-force computational screening of the entire chemical space is decidedly impossible. To alleviate the computational burden and accelerate rational molecular design, we here present an iterative deep learning workflow that combines (i) the density-functional tight-binding method for dynamic generation of property training data, (ii) a graph convolutional neural network surrogate model for rapid and reliable predictions of chemical and physical properties, and (iii) a masked language model. As proof of principle, we employ our workflow in the iterative generation of novel molecules with a target energy gap between the highest occupied molecular orbital (HOMO) and the lowest unoccupied molecular orbital (LUMO).

97 MATHEMATICS AND COMPUTING↗

ORNL_AISD_DL-HLgap

This dataset provides supplementary molecular dataset of Deep Learning Workflow for the Inverse Design of Molecules with Specific Optoelectronic Properties. The dataset comprises three main directories such as GDB-9_dataset, Low_HL_Gap_dataset, and High_HL_Gap_dataset which individually has csv files, smiles_txt files, pdb files and xyz files containing information of molecular structures, properties and coordinates generated from deep learning workflow using generative model, surrogate model and DFTB calculation results. GDB-9_dataset contains the molecular data extracted from the original GDB-9 dataset with additional data of DFTB HL gap, surrogate HL gap and molecular property analysis. (the number of atoms, aromaticity and double bond equivalent) Low_HL_Gap_dataset and High_HL_Gap_dataset contains series of dataset for different generations with further split to train and test dataset that were obtained from the iterative workflow described in the manuscript. Additional directory Chemiscope_visualization in Low_HL_Gap_dataset directory contains compressed json files to visualize molecules using chemiscope.org page or application to help readers examine generated molecules.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Two excited-state datasets for quantum chemical UV-vis spectra of organic molecules

Abstract We present two open-source datasets that provide time-dependent density-functional tight-binding (TD-DFTB) electronic excitation spectra of organic molecules. These datasets represent predictions of UV-vis absorption spectra performed on optimized geometries of the molecules in their electronic ground state. The GDB-9-Ex dataset contains a subset of 96,766 organic molecules from the original open-source GDB-9 dataset. The ORNL_AISD-Ex dataset consists of 10,502,904 organic molecules that contain between 5 and 71 non-hydrogen atoms. The data reveals the close correlation between the magnitude of the gaps between the highest occupied molecular orbital (HOMO) and the lowest unoccupied molecular orbital (LUMO), and the excitation energy of the lowest singlet excited state energies quantitatively. The chemical variability of the large number of molecules was examined with a topological fingerprint estimation based on extended-connectivity fingerprints (ECFPs) followed by uniform manifold approximation and projection (UMAP) for dimension reduction. Both datasets were generated using the DFTB+ software on the “Andes” cluster of the Oak Ridge Leadership Computing Facility (OLCF).

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Supplementary Material for ORNL_AISD-Ex

This dataset provides supplementary material for the previously published dataset ORNL_AISD-Ex (1), which is available at the following website: https://www.osti.gov/biblio/1907919 The dates comprises one compressed folder called ornl_aisd_ex.zip. The compressed folder ornl_aisd_ex.zip contains 1,000 CSV files, each of them titled ornl_aisd_ex_ID.csv, where ID is a number that ranges between 1 and 1,000. The information contained in ornl_aisd_ex_ID.csv corresponds to the information of the molecules compresses inside the ornl_aisd_ex_ID.tar.gz of the dataset ORNL_AISD-Ex. Each row in each .CSV file is associated with a molecule, and the columns contain the following information: 1) molecules ID 2) SMILES string representation 3) DFTB-PE (eV): formation energy 4) first 50 electronic excitation modes 5) oscillators strengths of the first 50 electronic excitation modes The compression of this information into CSV files will allow a more agile extraction and management of information to the users that do not have access to large scale HPC platforms. REFERENCES (1) Lupo Pasini, Massimiliano, Mehta, Kshitij, Yoo, Pilsun, and Irle, Stephan. ORNL_AISD-Ex: Quantum chemical prediction of UV/Vis absorption spectra for over 10 million organic molecules. United States: N. p., 2023. Web. doi:10.13139/OLCF/1907919.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗