Search NASA⌕ Search

Engineering topics

Cole, Jacqueline M.

Publications and source records attributed to Cole, Jacqueline M..

How Beneficial Is Pretraining on a Narrow Domain-Specific Corpus for Information Extraction about Photocatalytic Water Splitting?

Language models trained on domain-specific corpora have been employed to increase the performance in specialized tasks. However, little previous work has been reported on how specific a “domain-specific” corpus should be. Here, we test a number of language models trained on varyingly specific corpora by employing them in the task of extracting information from photocatalytic water splitting. We find that more specific corpora can benefit performance on downstream tasks. Furthermore, PhotocatalysisBERT, a pretrained model from scratch on scientific papers on photocatalytic water splitting, demonstrates improved performance over previous work in associating the correct photocatalyst with the correct photocatalytic activity during information extraction, achieving a precision of 60.8(+11.5)% and a recall of 37.2(+4.5)%.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A database of thermally activated delayed fluorescent molecules auto-generated from scientific literature with ChemDataExtractor

A database of thermally activated delayed fluorescent (TADF) molecules was automatically generated from the scientific literature. It consists of 25,482 data records with an overall precision of 82%. Among these, 5,349 records have chemical names in the form of SMILES strings which are represented with 91% accuracy; these are grouped in a subsidiary database. Each data record contains one of the following four properties: maximum emission wavelength (λ EM ), photoluminescence quantum yield (PLQY), singlet-triplet energy splitting (ΔE ST ), and delayed lifetime (τ D ). The databases were created through text mining using ChemDataExtractor, a chemistry-aware natural-language-processing toolkit, which has been adapted for TADF research. The text-mined corpus consisted of 2,733 papers from the Royal Society of Chemistry and Elsevier. To the best of our knowledge, these databases are the first databases that have been auto-generated for TADF molecules from existing publications. The databases have been publicly released for experimental and computational applications in the TADF research field.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Automated Construction of a Photocatalysis Dataset for Water-Splitting Applications

We present an automatically generated dataset of 15,755 records that were extracted from 47,357 papers. These records contain water-splitting activity in the presence of certain photocatalysts, along with additional information about the chemical reaction conditions under which this activity was recorded. These conditions include any co-catalysts and additives that were present during water splitting, the length of time for which the photocatalytic experiment was conducted, and the type of light source used, including its wavelength. Despite the text extraction of such a wide range of chemical reaction attributes, the dataset afforded good precision (71.2%) and recall (36.3%). These figures-of-merit were calculated based on a random sample of open-access papers from the corpus. Mining such a complex set of attributes required the development of novel techniques in knowledge extraction and interdependency resolution, leveraging inter- and intra-sentence relations, which are also described in this paper. We present a new version (version 2.2) of the chemistry-aware text-mining toolkit ChemDataExtractor, in which these new techniques are included.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

In‐Silico Device Performance Prediction of Cosensitizer Dye Pairs for Dye‐Sensitized Solar Cells

Abstract Endeavors in the field of dye‐sensitized solar cells (DSCs) have shown great promise when adopting a data‐driven approach to materials discovery, such as successful molecular‐scale predictions of light‐harvesting chromophores. However, predictions of DSC dyes would become much more sophisticated if a molecular‐to‐macroscopic DSC device prediction methodology existed. Thereby, a fully computational pipeline is presented that predicts device‐performance parameters of DSCs which contain varying dye combinations. Optimal pairing of complementary dyes is identified via a data‐driven workflow that affords cosensitized DSCs with maximum power‐conversion efficiencies. Six high‐performing DSC dyes are paired with partner dyes that are screened from a database of 8488 compounds using sequential heuristic filters. Existing models that predict short‐circuit‐current density ( J SC ) and open‐circuit voltage ( V OC ) parameters are adapted to predict singly sensitized and cosensitized DSC performance. The predictions for J sc values of singly sensitized devices match experimental literature values with comparable accuracy to more computationally costly methods. Five out of six dye pairings are predicted to have greater J SC values when cosensitized compared to their corresponding singly sensitized devices, including two pairs that show strong J sc boosts of +13% and +12% when cosensitized. Thus, the prospect of an entirely in‐silico prediction pipeline for DSC performance that can be used to realize the fully automated design of optimized cosensitized DSCs is demonstrated.

14 SOLAR ENERGY↗