Search NASA⌕ Search

SEARCH · Search NASA

Results for “machine-readable scientific literature”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Towards a machine-readable literature: finding relevant papers based on an uploaded powder diffraction pattern

A prototype application for machine-readable literature is investigated. The program is called pyDataRecognition and serves as an example of a data-driven literature search, where the literature search query is an experimental data set provided by the user. The user uploads a powder pattern together with the radiation wavelength. The program compares the user data to a database of existing powder patterns associated with published papers and produces a rank ordered according to their similarity score. The program returns the digital object identifier and full reference of top-ranked papers together with a stack plot of the user data alongside the top-five database entries. The paper describes the approach and explores successes and challenges.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Database of Stress-Strain Properties Auto-generated from the Scientific Literature using ChemDataExtractor

Abstract There has been an ongoing need for information-rich databases in the mechanical-engineering domain to aid in data-driven materials science. To address the lack of suitable property databases, this study employs the latest version of the chemistry-aware natural-language-processing (NLP) toolkit, ChemDataExtractor, to automatically curate a comprehensive materials database of key stress-strain properties. The database contains information about materials and their cognate properties: ultimate tensile strength, yield strength, fracture strength, Young’s modulus, and ductility values. 720,308 data records were extracted from the scientific literature and organized into machine-readable databases formats. The extracted data have an overall precision, recall and F-score of 82.03%, 92.13% and 86.79%, respectively. The resulting database has been made publicly available, aiming to facilitate data-driven research and accelerate advancements within the mechanical-engineering domain.

Kumar, Pankaj↗

The Goddard Infrared Astronomical Data Base

The contents, structure, and principal products of the Goddard Infrared Astronomical Data Base are briefly reviewed. The data base is a machine-readable compilation of data obtained by searching the astronomical literature both in scientific journals and in infrared survey catalogs. The current data base contains more than 140,000 individual observations of at least 30,000 different infrared sources. The principal products of the data base are the Catalog of Infrared Observations and its associated appendices and the Infrared Source Cross-Index.

Mead, Jaylee M.↗

Far infrared supplement: Catalog of infrared observations

The development of a new generation of orbital, airborne and ground-based infrared astronomical observatory facilities, including the infrared astronomical satellite (IRAS), the cosmic background explorer (COBE), the NASA Kuiper airborne observatory, and the NASA infrared telescope facility, intensified the need for a comprehensive, machine-readable data base and catalog of current infrared astronomical observations. The Infrared Astronomical Data Base and its principal data product, this catalog, comprise a machine-readable library of infrared (1 micrometer to 1000 micrometers) astronomical observations published in the scientific literature since 1965.

Gezari, D. Y.↗

A review and evaluation of the Langley Research Center's Scientific and Technical Information Program. Results of phase 5. Design and evaluation of STI systems: A selected, annotated bibliography

A selected, annotated bibliography of literature citations related to the design and evaluation of STI systems is presented. The use of manual and machine-readable literature searches; the review of numerous books, periodicals reports, and papers; and the selection and annotation of literature citations were required. The bibliography was produced because the information was needed to develop the methodology for the review and evaluation project, and a survey of the literature did not reveal the existence of a single published source of information pertinent to the subject. Approximately 200 citations are classified in four subject areas. The areas include information - general; information systems - design and evaluation, including information products and services; information - use and need; and information - economics.

Pinelli, T. E.↗

PDF Entity Annotation Tool (PEAT)

While different text mining approaches – including the use of Artificial Intelligence (AI) and other machine based methods - continue to expand at a rapid pace, the tools used by researchers to create the labeled datasets required for training, modeling, and evaluation remain rudimentary. Labeled datasets contain the target attributes the machine is going to learn; for example, training an algorithm to delineate between images of a car or truck would generally require a set of images with a quantitative description of the underlying features of each vehicle type. Development of labeled textual data that can be used to build natural language machine learning models for scientific literature is not currently integrated into existing manual workflows used by domain experts. Published literature is rich with important information, such as different types of embedded text, plots, and tables that can all be used as inputs to train ML/natural language processing (NLP) models, when extracted and prepared in machine readable formats. Currently, both normalized data extraction of use to domain experts and extraction to support development of ML/NLP models are labor intensive and cumbersome manual processes. Automatic extraction of data and information from formats such as PDFs that are optimized for layout and human readability, not machine readability. The PDF (Portable Document Format) Entity Annotation Tool (PEAT) was developed with the goal of allowing users to annotate publications within their current print format, while also allowing those annotations to be captured in a machine-readable format. One of the main issues with traditional annotation tools is that they require transforming the PDF into plain text to facilitate the annotation process. While doing so lessens the technical challenges of annotating data, the user loses all structure and provenance that was inherent in the underlying PDF. Also, textual data extraction from PDFs can be an error prone process. Challenges include identifying sequential blocks of text and a multitude of document formats (multiple columns, font encodings, etc.). As a result of these challenges, using existing tools for development of NLP/ML models directly from PDFs is difficult because the generated outputs are not interoperable. We created a system that allows annotations to be completed on the original PDF document structure, with no plain text extraction. The result is an application that allows for easier and more accurate annotations. In addition, by including a feature that grants the user the ability to easily create a schema, we have developed a system that can be used to annotate text for different domain-centric schemas of relevance to subject matter experts. Different knowledge domains require distinct schemas and annotation tags to support machine learning.

97 MATHEMATICS AND COMPUTING↗