Search NASASearch

SEARCH · Search NASA

Results for “Natural language processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

The second generation intelligent user interface for the crustal dynamics data information system

For the past decade, operations and research projects that support a major portion of NASA's overall mission have experienced a dramatic increase in the volume of generated data and resultant information that is unparalleled in the history of the agency. The effect of such an increase is that most of the science and engineering disciplines are undergoing an information glut, which has occurred, not only because of the amount, but also because of the type of data being collected. This information glut is growing exponentially and is expected to grow for the foreseeable future. Consequently, it is becoming physically and intellectually impossible to identify, access, modify, and analyze the most suitable information. Thus, the dilemma arises that the amount and complexity of information has exceeded and will continue to exceed, using present information systems, the ability of all the scientists and engineers to understand and take advantage of this information. As a result of this information problem, NASA has initiated the Intelligent Data Management (IDM) project to design and develop Advanced Information Management (IDM) project to design and develop Advanced Information Management Systems (AIMS). The first effort of the Project was the prototyping of an Intelligent User Interface (IUI) to an operational scientific database using expert systems, natural language processing, and graphics technologies. An overview of the IUI formulation and development for the second phase is presented.

Short, Nicholas, Jr.

An intelligent user interface for browsing satellite data catalogs

A large scale domain-independent spatial data management expert system that serves as a front-end to databases containing spatial data is described. This system is unique for two reasons. First, it uses spatial search techniques to generate a list of all the primary keys that fall within a user's spatial constraints prior to invoking the database management system, thus substantially decreasing the amount of time required to answer a user's query. Second, a domain-independent query expert system uses a domain-specific rule base to preprocess the user's English query, effectively mapping a broad class of queries into a smaller subset that can be handled by a commercial natural language processing system. The methods used by the spatial search module and the query expert system are explained, and the system architecture for the spatial data management expert system is described. The system is applied to data from the International Ultraviolet Explorer (IUE) satellite, and results are given.

Cromp, Robert F.

The development of an intelligent user interface for NASA's scientific databases

The National Space Science Data Center (NSSDC) has initiated an Intelligent Data Management (IDM) research effort which has as one of its components, the development of an Intelligent User Interface (IUI). The intent of the IUI effort is to develop a friendly and intelligent user interface service that is based on expert systems and natural language processing technologies. This paper presents the design concepts, development approach and evaluation of performance of a prototype Intelligent User Interface Subsystem (IUIS) supporting an operational database.

Campbell, William J.

Logic programming and metadata specifications

Artificial intelligence (AI) ideas and techniques are critical to the development of intelligent information systems that will be used to collect, manipulate, and retrieve the vast amounts of space data produced by 'Missions to Planet Earth.' Natural language processing, inference, and expert systems are at the core of this space application of AI. This paper presents logic programming as an AI tool that can support inference (the ability to draw conclusions from a set of complicated and interrelated facts). It reports on the use of logic programming in the study of metadata specifications for a small problem domain of airborne sensors, and the dataset characteristics and pointers that are needed for data access.

Lopez, Antonio M., Jr.

Proceedings of the 1993 Conference on Intelligent Computer-Aided Training and Virtual Environment Technology

The volume 2 proceedings from the 1993 Conference on Intelligent Computer-Aided Training and Virtual Environment Technology are presented. Topics discussed include intelligent computer assisted training (ICAT) systems architectures, ICAT educational and medical applications, virtual environment (VE) training and assessment, human factors engineering and VE, ICAT theory and natural language processing, ICAT military applications, VE engineering applications, ICAT knowledge acquisition processes and applications, and ICAT aerospace applications.

Hyde, Patricia R.

Methods for semi-automated indexing for high precision information retrieval

OBJECTIVE: To evaluate a new system, ISAID (Internet-based Semi-automated Indexing of Documents), and to generate textbook indexes that are more detailed and more useful to readers. DESIGN: Pilot evaluation: simple, nonrandomized trial comparing ISAID with manual indexing methods. Methods evaluation: randomized, cross-over trial comparing three versions of ISAID and usability survey. PARTICIPANTS: Pilot evaluation: two physicians. Methods evaluation: twelve physicians, each of whom used three different versions of the system for a total of 36 indexing sessions. MEASUREMENTS: Total index term tuples generated per document per minute (TPM), with and without adjustment for concordance with other subjects; inter-indexer consistency; ratings of the usability of the ISAID indexing system. RESULTS: Compared with manual methods, ISAID decreased indexing times greatly. Using three versions of ISAID, inter-indexer consistency ranged from 15% to 65% with a mean of 41%, 31%, and 40% for each of three documents. Subjects using the full version of ISAID were faster (average TPM: 5.6) and had higher rates of concordant index generation. There were substantial learning effects, despite our use of a training/run-in phase. Subjects using the full version of ISAID were much faster by the third indexing session (average TPM: 9.1). There was a statistically significant increase in three-subject concordant indexing rate using the full version of ISAID during the second indexing session (p < 0.05). SUMMARY: Users of the ISAID indexing system create complex, precise, and accurate indexing for full-text documents much faster than users of manual methods. Furthermore, the natural language processing methods that ISAID uses to suggest indexes contributes substantially to increased indexing speed and accuracy.

Evaluation Studies

Earth Science Deep Learning: Applications and Lessons Learned

Deep learning has revolutionized computer vision and natural language processing with various algorithms scaled using high-performance computing. At the NASA Marshall Space Flight Center (MSFC), the Data Science and Informatics Group (DSIG) has been using deep learning for a variety of Earth science applications. This paper provides examples of the applications and also addresses some of the challenges that were encountered.

Training

Reusing Data and Metadata to Create New Metadata Through Machine-Learning & Other Programmatic Methods

Recent improvements in natural language processing (NLP) enable metadata to be created programmatically from reused original metadata or even the dataset itself. Transfer-learning applied to NLP has greatly improved performance and reduced training data requirements. In this talk, we’ll compare machine-generated metadata to human-generated metadata and discuss characteristics of metadata and data archives that affect suitability for machine-learning reuse of metadata. Where as human-generated metadata is often populated once, populated from the perspective of data supplier, populated by many individuals with different words for the same thing, and limited in length, machine-generated metadata can be updated any number of times, generated from the perspective of any user, constrained to a standardized set of terms that can be evolved over time, and be any length required. Machine-learning generated metadata offers benefits but also additional needs in terms of version control, process transparency, human-computer interaction, and IT requirements. As a successful example, we’ll discuss how a dataset of abstracts and associated human-tagged keywords from a standardized list of several thousand keywords were used to create a machine-learning model that predicted keyword metadata for open-source code projects on code.nasa.gov. We’ll also discuss a less successful example from data.nasa.gov to show how data archive architecture and characteristics of initial metadata can be strong controls on how easy it is to leverage programmatic methods to reuse metadata to create additional metadata.

Gosses, Justin

Deep Learning Models for Planetary Seismicity Detection

Research in planetary seismology is fundamentally constrained by a lack of data. Seismo-logical science products of future missions can typically only be informed by theoretical signal/noise characteristics of the environment or likely Earth-analogues. Although objectives can be re-assessed after some initial data-collection upon lander arrival, transfer of high-resolution data back to Earth is costly on lander power usage. Over the last several years, development of GPU computing techniques and open-source high-level APIs have led to rapid advances in deep learning within the fields of computer vision, natural language processing, and collaborative filtering. These techniques are actively being adapted in seismology for a variety of tasks, including: earthquake detection, seismic phase discrimination, and ground-motion prediction. Until the recent detection of mars quakes during the Mars InSight mission, the only other measurements of seismicity recorded outside of Earth was on the Moon during the Apollo missions between 1969 to 1977. These unique data sets have been periodically revisited using new seismological methods, including ambient noise interferometry and Hidden Markov Models. Our objective is to develop a deep learning seismic detector and use it to catalog moonquakes from the Apollo 17 Lunar Seismic Profiling Experiment (LSPE) and compare the results with those obtained by other methods. Additionally, we will assess the accuracy tradeoff between using a training set of lunar data and one composed of Earth seismicity. In this document, we present preliminary results using a prototype classifier trained on a small set of earthquakes that was able to obtain detections for LSPE moonquakes with a greater accuracy than a recent study using Hidden Markov Models.

Civilini, F.

Trusted Communication: Utilizing Speech Communication to Enhance Human-Machine Teaming Success

An area of increasing interest for the next generation of aircraft is autonomy and the integration of increasingly autonomous systems into the national airspace. Such an integration requires humans to work closely with autonomous systems, forming teams. Our hypothesis is that a team composed of both humans and autonomous systems will operate better than either entity alone. We have existing procedures for certifying pilots to operate in the national airspace and are currently working on methods for validating the function of autonomous systems, however we have no method in place for assessing the interaction of these two disparate systems. Communication is one avenue. This paper will examine the use of language as a metric for ascertaining human-machine teaming effectiveness. A proof-of-concept of the application of two communication-based analysis techniques, Linguistic Inquiry and Word Count (LIWC) and Latent Semantic Analysis (LSA), for the prediction of success in human/chatbot teaming was conducted. By running these analyses over data from the 2014 and 2015 Loebner Prize competitions of human/chatbot teaming, numerical scores were obtained that can be associated with scores provided by human judges during the competition. Correlating their LIWC and LSA data with the scores provided by the judges, and using linear regression over this correlation, formulae were obtained that predict the score of human/chatbot interaction. These formulae were tested over the 2013 Loebner Prize transcripts, determining that, though there was strong correlation between predicted and actual scores, the predictive success of this method was not strong. However, with specialized topic spaces and lexica, as well as larger data sets, the predictive power of these metrics will improve. Given the importance of providing metrics for human-machine system team success and given the promise shown by the communication-basedLIWCand LSAmethods, continuing research in this area is necessary. After examining the potential for using communication and spoken language as a metric for the success of human/autonomous system teaming, this paper then examines aspects inherent to communication systems that may contribute to unreliability and reduced trust. Modern natural language processing tools rely on deep learning algorithms to create language rules that produce accurate results, but these rules are uninterpretable. The resulting blackbox system lacks transparency necessary for full validation and complete trust. Additionally, speech-based interfaces pose other difficulties to developing coordinated teamwork between humans and autonomous systems. Human communication is infrequently limited to speech only, instead usually relying on a combination of verbal, gestural, and general body language communication. Reducing an analysis of team effectiveness to a study of spoken language alone is problematic as it leaves these other equally important forms of communication out. This paper will examine these problems and the general deficiencies in speech-based metrics for human-machine teaming.

E L Meszaros

Autonomy Voice Assistant for NPAS (NASA Platform for Autonomous Systems)

A prototype voice interaction system, Autonomy Voice Assistant (AVA), is described in this paper. AVA is designed to seamlessly integrate into the NASA Platform for Autonomous Systems (NPAS), an autonomy software platform, and to enable an operator to interact with NPAS autonomy applications through voice conversations. By integrating VA with NPAS, a major enhancement to NPAS applications is facilitated, enabling interaction through natural language expressions. An AVA prototype has been designed incorporating two principles:(1) self-containment (no external data or computations required), and (2) a readily modifiable, reconfigurable, and flexible architecture. By using voice messages in an NPAS application, an additional layer of user interface capability is enabled, thereby enhancing a user’s overall experience. Advancements, over the past several decades in speech recognition and natural language processing technologies has made it possible for AVA to implement robust messaging capabilities while still being lightweight. The main objective of incorporating a voice assistant like AVA is to augment the number and effectiveness of interactions a user has with a system that typically uses mouse-based interaction, while simultaneously enriching the user experience and providing heightened system awareness.

Lucian Murdock

A Hybrid Approach to Labeling Datasets in Earth Science Publications

NASA Data Centers provide the public with thousands of datasets that result in published papers, reports, and conference proceedings. Collecting accurate metrics on usage of these datasets is key to connecting different areas of knowledge and evaluating the datasets’ impact. While most of the datasets have Digital Object Identifiers (DOIs) assigned, most publications do not cite them hampering the automated search of these publications. Instead, articles mention attributes like organization, instrument, mission, variable, or a publication describing the dataset. Often only domain experts can deduce the dataset that was used in the publication text. The lack of a citation slows the spread of information and reduces the research’s impact. With thousands of papers produced each year, an automated means of labeling datasets is critical. This paper explores a hybrid approach of heuristics and a Natural Language Processing (NLP) Named Entity Recognition (NER) model to find and label the datasets used within Earth Science papers. Heuristics are used to produce the labelled sentences and any potential dataset candidates that can be derived from a sentence. The heuristic labels the sentences with the names of mission, instrument, re-analysis models, and science keywords taken from the Global Change Master Directory (GCMD) ontology. Additionally, it uses those labels to generate the dataset citation candidates. If the mission, instrument, and variable are sufficient to create the citation for the dataset the citation and the label the domain expert reviews the output without going through the NLP model. If the extracted label is not sufficient to label the dataset on its own, the sentence and its associated dataset labels will be inputted into the NER model. The model outputs the labeled sentence and the potential dataset candidates with their associated probabilities. The domain expert then reviews the NER model’s output and the correct labels are determined. The newly labelled papers can then be used as additional training data. This creates an iterative process for the approach to continuously improve. Because all the possible mentions are gathered by the model, the domain expert can quickly and easily label the papers resulting in large time savings.

Jacob Atkins

Enhancing NASA Earth Science Data Discovery from Scientific Publications

Earth observations from space borne instruments have evolved explosively in the past decades. Following closely are reanalysis systems assimilating model and observational data, yielding even longer records and larger number of variables. Thanks to advances in internet technology, it is now easier than ever to visualize and analyze these data using web interfaces. On the other hand, it also becomes an increasingly daunting task to build upon the existing knowledge published in various peer reviewed sources, and navigate toward the most relevant data, analysis, and visualization. We present an analysis of a subset of publications that utilized a popular visualization web interface at the NASA Goddard Earth Science Data and Information Services Center. Known as "Giovanni", it allows researchers from wide backgrounds to work with hundreds of variables from space observations and assimilation systems. Since coming online more than a decade ago, Giovanni has been credited in more than 100 papers per year, and the total count now is estimated to be nearly 1,500. Many of these papers contain valuable information about when, where and how Giovanni has been used, and hence forge an opportunity to learn and share the knowledge of which variables were used for what research projects. The purpose of our work is to retrieve the information from the papers and organize it as a knowledge repository which links together datasets, variables, places, dates and phenomena all of which reflect the essence of the published research. Since the publications are unstructured texts, we use natural language processing along with machine learning methods in the retrieval process. One of the challenges is deciphering the dataset names, because in many cases researchers refer to variables, rather than the datasets containing them. To constrain the number of terms, we deploy Earth Science ontologies as dictionaries for the term extraction. We demonstrate that storing these terms and underlying ontologies, along with datasets, variables and papers in the knowledge graph database, enables various linkages between all these entities facilitating the data discovery. Thus, we are setting a qualitatively new stage in improvements of web data interfaces, where machine learning techniques are used to establish and optimize usage-based discovery of data.

Irina V Gerasimov

A Brief Introduction to AI/ML Applications of Air Traffic Management Data at NASA Ames

This presentation will give a brief overview of several AI/ML This presentation will give a brief overview of several AI/ML projects that NASA Ames interns are exploring in partnership with NASA Aeronautic Research Institute (NARI) and the FAA. NASA is interested in Natural Language Processing (NLP) of various legacy text and speech data within air traffic management e.g., Notices To Airmen (NOTAMs), Letters of Agreement (LoAs), Standard Operating Procedures (SOPs), and Air Traffic Control Center audio briefings. Since our focus is on applying state of the art AI/ML tools to legacy air traffic management data, we first showcase the different data sources of interest followed by a brief introduction to the techniques and language models used. We present some exciting preliminary results on each topic including both unsupervised learning techniques (e.g., clustering) and other modern language models (e.g., BERT) that help extract useful information from these data sources that are interpretable by both man and machine.

Air Traffic Management

Verb Sense Disambiguation for Densifying Knowledge Graphs in Earth Science

We begin with an ambitious goal: to create a knowledge graph that spans the entire discipline of Earth science. In order to achieve this, we need to apply Natural Language Processing (NLP) techniques on Earth science journal articles to extract their semantic components for the graph. When sentences from Earth science journal articles are broken down into their semantic components and loaded onto a graph, the relationships among these semantic components are represented by the verbs in the sentences. However, since there are multiple verbs in English that can be used to denote the same meaning, the knowledge graph can become sparse and so can the results when we query the graph. In order to ensure quality results, it would be desirable to consolidate similar verbs into a single "class". So, this is the problem at hand: how do we make sure that multiple verbs that mean the same thing are represented as a single class of verb in the knowledge graph? Or in other words, how do we distinguish which meaning a particular verb takes given a particular sentence? In this poster, we demonstrate a potential technique to solve this problem.

Ashish Acharya

TopiQAL: Topic-aware Question Answering using Scalable Domain-specific Supercomputers

We all have questions. About today's temperature, scores of our favorite baseball team, the Universe, and about vaccine for COVID-19. Life, physical, and natural scientists have been trying to find answers to various topics using scientific methods and experiments, while computer scientists have built language models as a tiny step towards automatically answering all of these questions across domains given a little bit of context. In this paper, we propose an architecture using state-of-the-art Natural Language Processing language models namely Topic Models and Bidirectional Encoder Representations from Transformers (BERT) that can transparently and automatically retrieve articles of relevance to questions across domains, and fetch answers to topical questions related to COVID-19 current and historical medical research literature. We demonstrate the benefits of using domain-specific supercomputers like Tensor Processing Units (TPUs), residing on cloud-based infrastructure, using which we could achieve significant gains in training and inference times, also with very minimal cost.

Penberthy, Scott

Integrating Human System Information with the Systems Platform for Aggregating and Relating Capabilities (SPARC)

Within Human Health and Performance, there exists a wealth of human system information that’s used a regular basis in support of NASA human exploration objectives, but the challenge is that all of this information was stored in multiple different locations and organized for specific uses, limiting its effectiveness and straining communication across multiple groups. To address this challenge, our project, the Systems Platform for Aggregating and Relating Capabilities (SPARC) was tasked with developing a new NASA internal web application that aggregates and relates multiple programs’ human system products, such as technical standards, program requirements and verifications, human system risks, research and evidence, and exploration capabilities, into one centralized platform that addresses the needs of human health and performance from multiple different perspectives. Using agile development methodologies, user-experience (UX) driven design principles, data visualization, and a strong emphasis on continuous improvement though consistent stakeholder engagement, the SPARC project released a beta version in less than 4 months, broadly launched version 1.0.0 Agency-wide three months after the beta, and has over 180 users in the first year of development. Our second year of development will see us moving from our initial capabilities to increasingly robust and complex integrations and visualizations, including Directed Acyclical Graphs (DAGs), natural language processing (NLP) for dynamic generation of relationships between the sources of truth, and an expansion into hierarchical levels of system design in support of the human exploration programs.

Data science

BERT-Based Topic Modeling and Information Retrieval to Support Fishbone Diagramming for Safe Integration of Unmanned Aircraft Systems in Wildfire Response

Recent concepts for emerging wildfire response operations have included unmanned aircraft systems (UAS) due to their increasing accessibility and capabilities. To integrate UAS into wildfire response safely, researchers have studied the use of large repositories of historic incident reports to improve the scope of root cause analysis. Recent work has emphasized applying state-of-the-art natural language processing techniques to extract useful information from these repositories. However, it has not yet been studied how these results can be interpreted and integrated into the systems engineering process. In this work, we propose a process in which Bidirectional Encoder Representations from Transformers (BERT)-based topic modeling and information retrieval are applied to a relevant set of documents in order to support the development of a fishbone diagram in a semiautomated process. High-level themes in the document set are identified using topic modeling, which are then refined and interpreted by a human analyst. Then, the themes are used to guide a finer search using information retrieval, which returns specific incident reports of relevance. This provides traceability to specific incidents as well as broader categorizations that comprise the fishbone branches. We apply the proposed process to relevant documents from NASA’s Aviation Safety Reporting System (ASRS). The proposed process is widely applicable when relevant documents are available, and the results from this study will be useful to identifying potential causes of wildfire response UAS incidents.

Hazard analysis