Search NASASearch

SEARCH · Search NASA

Results for “Natural Language Processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Wildfire Emergency Response Hazard Extraction and Analysis of Trends (HEAT) through Natural Language Processing and Time Series

A methodology for Hazard Extraction and Analysis of Trends (HEAT) is proposed and conducted on a data set of wildfire incident response forms, known as ICS-209-PLUS.The HEAT processes: (1) extract a set of hazards from a data set, (2) calculate hazard-relevant metrics in a primary analysis, (3) analyze trends over time in metrics using timeseries, and (4) examine potential explanations for metric trends using a secondary analysis. Hazards are extracted from narrative data in the ICS-209-PLUS based on a framework previously developed by the authors, using natural language processing. Metrics examined for each hazard include operational time to occurrence, rate of occurrence, frequency, and severity. Primary results include a taxonomy of hazards present in the data set with relevant quantitative metrics. The most frequent hazards identified are environmental and include hazardous terrain. Most hazards occur on average between 35-55% containment. Incidents with hazards tend to have a higher average severity score when compared to the average score for all incidents. Time series of the metrics and relevant predictors, including fire characteristics, fire intensity, and operations, are created to facilitate further analysis. Secondary results used to determine which factors best predict hazard frequency include a correlation matrix and regression analysis. These findings are relevant to safety for current, as well as emerging wildfire operations, and are an exploratory first step in developing historical data-driven risk assessment models.

Sequoia R. Andrade

Molecular property prediction for very large databases with natural language processing: a case study in ionic liquid design

The prospect of using artificial intelligence (AI) to accurately screen very large databases of compounds for multiple properties has yet to be realized. Here, we explore this possibility using ionic liquids (ILs) which offer unique physicochemical properties and excellent tunability, making them highly versatile solvents for various research applications. Screening millions of potential ILs for the best perfomance for use in specific tasks with experimental methods alone however, is impractical. Further, traditional’ physics-based computational chemistry is hindered by high computational cost. To address this challenge, we leverage a natural language processing (NLP)-based molecular embedding technique with advanced machine learning (ML) models to predict seven key IL properties: viscosity, density, ionic conductivity, surface tension, melting temperature, toxicity, and water solubility. Comprehensive datasets for these properties are obtained, then NLP featurization with Mol2vec is compared with other featurization techniques such as 2D Morgan fingerprints, and 3D quantum chemistry-derived sigma profiles. NLP-based featurization exhibited the best predictive performance, achieving the highest R 2 and lowest RMSE values for all the studied IL properties. Further, we present case studies of how ILs might be screened using combined property criteria for practical cases – lignocellulosic biomass processing, CO 2 capture, and optimal electrolytes for batteries – screening a novel database of ∼10.6 million generated feasible ILs. The results introduce NLP as a powerful tool for engineering many designer solvents with desirable properties for task specific applications.

Mohan, Mood [Oak Ridge National Laboratory (ORNL),

Search Enhancements using Natural Language Processing Techniques

NASA Goddard Earth Sciences Data and Information Services Center (GESDISC) is one of the 12 NASA Science Mission Directorate Data Centers. The main goal of GESDISC is to provide earth science data, information, and services to the earth science data community. Consequently, data discovery is at the center of our mission and our search engine is the primary tool for our users to interact, find, and access our data. Existing search approaches are largely focused on hard-matching of keywords in the search query with dataset metadata. Here we propose to expand the search by introducing a complementary natural language processing (NLP) search. At the heart of our proposed NLP search, we trained a joint embedding using scientific text corpus and a curated set of dataset metadata. The embedding learns the association between words in our dataset metadata and those of the scientific text corpus. This enables us to go beyond simple hard-matching of a query and data set metadata and have a notion of “similarity” between the search query and the datasets. We further integrated our NLP search into the Elastic Search (ES) framework leveraging similarity search capabilities offered through the “dense_vector” field type. Our preliminary evaluations show that our proposed NLP search has the potential to be utilized to complement the existing search engine and serve as a base for a dataset recommendation system.

Armin Mehrabian

Artificial intelligence, expert systems, computer vision, and natural language processing

An overview of artificial intelligence (AI), its core ingredients, and its applications is presented. The knowledge representation, logic, problem solving approaches, languages, and computers pertaining to AI are examined, and the state of the art in AI is reviewed. The use of AI in expert systems, computer vision, natural language processing, speech recognition and understanding, speech synthesis, problem solving, and planning is examined. Basic AI topics, including automation, search-oriented problem solving, knowledge representation, and computational logic, are discussed.

Gevarter, W. B.

Digital Assistance for System Requirement Discovery and Analysis using Machine Learning Natural Language Processing Algorithm

NASA’s Air Traffic Management-Exploration (ATM-X) Urban Air Mobility (UAM) Airspace Subproject is conducting research that evolves UAM airspace towards a highly automated and operationally flexible system of the future (see https://www.nasa.gov/uam-overview/ for more information). The complexity of UAM airspace, and its evolution through a series of transformative epochs, requires a planning tool to effectively organize, integrate, and communicate the research that will guide the evolution of UAM operations in the National Airspace System (NAS). The planning tool, called the UAM airspace research roadmap (or just roadmap), is being developed as a new system engineering methodology leveraging model based system engineering (MBSE) and machine learning natural language processing (ML NLP, or just NLP) capabilities. This presentation gives an overview of the NLP application within this system engineering methodology and will describe how it is being used to meet the ATM-X UAM Airspace Subproject’s overarching research goals.

ATM

Knowledge Discovery for Early Failure Assessment of Complex Engineered Systems Using Natural Language Processing

Emerging complex engineered systems may have unexpected safety issues due to novel operational environments, increasing autonomy, human-machine interaction, and other factors. To prevent failures in operation or testing that necessitate costly redesign, it is desirable to predict likely failure modes early in the design process. Text-based information about past engineering failures presents one possible solution by facilitating the retrieval of information that can inform new designs. However, identifying documents containing relevant information and extracting required information can be prohibitively time-consuming when implemented at scale. In this research, an automated natural language processing-based framework is proposed to discover relevant knowledge from documents containing failure-related design information. Documents containing usable information are filtered using sentiment analysis based on a custom lexicon specialized for engineering design and by filtering out documents containing only irrelevant topics. Next, from the identified usable documents, information relating to engineering failures, contributing factors that can be controlled at design time (“risk factors”), and recommended preventative actions are extracted. Semantic similarity is then used to group similar pieces of extracted information for improved generalizability. The proposed framework is applied to NASA’s Lessons Learned Information System (LLIS). The framework can be used to identify documents containing usable failure-related design information from other databases, extract relevant information from these documents, and generalize the acquired knowledge such that it can be applied to novel systems.

Sequoia R. Andrade

Natural Language Processing for Extracting Rich Disease Data Aligned To Satellite Meteorological Data

Global climate change is redefining our understanding of how diseases spread. In Sri Lanka, vector-borne diseases such as dengue fever, encephalitis, and leptospirosis historically surged during the monsoon seasons when temperatures were high enough for mosquito eggs to hatch. Unfortunately, due to rising temperatures and more erratic rainfall patterns, mosquito eggs can now hatch year-round and are increasingly unpredictable, leading to an alarmingly increasing number of hospitalizations and deaths. More data is needed to adapt our response to these diseases in an increasingly warmer world. In the contemporary landscape, a wealth of disease information is available, yet accessibility remains limited due to unstructured data formats such as PDFs. Therefore, converting unstructured disease reports into structured formats is necessary for effectively leveraging data. This paper introduces a comprehensive framework for collecting unstructured disease reports and transforming them into analyzable formats. By creating separate models tailored to each data format, we can ensure accuracy compared to general models. These straightforward models enhance accessibility and empower other researchers to use our tools. The returned structured data can then be harnessed for analysis, statistical purposes, and informing evidence-based public health interventions, thus facilitating more informed decision-making in healthcare. We deploy this framework to produce geospatial data for Sri Lanka and Brazil for many different conditions and align these data with satellite environmental data, providing for the first time a structured, aligned powerful dataset for disease modeling.

Open Source Open Science

Natural Language Processing Analysis of Notices to Airmen for Air Traffic Management Optimization

With new emerging technologies in the field of NLP, we explore their applications to digitize and analyze heritage Air Traffic Management (ATM) documents for planning and optimizing airspace operations. Specifically, this research focuses on harvesting semi-structured or un-structured information contained in Notices to Airmen (NOTAMs). Using NLP and other advanced data analytics, we will construct a data-driven framework which facilitates finding language patterns and the use of pretrained language models for classification and extraction of useful airspace constraints and restrictions. These may lead to tools that assist airspace users in understanding the constraints more efficiently, contributing to better route planning and safer execution. This paper explores three workflows entailing different NLP tasks. First, unsupervised techniques like word embedding and topic modeling are used for pattern finding and document classification. Second, a dataset is created by extracting information from the semi-structured NOTAM format as metadata for categorizing, visualizing, and extracting key entities driving NOTAM content. Third, modern pre-built deep learning based transformer models such as BERT, RoBERTa, and XLNet are evaluated on the question answering task, an even more robust approach to information extraction, as well as their respective fine-tuning tasks. In this work we include various performance metrics for the trained models to evaluate both accuracy and precision and we show that the models can be generalized for their respective tasks. The research work developed shows promise in uncovering trends in digital NOTAMs in the NAS and also offers a new framework for digitizing and inferring insights from free-form legacy NOTAMs, that are yet to be digitized.

Natural Language Processing

Natural Language Processing (NLP) Analysis of NOTAMs for Air Traffic Management Optimization

With new emerging technologies in the field of NLP, we explore their applications to digitize and analyze heritage Air Traffic Management (ATM) documents for planning and optimizing airspace operations. Specifically, this research focuses on harvesting semi-structured or un-structured information contained in Notices to Airmen (NOTAMs). Using NLP and other advanced data analytics, we will construct a data-driven framework which facilitates finding language patterns and the use of pretrained language models for classification and extraction of useful airspace constraints and restrictions. These may lead to tools that assist airspace users in understanding the constraints more efficiently, contributing to better route planning and safer execution. This paper explores three workflows entailing different NLP tasks. First, unsupervised techniques like word embedding and topic modeling are used for pattern finding and document classification. Second, a dataset is created by extracting information from the semi-structured NOTAM format as metadata for categorizing, visualizing, and extracting key entities driving NOTAM content. Third, modern pre-built deep learning based transformer models such as BERT, RoBERTa, and XLNet are evaluated on the question answering task, an even more robust approach to information extraction, as well as their respective fine-tuning tasks. In this work we include various performance metrics for the trained models to evaluate both accuracy and precision and we show that the models can be generalized for their respective tasks. The research work developed shows promise in uncovering trends in digital NOTAMs in the NAS and also offers a new framework for digitizing and inferring insights from free-form legacy NOTAMs, that are yet to be digitized. Video is an mp4 download, with a play time of 9 min 35 secs.

Natural Language Processing

Improving NASA Earth Science Data and Information Access Through Natural Language Processing Based Data Analysis and Visualization

NASA: The Research Access initiative is part of the agency's framework for increasing public access to scientific publications and digital scientific data. The initiative follows the release of White House Office of Science and Technology Policy's (OSTP) memorandum "Increasing Access to the Results of Federally Funded Research," to ensure federally funded research is available to the public within one year of publication. NASA answered the mandate by creating an agency plan entitled "NASA Plan for Increasing Access to the Results of Scientific Research" and associated policy, NPD 2230.1, Research Data and Publication Access. Principles in NASA SMD Strategic Plan for Scientific Data and Computing: Continued free and open access to scientific data for any use. Improved ease of use and discoverability. Enhanced science applications and new use cases. Incorporates best practices and "state of the art" through partnerships. Earth Data and Systems are Evolving: Increasing archive and file sizes. More complicated data structures. More user-friendly and data services. What is the future direction?

Liu, Zhong

Natural Language Processing to Inform Agent-Based Modeling: With Application to Modeling Adoption of Medium-Duty Electric Vehicles

Agent-based socio-technical modeling of medium- and heavy-duty (MDHD) electric vehicle (EV) adoption has the potential to provide analysis, prediction, and gui. This paper describes new applications of text analysis developed through machine learning (ML) to build and understand relevant topics and their saliency in the published discourse on adoption of MDHD EVs. This work contributes to the state of the art in topic mining models by defining a new metric of topic ranking (START) that quantifies the importance of predefined topics within the corpus using weighted results for predefined topics from two topic modeling approaches: Latent Dirichlet Allocation (LDA) and BERTopic. The START metric is then demonstrated in practice to model how academia and industry view the EV adoption process based on the respective texts published by these groups. Results show that academic literature places more emphasis on categories of interests such as norms/attitudes and adopter knowledge, while trade journals tend to emphasize long-term cost more than academia. The two bodies of literature agree on the importance of policy and incentives in MDHD EV adoption. Together these results illustrate the potential to use ML-based text analysis to populate the characteristics of agent-based socio-technical models.

Electric vehicle adoption, fleet electrification,

Laboratory process control using natural language commands from a personal computer

PC software is described which provides flexible natural language process control capability with an IBM PC or compatible machine. Hardware requirements include the PC, and suitable hardware interfaces to all controlled devices. Software required includes the Microsoft Disk Operating System (MS-DOS) operating system, a PC-based FORTRAN-77 compiler, and user-written device drivers. Instructions for use of the software are given as well as a description of an application of the system.

Will, Herbert A.

Towards an Aviation Large Language Model by Fine-tuning and Evaluating Transformers

In the aviation domain, there are many applications for machine learning and artificial intelligence tools that utilize natural language. For example, there is a desire to know the commonalities in written safety reports such as voluntary post incidents reports or aerial wildfire operations reports to better understand the risks present. Another use-case is the possibility of extracting airspace procedures and constraints currently written in documents such as Letters of Agreement. These applications can benefit from the use of state-of-the-art natural language processing techniques when adapted to the language/phraseology specific to the aviation domain. This paper evaluates the viability of adaptation of NLP tools to the aviation domain by fine-tuning transformer based models using aviation data sets. In 2018, a novel language model based on neural units (also called transformers) was created and became known as “Bidirectional Encoder Representations from Transformers” or BERT. This architecture combined with large amounts of English training data and innovative semi-supervised training tasks set the standard for what would later emerge as Large Language Models. The performance of these models was further improved by hyperparameter tuning and refinement of the semi-supervised training task and resulted in “Robustly Optimized BERT Pre-training Approach through hyperparameter tuning” or RoBERTa models. These pre-trained Large Language Models proved to be useful for a wide variety of natural language processing tasks such as text classification and question answering through a process called fine-tuning. The transformer architecture with pre-trained weights served as the basis with the last few layers replaced with layers fine-tuned to perform a new task e.g., a layer that provides a label for the entire input text. This process of fine-tuning can also be used to adapt the models to new domains; e.g., BioBERT started with the pre-trained BERT model and was completed by additional fine-tuning and training on biomedical documents. Transformer-based architectures can also be used to create rich representations of text called embeddings which can serve as the input to other machine learning models. This allows simpler algorithms such as logistic regression to use context-rich representations of the text while still remaining quick to train and evaluate. In the world of aviation, there is a growing demand for natural language processing and understanding but the domain presents unique challenges. Due to the technical content (and specialized language) of most aviation documents, fine-tuning pre-trained Large Language Models to specific tasks has not met the benchmark on natural language processing tasks set by simpler models trained from scratch on the data. To address this deficiency, this paper evaluates the improvements from fine-tuning a Large Language Model on a large set of aviation documents using the original semi-supervised training tasks before performing specific natural language tasks. In fine-tuning, a domain-specific dataset is used on the original training task but with the pre-trained Large Language Model instead of starting from a random initialization. This approach allows the model to be adapted to the specific domain language without discarding the information gained from training on general English data. This paper utilized two major dataset types to train and assess the RoBERTa fine-tuning performance. The first are 7,057 Letters of Agreement which are Federal Aviation Administration (FAA) documents that formalize airspace operations across the national airspace system. They contain many examples of ‘aviation English’ using domain specific terminology and phrasing which serves as a representative basis to perform the semi-supervised fine-tuning. The second type is the 494 document classification labels to be used for evaluation. This down-stream evaluation aims to show the performance of the fine-tuned model, better understand how much data is needed for an effective fine-tuning, and how fine-tuning can be adapted for different applications in-the domain. After semi-supervised training, evaluation begins by encoding the documents for classification using the fine-tuned RoBERTa model. Then a logistic regression classifier is trained to label the document type and compared against our ground truth labels. This currently leads to a 82.8% accuracy on 10-fold cross validation showing improvement over baseline RoBERTa which achieved 81.0%. We plan to measure the improvements on additional tasks and it is expected that these improvements will lead to more robust models that can tackle the natural language processing challenges present in aviation datasets.

ATM

Enhancing Air Traffic Control Planning with Automatic Speech Recognition

The decisions made during the Federal Aviation Administration Air Traffic Control System Command Center's planning teleconferences hold significant sway over the National Airspace System. Held every two hours, these teleconferences convene air traffic managers and stakeholders from across the nation to discuss airspace conditions, weather, and constraints, leading to the formulation and adjustment of traffic management initiatives. Given the critical nature of these decisions, the need for accurate and efficient record-keeping is paramount. In recent years, the application of automatic speech recognition has gained popularity across diverse industries, including aviation. While traditional applications focus on transcribing air traffic control communication, this paper explores a unique application of automatic speech recognition by converting the audio from planning teleconferences into text transcriptions. This innovative approach addresses key challenges in the field, presenting potential benefits for quality assurance, real-time participation, and downstream natural language processing tasks. A notable breakthrough in the machine learning community, namely the transformer neural network architecture, forms the backbone of the proposed solution in this paper. The transformer architecture's role in this research represents a paradigm shift in the efficiency of automatic speech recognition models. By reducing the amount of in-domain training data required, this architecture allows for the fine-tuning of such models like Whisper, originally pretrained on vast English speech datasets. The adaptability of the transformer architecture proves invaluable in capturing the nuances of aviation terminology and specific language used in planning teleconferences. Leveraging the Whisper model as a baseline, our research details the fine-tuning and validation using a dataset comprising 20 hours of meticulously transcribed planning teleconferences. Notably, the baseline pretrained Whisper model exhibited a word error rate of 18.77%. Through the fine-tuning process, the model achieved a substantial improvement, demonstrating an impressive performance with a reduced word error rate of 6.82%. This substantial decrease in WER not only highlights the effectiveness of the transformer architecture but also emphasizes the practical advancements achieved through the application of automatic speech recognition in this specific domain. The utilization of automatic speech recognition in planning teleconferences in this work introduces several novelties. Firstly, the creation of text transcriptions offers a valuable tool for quality assurance and facilitates the efficient review of teleconferences. This is an important aspect of the proposed solution, given the time-sensitive and high-stakes nature of decisions made during these meetings. Furthermore, text-searchable transcriptions provide a streamlined approach for locating and validating critical information, potentially saving hours of manual effort in searching through audio recordings. Moreover, our research identifies a key use case for external facilities and stakeholders. In situations where attendance at the planning teleconference is not feasible, having access to text transcriptions in real-time or shortly after the teleconference ends, proves to be a time-saving and informative resource. This feature enhances collaboration and ensures that stakeholders can stay abreast of important discussions and decisions even in their absence. Despite the efficiency gains facilitated by the transformer architecture in automatic speech recognition technology, it is essential to acknowledge the human factors in data creation. Subject matter experts play a crucial role in accurately transcribing planning teleconferences due to the specificity and complexity of the information discussed. The research dataset, consisting of 20 hours of transcribed planning teleconferences, forms the foundation for fine-tuning and validating the Whisper model. The achieved word error rate of 6.82% demonstrates promising advancements, particularly in recognizing essential aviation terminology within the teleconferences. In conclusion, this paper presents a comprehensive exploration of the application of automatic speech recognition in Air Traffic Control System Command Center planning teleconferences, leveraging the transformer architecture for enhanced efficiency. The novel contributions lie in the improved accessibility of decision-making records, real-time participation opportunities for external stakeholders, and the potential for downstream natural language processing advancements. As the aviation industry continues to evolve, the integration of automatic speech recognition technologies holds the promise of revolutionizing decision-making processes and contributing to the overall safety and efficiency of air traffic management.

ATM

MIKA: Manager for Intelligent Knowledge Access Toolkit for Engineering Knowledge Discovery and Information Retrieval

Repositories of safety reports are often underutilized and only analyzed manually by trained experts, despite safety management systems requiring reports. These collections of documents contain a wealth of information from past projects and operations that could improve system safety and design. Advances in natural language processing techniques have improved information extraction and retrieval in consumer technology, biomedicine, and finance, for instance, but have not been applied to engineering documents on the same scale. To this end, the Manager for Intelligent Knowledge Access (MIKA) open-source toolkit has been developed for rapid knowledge discovery and information retrieval in safety engineering applications. The MIKA toolkit uses state-of-the-art natural language processing algorithms and allows a user to apply these methods to their own dataset. This paper describes the MIKA toolkit and its two primary capabilities, knowledge discovery and information retrieval, and demonstrates the toolkit via a case study on National Transportation Safety Board (NTSB) reports.

Machine Learning

MIKA: Manager for Intelligent Knowledge Access Toolkit for Engineering Knowledge Discovery and Information Retrieval

Repositories of safety reports are often underutilized and only analyzed manually by trained experts, despite safety management systems requiring reports. These collections of documents contain a wealth of information from past projects and operations that could improve system safety and design. Advances in natural language processing techniques have improved information extraction and retrieval in consumer technology, biomedicine, and finance, for instance, but have not been applied to engineering documents on the same scale. To this end, the Manager for Intelligent Knowledge Access (MIKA) open-source toolkit has been developed for rapid knowledge discovery and information retrieval in safety engineering applications. The MIKA toolkit uses state-of-the-art natural language processing algorithms and allows a user to apply these methods to their own dataset. This paper describes the MIKA toolkit and its two primary capabilities, knowledge discovery and information retrieval, and demonstrates the toolkit via a case study on National Transportation Safety Board (NTSB) reports.

Systems Engineering

Digitizing Named Entities Found Within Letters of Agreement

Letters of Agreement (LOAs) are text-based air traffic control documents that contain procedures and actions agreed upon by the different parties, typically two or more FAA facilities, that are subject to an agreement. The documents contain among other things generic constraints, which are explicit and implicit combinations of procedures that limit a flight’s trajectory and affects pilot actions. For example, a controller may be required, to assign a specific altitude to an aircraft crossing the boundary between two airspaces. Although LOA generic constraints directly impact the trajectory of an aircraft, they are not currently available in a digital form that can be used for (or directly ingested into automated) flight planning. Instead, the constraints are manually input into an onboard or ground based system. LOA documents are primarily stored at a controlling facility and the generic constraints are implemented by experienced air traffic controllers and pilots primarily using voice instructions. This increases the workload of the controllers, likelihood of error (e.g., due to noisy communication) and makes it impractical for implementation with unmanned aircraft. Therefore, steps must be taken to make existing constraints machine interpretable to enable e.g., automated handoffs which in turn would reduce controller workload. With recent advances in natural language processing, especially the rise in digitization of text documents (e.g., medical documents) and automated extraction of information therein, it is now possible to extract flight specific constraints from LOAs. The goal of this work is to digitize named entities through a combination of natural language processing tasks: named entity disambiguation, toponym resolution, and numeric parsing to extract general constraint components contained within LOAs, herein referred to as Entity Enhancement (EE). Starting with a small list of named entities (e.g., ARTCC, Tower, Altitude and Speed), EE can extract the named entities while simultaneously converting the string-based output into a digital format using an ensemble of processes like rule-based gazetteers and syntactic-lexical patterns. The digital format contains a diverse set of information based on the entity label in question, ranging from standardized facility names to units of measure (e.g., feet) and other numeric information. Upon validating our approach using a truth dataset, we show an overall F1-Score of 0.71 for the extraction process. Looking beyond entity enhancement, we are also working towards the goal of completely digitizing the general constraints by performing EE and fitting them into a standardized exchange model (XM) such as the Aeronautical Information Exchange Model (AIXM). This will allow for easy distribution and dissemination of LOA constraints to air users, better searchability within documents, and enable ingestion into automated flight planning. Finally, we show a preliminary version of the proposed XM architecture and demonstrate how the model can be populated from the EE output.

Stephen S. B. Clarke