Search NASASearch

SEARCH · Search NASA

Results for “text mining”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Experiences with Text Mining Large Collections of Unstructured Systems Development Artifacts at JPL

Often repositories of systems engineering artifacts at NASA's Jet Propulsion Laboratory (JPL) are so large and poorly structured that they have outgrown our capability to effectively manually process their contents to extract useful information. Sophisticated text mining methods and tools seem a quick, low-effort approach to automating our limited manual efforts. Our experiences of exploring such methods mainly in three areas including historical risk analysis, defect identification based on requirements analysis, and over-time analysis of system anomalies at JPL, have shown that obtaining useful results requires substantial unanticipated efforts - from preprocessing the data to transforming the output for practical applications. We have not observed any quick 'wins' or realized benefit from short-term effort avoidance through automation in this area. Surprisingly we have realized a number of unexpected long-term benefits from the process of applying text mining to our repositories. This paper elaborates some of these benefits and our important lessons learned from the process of preparing and applying text mining to large unstructured system artifacts at JPL aiming to benefit future TM applications in similar problem domains and also in hope for being extended to broader areas of applications.

text mining

Identification of Security related Bug Reports via Text Mining using Supervised and Unsupervised Classification

This paper is focused on automated classification of software bug reports to security and non-security related, using both supervised and unsupervised approaches. For both approaches, three types of feature vectors are used. For supervised learning, we experiment with multiple learning algorithms and training sets with different sizes. Furthermore, we propose a novel unsupervised approach based on anomaly detection. The evaluated is based on three NASA datasets. The results show that supervised classification is affected more by the learning algorithms than by feature vectors and using only 25% of the data for training provides as good results as if 90% of data are used for training. Both supervised and unsupervised learning can be used for identification of security bug reports; the former slightly outperforms the latter at the expense of labeling the testing set. In general, the performance differs across datasets, mainly due to the different amounts of security related information.

Goseva-Popstojanova, Katerina

Semantic Theme Analysis of Pilot Incident Reports

Pilots report accidents or incidents during take-off, on flight and landing to airline authorities and Federal aviation authority as well. The description of pilot reports for an incident contains technical terms related to Flight instruments and operations. Normal text mining approaches collect keywords from text documents and relate them among documents that are stored in database. Present approach will extract specific theme analysis of incident reports and semantically relate hierarchy of terms assigning weights of themes. Once the theme extraction has been performed for a given document, a unique key can be assigned to that document to cross linking the documents. Semantic linking will be used to categorize the documents based on specific rules that can help an end-user to analyze certain types of accidents. This presentation outlines the architecture of text mining for pilot incident reports for autonomous categorization of pilot incident reports using semantic theme analysis.

Thirumalainambi, Rajkumar

Enabling the Discovery of Recurring Anomalies in Aerospace System Problem Reports using High-Dimensional Clustering Techniques

This paper describes the results of a significant research and development effort conducted at NASA Ames Research Center to develop new text mining techniques to discover anomalies in free-text reports regarding system health and safety of two aerospace systems. We discuss two problems of significant importance in the aviation industry. The first problem is that of automatic anomaly discovery about an aerospace system through the analysis of tens of thousands of free-text problem reports that are written about the system. The second problem that we address is that of automatic discovery of recurring anomalies, i.e., anomalies that may be described m different ways by different authors, at varying times and under varying conditions, but that are truly about the same part of the system. The intent of recurring anomaly identification is to determine project or system weakness or high-risk issues. The discovery of recurring anomalies is a key goal in building safe, reliable, and cost-effective aerospace systems. We address the anomaly discovery problem on thousands of free-text reports using two strategies: (1) as an unsupervised learning problem where an algorithm takes free-text reports as input and automatically groups them into different bins, where each bin corresponds to a different unknown anomaly category; and (2) as a supervised learning problem where the algorithm classifies the free-text reports into one of a number of known anomaly categories. We then discuss the application of these methods to the problem of discovering recurring anomalies. In fact the special nature of recurring anomalies (very small cluster sizes) requires incorporating new methods and measures to enhance the original approach for anomaly detection. ?& pant 0-

Srivastava, Ashok, N.

Semantic Annotation of Complex Text Structures in Problem Reports

Text analysis is important for effective information retrieval from databases where the critical information is embedded in text fields. Aerospace safety depends on effective retrieval of relevant and related problem reports for the purpose of trend analysis. The complex text syntax in problem descriptions has limited statistical text mining of problem reports. The presentation describes an intelligent tagging approach that applies syntactic and then semantic analysis to overcome this problem. The tags identify types of problems and equipment that are embedded in the text descriptions. The power of these tags is illustrated in a faceted searching and browsing interface for problem report trending that combines automatically generated tags with database code fields and temporal information.

Malin, Jane T.

Topic Modeling Tool for PeTaL (Periodic Table of Life)

A topic modeling tool is constructed for the purpose of providing insights from biology to the engineer within the framework of PeTaL (Periodic Table of Life). The machine learning text mining tools–latent Dirichlet allocation (LDA) and nonnegative matrix factorization (NMF) with Kullback-Leibler (KL) divergence—are used to provide topic clusters to the user. Topic clusters are the underlying themes of a paper. For the text modeling problem, NMF-KL is the equivalent of probabilistic latent semantic analysis. Both LDA and NMF-KL are top-performing modeling tools. These tools are used to identify biological specimens relevant to the user. Various organisms solve a particular survival problem in nature differently. The topic clusters allow people without domain expertise to find these cross-topic themes in the body of documents and then branch out and examine papers whose target organisms solve the engineer’s problem. Abstracts from the Journal of Experimental Biology were used as input for the clustering tool in addition to a curated set of articles for validation. The tool is able to accept alternate input sources.

Machine learning

Using Perilog to Explore "Decision Making at NASA"

Perilog, a context intensive text mining system, is used as a discovery tool to explore topics and concerns in "Decision Making at NASA," chapter 6 of the Columbia Accident Investigation Board (CAIB) Report, Volume I. Two examples illustrate how Perilog can be used to discover highly significant safety-related information in the text without prior knowledge of the contents of the document. A third example illustrates how "if-then" statements found by Perilog can be used in logical analysis of decision making. In addition, in order to serve as a guide for future work, the technical details of preparing a PDF document for input to Perilog are included in an appendix.

McGreevy, Michael W.

Understanding the International Space Station Crew Perspective following Long-Duration Missions through Data Analytics & Visualization of Crew Feedback

The International Space Station (ISS) first became a home and research laboratory for NASA and International Partner crewmembers over 16 years ago. Each ISS mission lasts approximately 6 months and consists of three to six crewmembers. After returning to Earth, most crewmembers participate in an extensive series of 30+ debriefs intended to further understand life onboard ISS and allow crews to reflect on their experiences. Examples of debrief data collected include ISS crew feedback about sleep, dining, payload science, scheduling and time planning, health & safety, and maintenance. The Flight Crew Integration (FCI) Operational Habitability (OpsHab) team, based at Johnson Space Center (JSC), is a small group of Human Factors engineers and one stenographer that has worked collaboratively with the NASA Astronaut office and ISS Program to collect, maintain, disseminate and analyze this data. The database provides an exceptional and unique resource for understanding the "crew perspective" on long duration space missions. Data is formatted and categorized to allow for ease of search, reporting, and ultimately trending, in order to understand lessons learned, recurring issues and efficiencies gained over time. Recently, the FCI OpsHab team began collaborating with the NASA JSC Knowledge Management team to provide analytical analysis and visualization of these over 75,000 crew comments in order to better ascertain the crew's perspective on long duration spaceflight and gain insight on changes over time. In this initial phase of study, a text mining framework was used to cluster similar comments and develop measures of similarity useful for identifying relevant topics affecting crew health or performance, locating similar comments when a particular issue or item of operational interest is identified, and providing search capabilities to identify information pertinent to future spaceflight systems and processes for things like procedure development and training. In addition, the comments were scored for sentiment using a polarity scoring algorithm to identify both positive and negative comments for particular groups and clusters, allowing the team to make analytically informed decisions regarding future hardware and operating procedures. The use of polarity scoring with time series analysis was used to provide insight into how crew health and habitability is changing throughout various spaceflight increments or the station lifecycle as a whole. Finally, a visualization framework was developed to address the needs of the end users to search for and analyze comments by user, category or mission. This paper will discuss how the use of an analytical framework in conjunction with the current human interface, improved the understanding of crew perspective and shortened the time for analysis allowing for more informed decisions and rapid development of improvements. These methods are significantly optimizing the way that this valuable data can be assessed and applied to current and future spaceflight design and development. This collaboration allows the FCI OpsHab team to effectively analyze and share data in a more automated and timely fashion. Trends are no longer derived manually and can be illustrated effectively and accurately with these evolving techniques to an ever growing group of human spaceflight end users.

Bryant, Cody

Mining and beneficiation: A review of possible lunar applications

Successful exploration of Mars and outer space may require base stations strategically located on the Moon. Such bases must develop a certain self-sufficiency, particularly in the critical life support materials, fuel components, and construction materials. Technology is reviewed for the first steps in lunar resource recovery-mining and beneficiation. The topic is covered in three main categories: site selection; mining; and beneficiation. It will also include (in less detail) in-situ processes. The text described mining technology ranging from simple diggings and hauling vehicles (the strawman) to more specialized technology including underground excavation methods. The section of beneficiation emphasizes dry separation techniques and methods of sorting the ore by particle size. In-situ processes, chemical and thermal, are identified to stimulate further thinking by future researchers.

Chamberlain, Peter G.

Assessing the use of UAS-related terms in ASRS using Seed Topic Modeling

Context: The NASA Aviation Safety Reporting System (ASRS) is a voluntary confidential aviation safety reporting system. The ASRS receives reports from pilots, air traffic controllers, flight attendants and other involved in aviation operations. The reports are de-identified and coded by ASRS expert safety analysts. The de-identified reports are then disseminated to the aviation community in a number of ways including entry into an online database. Augmenting the discovery of topics of user interest in this online database would therefore be beneficial to the community it serves. Aim: We propose and execute an experiment to assess the use of seed term topic modeling using the database narratives to identify UAS reports. The use of seed term topic modeling would enable users to identify groups of related narratives associated to a topic of their interest. Method: We use a newly curated field in ASRS reports which identify UAS from non-UAS reports in combination of different set of UAS related terms to assess if seed topic modeling can be used in ASRS.

LDA

Assessing the Use of UAS-Related Terms in ASRS Using Seed Topic Modeling

Context: The NASA Aviation Safety Reporting System (ASRS) is a voluntary confidential system that disseminates reports received from personnel involved in aviation operations after de-identifying them. These reports are used by the community to improve overall aviation system safety. Aim: We propose and execute an experiment to assess the use of seed term topic modeling over the database narratives to identify Unmanned Aircraft System (UAS) reports. The use of seed term topic modeling enables users to identify groups of conceptually similar narratives associated to a topic of their interest. Method: We use a collection of narratives, expert-selected words, and report metadata that separates UAS from non-UAS reports to assess if seed topic modeling can be used to improve ASRS searches. Results: For simpler queries, seed topic search observes a higher recall and lower precision than the existing DBOL (DataBase OnLine) search in operation. However, the best results are obtained when seed topic search is used as a search suggestion system to be executed on the DBOL. Conclusion: Utilizing a combination of both the existing method and the proposed method, users can expand their search vocabulary about subjects of interest while improving the quality of results.

Text Mining

Assessing the Use of UAS-Related Terms in ASRS using Seeds for Topic Modeling

Context: The NASA Aviation Safety Reporting System (ASRS) is a voluntary confidential system that disseminates reports received from personnel involved in aviation operations after de-identifying them. These reports are used by the community to improve overall aviation system safety. Aim: We propose and execute an experiment to assess the use of seed term topic modeling over the database narratives to identify Unmanned Aircraft System (UAS) reports. The use of seed term topic modeling enables users to identify groups of conceptually similar narratives associated to a topic of their interest. Method: We use a collection of narratives, expert-selected words, and report metadata that separates UAS from non-UAS reports to assess if seed topic modeling can be used to improve ASRS searches. Results: For simpler queries, seed topic search observes a higher recall and lower precision than the existing DBOL (DataBase OnLine) search in operation. However, the best results are obtained when seed topic search is used as a search suggestion system to be executed on the DBOL. Conclusion: Utilizing a combination of both the existing method and the proposed method, users can expand their search vocabulary about subjects of interest while improving the quality of results.

LDA

An Approach to Identifying Aspects of Positive Pilot Behavior within the Aviation Safety Reporting System

The National Airspace System (NAS) is constantly evolving as air traffic continues to ramp up to pre-pandemic numbers and projected to grow to unprecedented levels in the coming years. As well as increasing demand to the current system, emerging operations such as Unmanned Autonomous Systems are also expected to add to complexity in the airspace. To address these issues, the industry and government agencies supporting the NAS will need to rely upon additional automation and new technologies to address future operational requirements, while continuing to be a world-leading safe transportation system. As these new technologies are implemented, the system continues to rely on human pilots and controllers in the loop to monitor the system and intervene in situations the automation cannot handle. The goal of proactively addressing safety is of foremost concern to ensure passenger confidence. The industry has implemented various Safety Monitoring Systems to identify safety risks and proactively address them before they result in a serious incident or accident. One such program is the Aviation Safety Reporting System (ASRS). ASRS is a long-established system where pilots and controllers voluntarily and anonymously report safety incidents they experienced and observed during line operations by providing rich text narratives describing the events, the environment, and conditions leading to the safety event of concern. These narratives provide insight and context around events of interest and can be used to identify emerging problems. They can trigger investigations within Flight Operational Quality Assurance or Flight Data Monitoring programs. However, this process typically focuses on the adverse events and the unsafe aspects of the operations surrounding the reported or detected events. This perspective of investigating factors that went wrong around an adverse event is commonly referred to as Safety I. Alternatively, characterizing successful actions that operators perform every day under varying conditions that keep the system within safe operating bounds is a concept referred to as Safety II. The benefit of the Safety II view is that the scope is much larger than that of Safety I since a vast majority of the operations result in successful flights. Many of the successful techniques used to manage operational threats are not documented in standard operating procedures or taught during training. They are typically acquired over time by working with experienced pilots during line operations or in many cases after experiencing a problem for the first time and reacting to it in situ, drawing from years of experience to manage the threat. In an attempt to quantify these positive actions, we are proposing an approach to extracting key behaviors within ASRS reports that can support the Safety II concept. Our analysis assumes that ASRS reports contain some descriptions of corrective actions that operators performed to prevent a situation from leading to an accident. Leveraging recent advances in Natural Language Process modeling, we have developed an approach to extract positive sentiment from reports, embed these positive statements in a vector space where they can be numerically analyzed, and clustering these statements into similar contextual categories. From these contextualized categories we can attempt to summarized and distilled aspects of the positive behavior. The goal is to identify categories of behavior that describe consistent operator techniques that supports the Safety II concept. With this information, airlines may enable learning from these positive actions, or address procedures that need to be changed to avoid having pilots implement a workaround. These insights can provide a lens into what is “going right” in the operations that may otherwise not be known widely within the community. It is envisioned that this approach can be extended to other narrative programs such as Line Operation Safety Audit or Learning Improvement Team reports where similar observed behavior can be analyzed to extract positive actions and inform the overall operations.

NLP

Using a Knowledge Graph to Discover Earth Science Information

Knowledge graphs link key entities within a specific domain to other entities via relationships. Researchers are able to mine these relationships from numerous sources to infer new knowledge. Text extraction from peer-reviewed papers and scientific reports are untapped resources that can be leveraged by knowledge graphs to accelerate scientific discovery.

Freitag, Brian

Data Mining SIAM Presentation

This viewgraph document describes the data mining system developed at NASA Ames. Many NASA programs have large numbers (and types) of problem reports.These free text reports are written by a number of different people, thus the emphasis and wording vary considerably With so much data to sift through, analysts (subject experts) need help identifying any possible safety issues or concerns and help them confirm that they haven't missed important problems. Unsupervised clustering is the initial step to accomplish this; We think we can go much farther, specifically, identify possible recurring anomalies. Recurring anomalies may be indicators of larger systemic problems. The requirement to identify these anomalies has led to the development of Recurring Anomaly Discovery System (ReADS).

Srivastava, Ashok

Searching for 'Unknown Unknowns'

The NASA Engineering and Safety Center (NESC) was established to improve safety through engineering excellence within NASA programs and projects. As part of this goal, methods are being investigated to enable the NESC to become proactive in identifying areas that may be precursors to future problems. The goal is to find unknown indicators of future problems, not to duplicate the program-specific trending efforts. The data that is critical for detecting these indicators exist in a plethora of dissimilar non-conformance and other databases (without a common format or taxonomy). In fact, much of the data is unstructured text. However, one common database is not required if the right standards and electronic tools are employed. Electronic data mining is a particularly promising tool for this effort into unsupervised learning of common factors. This work in progress began with a systematic evaluation of available data mining software packages, based on documented decision techniques using weighted criteria. The four packages, which were perceived to have the most promise for NASA applications, are being benchmarked and evaluated by independent contractors. Preliminary recommendations for "best practices" in data mining and trending are provided. Final results and recommendations should be available in the Fall 2005. This critical first step in identifying "unknown unknowns" before they become problems is applicable to any set of engineering or programmatic data.

Parsons, Vickie S.

Exploration Clinical Decision Support System: Medical Data Architecture

The Exploration Clinical Decision Support (ECDS) System project is intended to enhance the Exploration Medical Capability (ExMC) Element for extended duration, deep-space mission planning in HRP. A major development guideline is the Risk of "Adverse Health Outcomes & Decrements in Performance due to Limitations of In-flight Medical Conditions". ECDS attempts to mitigate that Risk by providing crew-specific health information, actionable insight, crew guidance and advice based on computational algorithmic analysis. The availability of inflight health diagnostic computational methods has been identified as an essential capability for human exploration missions. Inflight electronic health data sources are often heterogeneous, and thus may be isolated or not examined as an aggregate whole. The ECDS System objective provides both a data architecture that collects and manages disparate health data, and an active knowledge system that analyzes health evidence to deliver case-specific advice. A single, cohesive space-ready decision support capability that considers all exploration clinical measurements is not commercially available at present. Hence, this Task is a newly coordinated development effort by which ECDS and its supporting data infrastructure will demonstrate the feasibility of intelligent data mining and predictive modeling as a biomedical diagnostic support mechanism on manned exploration missions. The initial step towards ground and flight demonstrations has been the research and development of both image and clinical text-based computer-aided patient diagnosis. Human anatomical images displaying abnormal/pathological features have been annotated using controlled terminology templates, marked-up, and then stored in compliance with the AIM standard. These images have been filtered and disease characterized based on machine learning of semantic and quantitative feature vectors. The next phase will evaluate disease treatment response via quantitative linear dimension biomarkers that enable image content-based retrieval and criteria assessment. In addition, a data mining engine (DME) is applied to cross-sectional adult surveys for predicting occurrence of renal calculi, ranked by statistical significance of demographics and specific food ingestion. In addition to this precursor space flight algorithm training, the DME will utilize a feature-engineering capability for unstructured clinical text classification health discovery. The ECDS backbone is a proposed multi-tier modular architecture providing data messaging protocols, storage, management and real-time patient data access. Technology demonstrations and success metrics will be finalized in FY16.

Biomedical support