SEARCH · Search NASA
Results for “database mining”
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Knowledge Discovery in Databases 2-Hour Session
Knowledge Discovery in Databases, or Database Mining, is the process of information extraction from very large databases.
(abstract) Modeling Protein Families and Human Genes: Hidden Markov Models and a Little Beyond
We will first give a brief overview of Hidden Markov Models (HMMs) and their use in Computational Molecular Biology. In particular, we will describe a detailed application of HMMs to the G-Protein-Coupled-Receptor Superfamily. We will also describe a number of analytical results on HMMs that can be used in discrimination tests and database mining. We will then discuss the limitations of HMMs and some new directions of research. We will conclude with some recent results on the application of HMMs to human gene modeling and parsing.
DancePartner: Python Package to Mine Multiomics Relationship Networks from Literature and Databases
A goal of multi-omics experiments is to understand how mechanistic molecular biology is altered between conditions, typically a control group and experimental groups. Oftentimes this involves studying changes in biomolecule relationships (e.g. interactions, metabolic relationships) of several types of biomolecules (e.g. proteins, lipids, metabolites). Though several databases contain relationships between biomolecules, understudied species may have little to no relationship information in databases and thus must be mined from literature. There are several challenges to literature mining, including automated full-text extraction, duplicate biomolecule term collapsing, and implementing complex machine learning tools. To make relationship extraction more accessible to the community, a python package called DancePartner was developed to allow for the extraction of relationships from literature and databases, with functions to map biomolecule synonyms to standardized identifiers and visualize and characterize the resulting multi-omics network. Here, in this study, an example dataset involving Caenorhabditis elegans is presented, where relationships are mined from 1443 publications using DancePartner. These relationships are combined with relationships from KEGG, WikiPathways, UniProt, and LipidMaps, and visualized.
Extracting Lessons of Resilience Using Machine Mining of the ASRS Database
NASA’s Aviation Safety Reporting System (ASRS) database is the world's largest repository of voluntary, confidential safety information provided by aviation's frontline personnel, including pilots, air traffic controllers, mechanics, flight attendants, dispatchers, and other members of the aviation community and the public. The database contains close to 2 million narratives, many of which describe everyday situations in which people saved the day. In these situations, people’s resilient behavior solved a problem, dealt with a malfunction, and maintained a safe operation despite a serious perturbation. To be able to extract lessons of such resilience from this large database, the use of machine learning algorithms is being explored. In this report, we describe a comparison between two such algorithms: Perilog and Word2Vec. An identical search using both programs was done on a database containing approximately 470,000 ASRS reports submitted between 1988 and 2022. The comparison reveals some of the strength and weaknesses of each algorithm as well as the challenges inherent in using such algorithms to extract lessons of resilience from the ASRS database.
Use of Business Intelligence Tools in the DSN
JPL has operated the Deep Space Network (DSN) on behalf of NASA since the 1960's. Over the last two decades, the DSN budget has generally declined in real-year dollars while the aging assets required more attention, and the missions became more complex. As a result, the DSN budget has been increasingly consumed by Operations and Maintenance (O&M), significantly reducing the funding wedge available for technology investment and for enhancing the DSN capability and capacity. Responding to this budget squeeze, the DSN launched an effort to improve the cost-efficiency of the O&M. In this paper we: elaborate on the methodology adopted to understand "where the time and money are used"-surprisingly, most of the data required for metrics development was readily available in existing databases-we have used commercial Business Intelligence (BI) tools to mine the databases and automatically extract the metrics (including trends) and distribute them weekly to interested parties; describe the DSN-specific effort to convert the intuitive understanding of "where the time is spent" into meaningful and actionable metrics that quantify use of resources, highlight candidate areas of improvement, and establish trends; and discuss the use of the BI-derived metrics-one of the most fascinating processes was the dramatic improvement in some areas of operations when the metrics were shared with the operators-the visibility of the metrics, and a self-induced competition, caused almost immediate improvement in some areas. While the near-term use of the metrics is to quantify the processes and track the improvement, these techniques will be just as useful in monitoring the process, e.g. as an input to a lean-six-sigma process.
Reference Material Kydex(registered trademark)-100 Test Data Message for Flammability Testing
The Marshall Space Flight Center (MSFC) Materials and Processes Technical Information System (MAPTIS) database contains, as an engineering resource, a large amount of material test data carefully obtained and recorded over a number of years. Flammability test data obtained using Test 1 of NASA-STD-6001 is a significant component of this database. NASA-STD-6001 recommends that Kydex 100 be used as a reference material for testing certification and for comparison between test facilities in the round-robin certification testing that occurs every 2 years. As a result of these regular activities, a large volume of test data is recorded within the MAPTIS database. The activity described in this technical report was undertaken to mine the database, recover flammability (Test 1) Kydex 100 data, and review the lessons learned from analysis of these data.
Kaona: Deep Searching and Curating Data from Aviation Safety Reporting Systems
Context: Several works in the literature have examined how safety narrative databases can be leveraged to share lessons learned. However, less attention has been given to augmenting existing processes for mining these safety reporting system databases. Aim: In this work, we introduce Kaona: An interface that weaves machine learning in existing aviation safety database mining activities. Method: We provide a use case of search, curation and newsletter writing to showcase how Kaona features build on existing processes and on its own to enhance information retrieval, curation and synthesis of narratives. Results: We created two instances of Kaona internally for evaluation, one using publicly available NASA’s ASRS narratives and another using publicly available C3RS narratives. Data ranged from 1998 to 2024. Conclusion: Our tool provides a new way to explore safety narratives, serving to re-imagine how text databases can benefit of novel information retrieval mechanisms in the era of large language models.
LIB Design Module for Grid Energy System Application
We will employ a machine learning approach with intelligent data mining and database construction to analyze enormous data repositories for identifying and extracting geographic-dependent cell design specifications from publicly accessible grid-scale energy storage usage databases in an automated way at scale.
Data Mining of Historical Human Data to Assess the Risk of Injury due to Dynamic Loads
The NASA Occupant Protection Group is charged with ensuring crewmembers are protected during all dynamic phases of spaceflight. Previous work with outside experts has led to the development of a definition of acceptable risk (DAR) for space capsule vehicles. The DAR defines allowable probability rates for various categories of injuries. An important question is how to validate these probabilities for a given vehicle. One approach is to impact test human volunteers under projected nominal landing loads. The main drawback is the large number of subject tests required to attain a reasonable level of confidence that the injury probability rates would meet those outlined in the DAR. An alternative is to mine existing databases containing human responses to impact. Testing an anthropomorphic test device (ATD) at the same human‐exposure levels could yield a range of ATD responses that would meet DAR. As one aspect of future vehicle validation, the ATD could be tested in the vehicle's seat and suit configuration at nominal landing loads and compared with the ATD responses supported by the human data set. This approach could reduce the number of human‐volunteer tests NASA would need to conduct to validate that a vehicle meets occupant protection standards. METHODS: The U.S. Air Force has recorded hundreds of human responses to frontal, lateral, and spinal impacts at many acceleration levels and pulse durations. All of this data are stored on the Collaborative Biomechanics Data Network (CBDN), which is maintained by the Wright Patterson Air Force Base (WPAFB). The test device for human occupant restraint (THOR) ATD was impact tested on WPAFB's horizontal impulse accelerator (HIA) matching human‐volunteer exposures on the HIA to 5 frontal and 3 spinal loading conditions. No human injuries occurred as a result of these impact conditions. Peak THOR response variables for neck axial tension and compression, and thoracic‐spine axial compression were collected. Maximal chest deflection was determined from motion capture video of the impact test. HIC‐ 15 and BRIC were calculated from head acceleration responses. Given the number of human subjects for each test condition a confidence interval of injury probability will be obtained. RESULTS: Results will be discussed in terms of injury‐risk probability estimates based on the human data set evaluated. Also, gaps in the data set will be identified. These gaps could be one of two types. One is areas where additional THOR testing would increase the comparable human data set, thereby improving confidence in the injury probability rate. The other is where additional human testing would assist in obtaining information on other acceleration levels or directions. DISCUSSION: The historical human data showed validity of the THOR ATD for supplemental testing. The historical human data are limited in scope, however. Further data are needed to characterize the effects of sex, age, anthropometry, and deconditioning due to spaceflight on risk of injury
FJET Database Project: Extract, Transform, and Load
The Data Mining & Knowledge Management team at Kennedy Space Center is providing data management services to the Frangible Joint Empirical Test (FJET) project at Langley Research Center (LARC). FJET is a project under the NASA Engineering and Safety Center (NESC). The purpose of FJET is to conduct an assessment of mild detonating fuse (MDF) frangible joints (FJs) for human spacecraft separation tasks in support of the NASA Commercial Crew Program. The Data Mining & Knowledge Management team has been tasked with creating and managing a database for the efficient storage and retrieval of FJET test data. This paper details the Extract, Transform, and Load (ETL) process as it is related to gathering FJET test data into a Microsoft SQL relational database, and making that data available to the data users. Lessons learned, procedures implemented, and programming code samples are discussed to help detail the learning experienced as the Data Mining & Knowledge Management team adapted to changing requirements and new technology while maintaining flexibility of design in various aspects of the data management project.
Descriptors of water aggregation
For this work, we rely on a total of 23 (cluster size, 8 structural, and 14 connectivity) descriptors to investigate structural patterns and connectivity motifs associated with water cluster aggregation. In addition to the cluster size n (number of molecules), the 8 structural descriptors can be further categorized into (i) one-body (intramolecular): covalent OH bond length (r OH ) and HOH bond angle (θ HOH ), (ii) two-body: OO distance (r OO ), OHO angle (θ OHO ), and HOOX dihedral angle ($\phi$ HOOX ), where X lies on the bisector of the HOH angle, (iii) three-body: OOO angle (θ OOO ), and (iv) many-body: modified tetrahedral order parameter (q) to account for two-, three-, four-, five-coordinated molecules (q m , m = 2, 3, 4, 5) and radius of gyration (R g ). The 14 connectivity descriptors are all many-body in nature and consist of the AD, AAD, ADD, AADD, AAAD, AAADD adjacencies [number of hydrogen bonds accepted (A) and donated (D) by each water molecule], Wiener index, Average Shortest Path Length, hydrogen bond saturation (% HB), and number of non-short-circuited three-membered cycles, four-membered cycles, five-membered cycles, six-membered cycles, and seven-membered cycles. We mined a previously reported database of 4 948 959 water cluster minima for (H 2 O) n , n = 3–25 to analyze the evolution and correlation of these descriptors for the clusters within 5 kcal/mol of the putative minima. It was found that r OH and % HB correlated strongly with cluster size n, which was identified as the strongest predictor of energetic stability. Marked changes in the adjacencies and cycle count were observed, lending insight into changes in the hydrogen bond network upon aggregation. A Principal Component Analysis (PCA) was employed to identify descriptor dependencies and group clusters into specific structural patterns across different cluster sizes. The results of this study inform our understanding of how water clusters evolve in size and what appropriate descriptors of their structural and connectivity patterns are with respect to system size, stability, and similarity. The approach described in this study is general and can be easily extended to other hydrogen-bonded systems.
Unifying Quantum Materials Modeling and Experiments: The Role of Machine Learning Interatomic Potentials
Computational experiments have emerged as a powerful complement to traditional experiments in the design of new materials. The development of machine learning (ML) and deep learning techniques, combined with database construction and data mining, has significantly enhanced traditional quantum mechanical methods. This synergy enables the rapid development of structure-property relationships. In this talk, I will discuss our recent efforts in applying Machine Learning Interatomic Potentials (MLIAPs) to accelerate materials modeling across various material classes and challenging applications where traditional methods fall short. First, I will highlight the success of MLIAPs in accurately modeling the melting behavior of complex materials. Our results demonstrate high fidelity with experimental observations and also with calculated reference melting temperatures. In the second application, I will discuss how MLIAPs are trained and applied to elucidate the interplay between segregation tendencies and surface reconstructions in CuNi alloys under oxidizing conditions. A key factor in the success of these MLIAP applications is the design of minimalistic yet flexible datasets along with a computational framework for training MLIAPs.
GeneLab: A Systems Biology Platform for Omics Analysis
NASA's GeneLab includes an open-access repository of some 200+ omics datasets generated by biological experiments relevant to spaceflight (including simulated cosmic radiation and microgravity). In order to maximize the intelligibility of these data, particularly for users with limited bioinformatics knowledge, GeneLab is now transforming the data in the repository into actual biological and physiological knowledge of the genetic and proteomic signatures found in these samples. This processed data is being derived by establishing standard data analysis workflows vetted by 114 scientists who are members of the four GeneLab Analysis Working Groups (Animal AWG, Plant AWG, Microbe AWG, Multi-Omics AWG). AWG members from institutes spanning the U.S. and four other countries participate on a voluntary basis. The AWGs meet monthly to discuss data mining, compare results and interpretations, and test forthcoming releases of the GeneLab Data Systems (GLDS). GLDS version 3.0 has been available to the general public since October 1st 2018, and has been providing a professional state-of-the-art bioinformatics platform for everyone in the space biology community to upload their data into a space biology omics data commons, to process their data with vetted standard workflows and to compare to existing analyses. The user interface for the platform is being designed to be accessible to a broad variety of users including those with limited bioinformatics experience, including high school and college students who can use it to learn about omics data analysis and space biology. As such, Genelab will constitute a powerful general public outreach capability of NASA and the Space Biology community at large. Data mining of the GeneLab database by the AWG has already started generating very interesting findings, including reports linking specific spaceflight conditions such as radiation, microgravity or carbon dioxide levels to molecular changes seen across various species. In this presentation, we will report on the current and future objectives for GeneLab, and review recent studies reported by the various AWGs relating molecular changes observed in various animal models and tissue with microgravity, radiation, circadian rhythm, hydration and carbon dioxide conditions.
GeneLab: A Systems Biology Platform for Omics Analysis: Disseminate and Reuse Data, Tools, and Samples Post-Project
NASA's GeneLab includes an open-access repository of some 200 plus omics datasets generated by biological experiments relevant to spaceflight (including simulated cosmic radiation and microgravity). In order to maximize the intelligibility of these data, particularly for users with limited bioinformatics knowledge, GeneLab is now transforming the data in the repository into actual biological and physiological knowledge of the genetic and proteomic signatures found in these samples. This processed data is being derived by establishing standard data analysis workflows vetted by 114 scientists who are members of the four GeneLab Analysis Working Groups (Animal AWG, Plant AWG, Microbe AWG, Multi-Omics AWG). AWG members from institutes spanning the U.S. and four other countries participate on a voluntary basis. The AWGs meet monthly to discuss data mining, compare results and interpretations, and test forthcoming releases of the GeneLab Data Systems (GLDS). GLDS version 3.0 has been available to the general public since October 1st 2018, and has been providing a professional state-of-the-art bioinformatics platform for everyone in the space biology community to upload their data into a space biology omics data commons, to process their data with vetted standard workflows and to compare to existing analyses. The user interface for the platform is being designed to be accessible to a broad variety of users including those with limited bioinformatics experience, including high school and college students who can use it to learn about omics data analysis and space biology. As such, Genelab will constitute a powerful general public outreach capability of NASA and the Space Biology community at large. Data mining of the GeneLab database by the AWG has already started generating very interesting findings, including reports linking specific spaceflight conditions such as radiation, microgravity or carbon dioxide levels to molecular changes seen across various species. In this presentation, we will report on the current and future objectives for GeneLab, and review recent studies reported by the various AWGs relating molecular changes observed in various animal models and tissue with microgravity, radiation, circadian rhythm, hydration and carbon dioxide conditions.
Improve Data Mining and Knowledge Discovery Through the Use of MatLab
Data mining is widely used to mine business, engineering, and scientific data. Data mining uses pattern based queries, searches, or other analyses of one or more electronic databases/datasets in order to discover or locate a predictive pattern or anomaly indicative of system failure, criminal or terrorist activity, etc. There are various algorithms, techniques and methods used to mine data; including neural networks, genetic algorithms, decision trees, nearest neighbor method, rule induction association analysis, slice and dice, segmentation, and clustering. These algorithms, techniques and methods used to detect patterns in a dataset, have been used in the development of numerous open source and commercially available products and technology for data mining. Data mining is best realized when latent information in a large quantity of data stored is discovered. No one technique solves all data mining problems; challenges are to select algorithms or methods appropriate to strengthen data/text mining and trending within given datasets. In recent years, throughout industry, academia and government agencies, thousands of data systems have been designed and tailored to serve specific engineering and business needs. Many of these systems use databases with relational algebra and structured query language to categorize and retrieve data. In these systems, data analyses are limited and require prior explicit knowledge of metadata and database relations; lacking exploratory data mining and discoveries of latent information. This presentation introduces MatLab(R) (MATrix LABoratory), an engineering and scientific data analyses tool to perform data mining. MatLab was originally intended to perform purely numerical calculations (a glorified calculator). Now, in addition to having hundreds of mathematical functions, it is a programming language with hundreds built in standard functions and numerous available toolboxes. MatLab's ease of data processing, visualization and its enormous availability of built in functionalities and toolboxes make it suitable to perform numerical computations and simulations as well as a data mining tool. Engineers and scientists can take advantage of the readily available functions/toolboxes to gain wider insight in their perspective data mining experiments.
Improve Data Mining and Knowledge Discovery through the use of MatLab
Data mining is widely used to mine business, engineering, and scientific data. Data mining uses pattern based queries, searches, or other analyses of one or more electronic databases/datasets in order to discover or locate a predictive pattern or anomaly indicative of system failure, criminal or terrorist activity, etc. There are various algorithms, techniques and methods used to mine data; including neural networks, genetic algorithms, decision trees, nearest neighbor method, rule induction association analysis, slice and dice, segmentation, and clustering. These algorithms, techniques and methods used to detect patterns in a dataset, have been used in the development of numerous open source and commercially available products and technology for data mining. Data mining is best realized when latent information in a large quantity of data stored is discovered. No one technique solves all data mining problems; challenges are to select algorithms or methods appropriate to strengthen data/text mining and trending within given datasets. In recent years, throughout industry, academia and government agencies, thousands of data systems have been designed and tailored to serve specific engineering and business needs. Many of these systems use databases with relational algebra and structured query language to categorize and retrieve data. In these systems, data analyses are limited and require prior explicit knowledge of metadata and database relations; lacking exploratory data mining and discoveries of latent information. This presentation introduces MatLab(TradeMark)(MATrix LABoratory), an engineering and scientific data analyses tool to perform data mining. MatLab was originally intended to perform purely numerical calculations (a glorified calculator). Now, in addition to having hundreds of mathematical functions, it is a programming language with hundreds built in standard functions and numerous available toolboxes. MatLab's ease of data processing, visualization and its enormous availability of built in functionalities and toolboxes make it suitable to perform numerical computations and simulations as well as a data mining tool. Engineers and scientists can take advantage of the readily available functions/toolboxes to gain wider insight in their perspective data mining experiments.
High Performance EVA Glove Collaboration: Glove Injury Data Mining Effort
Human hands play a significant role during extravehicular activity (EVA) missions and Neutral Buoyancy Lab (NBL) training events, as they are needed for translating and performing tasks in the weightless environment. It is because of this high frequency usage that hand- and arm-related injuries and discomfort are known to occur during training in the NBL and while conducting EVAs. Hand-related injuries and discomforts have been occurring to crewmembers since the days of Apollo. While there have been numerous engineering changes to the glove design, hand-related issues still persist. The primary objectives of this study are therefore to: 1) document all known EVA glove-related injuries and the circumstances of these incidents, 2) determine likely risk factors, and 3) recommend ergonomic mitigations or design strategies that can be implemented in the current and future glove designs. METHODS: The investigator team conducted an initial set of literature reviews, data mining of Lifetime Surveillance of Astronaut Health (LSAH) databases, and data distribution analyses to understand the ergonomic issues related to glove-related injuries and discomforts. The investigation focused on the injuries and discomforts of U.S. crewmembers who had worn pressurized suits and experienced glove-related incidents during the 1980 to 2010 time frame, either during training or on-orbit EVA. In addition to data mining of the LSAH database, the other objective of the study was to find complimentary sources of information such as training experience, EVA experience, suit-related sizing data, and hand-arm anthropometric data to be tied to the injury data from LSAH. RESULTS: Past studies indicated that the hand was the most frequently injured part of the body during both EVA and NBL training. This study effort thus focused primarily on crew training data in the NBL between 2002 and 2010. Of the 87 recorded training incidents, 19 occurred to women and 68 to men. While crew ages ranged from thirties to fifties, the age category most affected was in the forties range. Incident rate calculations (incidents per 100 training runs) revealed that the 2002, 2003, and 2004 time periods registered the highest reported incident rate levels (3.4, 6.1, and 4.1 respectively) when compared to the following years (all ≤ 1.0). In addition to general hand-arm discomfort being the highest reported result from training, specific types of hand injuries or symptoms included erythema, fingernail delamination, abrasions, muscle soreness/fatigue, paresthesia, bruising, blanching, and edema. Specific body locations most affected by hand injuries included the metacarpophalangeal joints, fingernails, finger crotches, fingers in general, interphalangeal joints, and fingertips. Causes of injuries reported in the LSAH data were primarily attributed to the forces that the gloved hands were exposed to due to hand intensive tasks and/or poor glove sizing. DISCUSSION: Although the age data indicate that most injuries are reported by male crewmembers in their forties, that is also the dominant gender and age range of most EVA crew therefore it is not an unexpected finding. Age and gender analysis will continue as more details on the uninjured population is accrued. While there is a reasonable mechanism to link training quantity to injury, the results were inconsistent and point to the need for a consistent method of suit-related injury screening and documentation. For instance, the high-incident rate levels for the years 2002 to 2004 could be attributed to a comprehensive medical review of crewmembers post-NBL EVA training that occurred from July 19, 2002 to January 16, 2004. Furthermore, there could have been increased awareness from an investigation at the NBL. These investigations may have temporarily increased the fidelity of reported injuries and discomforts during these dates as compared to surrounding years, when injury signs and symptom were no longer actively being investigated but rather voluntarily reported. Data mining for possible mechanistic factors continues and includes more detailed training timelines, hand anthropometry, and suit sizing information. The limited published data looking at hand-arm anthropometry correlated hand-anthropometry metrics with injuries stemming from glove design and operation. Future work will include further evaluation of body sizing and fit in relation to hand injury incidents.