Search NASA⌕ Search

SEARCH · Search NASA

Results for “database mining”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

(abstract) Modeling Protein Families and Human Genes: Hidden Markov Models and a Little Beyond

We will first give a brief overview of Hidden Markov Models (HMMs) and their use in Computational Molecular Biology. In particular, we will describe a detailed application of HMMs to the G-Protein-Coupled-Receptor Superfamily. We will also describe a number of analytical results on HMMs that can be used in discrimination tests and database mining. We will then discuss the limitations of HMMs and some new directions of research. We will conclude with some recent results on the application of HMMs to human gene modeling and parsing.

Hidden Markov Models HMMs proteins computational m↗

Extracting Lessons of Resilience Using Machine Mining of the ASRS Database

NASA’s Aviation Safety Reporting System (ASRS) database is the world's largest repository of voluntary, confidential safety information provided by aviation's frontline personnel, including pilots, air traffic controllers, mechanics, flight attendants, dispatchers, and other members of the aviation community and the public. The database contains close to 2 million narratives, many of which describe everyday situations in which people saved the day. In these situations, people’s resilient behavior solved a problem, dealt with a malfunction, and maintained a safe operation despite a serious perturbation. To be able to extract lessons of such resilience from this large database, the use of machine learning algorithms is being explored. In this report, we describe a comparison between two such algorithms: Perilog and Word2Vec. An identical search using both programs was done on a database containing approximately 470,000 ASRS reports submitted between 1988 and 2022. The comparison reveals some of the strength and weaknesses of each algorithm as well as the challenges inherent in using such algorithms to extract lessons of resilience from the ASRS database.

resilience↗

Use of Business Intelligence Tools in the DSN

JPL has operated the Deep Space Network (DSN) on behalf of NASA since the 1960's. Over the last two decades, the DSN budget has generally declined in real-year dollars while the aging assets required more attention, and the missions became more complex. As a result, the DSN budget has been increasingly consumed by Operations and Maintenance (O&M), significantly reducing the funding wedge available for technology investment and for enhancing the DSN capability and capacity. Responding to this budget squeeze, the DSN launched an effort to improve the cost-efficiency of the O&M. In this paper we: elaborate on the methodology adopted to understand "where the time and money are used"-surprisingly, most of the data required for metrics development was readily available in existing databases-we have used commercial Business Intelligence (BI) tools to mine the databases and automatically extract the metrics (including trends) and distribute them weekly to interested parties; describe the DSN-specific effort to convert the intuitive understanding of "where the time is spent" into meaningful and actionable metrics that quantify use of resources, highlight candidate areas of improvement, and establish trends; and discuss the use of the BI-derived metrics-one of the most fascinating processes was the dramatic improvement in some areas of operations when the metrics were shared with the operators-the visibility of the metrics, and a self-induced competition, caused almost immediate improvement in some areas. While the near-term use of the metrics is to quantify the processes and track the improvement, these techniques will be just as useful in monitoring the process, e.g. as an input to a lean-six-sigma process.

Metrics↗

Reference Material Kydex(registered trademark)-100 Test Data Message for Flammability Testing

The Marshall Space Flight Center (MSFC) Materials and Processes Technical Information System (MAPTIS) database contains, as an engineering resource, a large amount of material test data carefully obtained and recorded over a number of years. Flammability test data obtained using Test 1 of NASA-STD-6001 is a significant component of this database. NASA-STD-6001 recommends that Kydex 100 be used as a reference material for testing certification and for comparison between test facilities in the round-robin certification testing that occurs every 2 years. As a result of these regular activities, a large volume of test data is recorded within the MAPTIS database. The activity described in this technical report was undertaken to mine the database, recover flammability (Test 1) Kydex 100 data, and review the lessons learned from analysis of these data.

Engel, Carl D.↗

Kaona: Deep Searching and Curating Data from Aviation Safety Reporting Systems

Context: Several works in the literature have examined how safety narrative databases can be leveraged to share lessons learned. However, less attention has been given to augmenting existing processes for mining these safety reporting system databases. Aim: In this work, we introduce Kaona: An interface that weaves machine learning in existing aviation safety database mining activities. Method: We provide a use case of search, curation and newsletter writing to showcase how Kaona features build on existing processes and on its own to enhance information retrieval, curation and synthesis of narratives. Results: We created two instances of Kaona internally for evaluation, one using publicly available NASA’s ASRS narratives and another using publicly available C3RS narratives. Data ranged from 1998 to 2024. Conclusion: Our tool provides a new way to explore safety narratives, serving to re-imagine how text databases can benefit of novel information retrieval mechanisms in the era of large language models.

ASRS↗

Data Mining of Historical Human Data to Assess the Risk of Injury due to Dynamic Loads

The NASA Occupant Protection Group is charged with ensuring crewmembers are protected during all dynamic phases of spaceflight. Previous work with outside experts has led to the development of a definition of acceptable risk (DAR) for space capsule vehicles. The DAR defines allowable probability rates for various categories of injuries. An important question is how to validate these probabilities for a given vehicle. One approach is to impact test human volunteers under projected nominal landing loads. The main drawback is the large number of subject tests required to attain a reasonable level of confidence that the injury probability rates would meet those outlined in the DAR. An alternative is to mine existing databases containing human responses to impact. Testing an anthropomorphic test device (ATD) at the same human‐exposure levels could yield a range of ATD responses that would meet DAR. As one aspect of future vehicle validation, the ATD could be tested in the vehicle's seat and suit configuration at nominal landing loads and compared with the ATD responses supported by the human data set. This approach could reduce the number of human‐volunteer tests NASA would need to conduct to validate that a vehicle meets occupant protection standards. METHODS: The U.S. Air Force has recorded hundreds of human responses to frontal, lateral, and spinal impacts at many acceleration levels and pulse durations. All of this data are stored on the Collaborative Biomechanics Data Network (CBDN), which is maintained by the Wright Patterson Air Force Base (WPAFB). The test device for human occupant restraint (THOR) ATD was impact tested on WPAFB's horizontal impulse accelerator (HIA) matching human‐volunteer exposures on the HIA to 5 frontal and 3 spinal loading conditions. No human injuries occurred as a result of these impact conditions. Peak THOR response variables for neck axial tension and compression, and thoracic‐spine axial compression were collected. Maximal chest deflection was determined from motion capture video of the impact test. HIC‐ 15 and BRIC were calculated from head acceleration responses. Given the number of human subjects for each test condition a confidence interval of injury probability will be obtained. RESULTS: Results will be discussed in terms of injury‐risk probability estimates based on the human data set evaluated. Also, gaps in the data set will be identified. These gaps could be one of two types. One is areas where additional THOR testing would increase the comparable human data set, thereby improving confidence in the injury probability rate. The other is where additional human testing would assist in obtaining information on other acceleration levels or directions. DISCUSSION: The historical human data showed validity of the THOR ATD for supplemental testing. The historical human data are limited in scope, however. Further data are needed to characterize the effects of sex, age, anthropometry, and deconditioning due to spaceflight on risk of injury

Wells, Jesica↗

FJET Database Project: Extract, Transform, and Load

The Data Mining & Knowledge Management team at Kennedy Space Center is providing data management services to the Frangible Joint Empirical Test (FJET) project at Langley Research Center (LARC). FJET is a project under the NASA Engineering and Safety Center (NESC). The purpose of FJET is to conduct an assessment of mild detonating fuse (MDF) frangible joints (FJs) for human spacecraft separation tasks in support of the NASA Commercial Crew Program. The Data Mining & Knowledge Management team has been tasked with creating and managing a database for the efficient storage and retrieval of FJET test data. This paper details the Extract, Transform, and Load (ETL) process as it is related to gathering FJET test data into a Microsoft SQL relational database, and making that data available to the data users. Lessons learned, procedures implemented, and programming code samples are discussed to help detail the learning experienced as the Data Mining & Knowledge Management team adapted to changing requirements and new technology while maintaining flexibility of design in various aspects of the data management project.

excel vba↗

GeneLab: A Systems Biology Platform for Omics Analysis

NASA's GeneLab includes an open-access repository of some 200+ omics datasets generated by biological experiments relevant to spaceflight (including simulated cosmic radiation and microgravity). In order to maximize the intelligibility of these data, particularly for users with limited bioinformatics knowledge, GeneLab is now transforming the data in the repository into actual biological and physiological knowledge of the genetic and proteomic signatures found in these samples. This processed data is being derived by establishing standard data analysis workflows vetted by 114 scientists who are members of the four GeneLab Analysis Working Groups (Animal AWG, Plant AWG, Microbe AWG, Multi-Omics AWG). AWG members from institutes spanning the U.S. and four other countries participate on a voluntary basis. The AWGs meet monthly to discuss data mining, compare results and interpretations, and test forthcoming releases of the GeneLab Data Systems (GLDS). GLDS version 3.0 has been available to the general public since October 1st 2018, and has been providing a professional state-of-the-art bioinformatics platform for everyone in the space biology community to upload their data into a space biology omics data commons, to process their data with vetted standard workflows and to compare to existing analyses. The user interface for the platform is being designed to be accessible to a broad variety of users including those with limited bioinformatics experience, including high school and college students who can use it to learn about omics data analysis and space biology. As such, Genelab will constitute a powerful general public outreach capability of NASA and the Space Biology community at large. Data mining of the GeneLab database by the AWG has already started generating very interesting findings, including reports linking specific spaceflight conditions such as radiation, microgravity or carbon dioxide levels to molecular changes seen across various species. In this presentation, we will report on the current and future objectives for GeneLab, and review recent studies reported by the various AWGs relating molecular changes observed in various animal models and tissue with microgravity, radiation, circadian rhythm, hydration and carbon dioxide conditions.

Omics↗

GeneLab: A Systems Biology Platform for Omics Analysis: Disseminate and Reuse Data, Tools, and Samples Post-Project

NASA's GeneLab includes an open-access repository of some 200 plus omics datasets generated by biological experiments relevant to spaceflight (including simulated cosmic radiation and microgravity). In order to maximize the intelligibility of these data, particularly for users with limited bioinformatics knowledge, GeneLab is now transforming the data in the repository into actual biological and physiological knowledge of the genetic and proteomic signatures found in these samples. This processed data is being derived by establishing standard data analysis workflows vetted by 114 scientists who are members of the four GeneLab Analysis Working Groups (Animal AWG, Plant AWG, Microbe AWG, Multi-Omics AWG). AWG members from institutes spanning the U.S. and four other countries participate on a voluntary basis. The AWGs meet monthly to discuss data mining, compare results and interpretations, and test forthcoming releases of the GeneLab Data Systems (GLDS). GLDS version 3.0 has been available to the general public since October 1st 2018, and has been providing a professional state-of-the-art bioinformatics platform for everyone in the space biology community to upload their data into a space biology omics data commons, to process their data with vetted standard workflows and to compare to existing analyses. The user interface for the platform is being designed to be accessible to a broad variety of users including those with limited bioinformatics experience, including high school and college students who can use it to learn about omics data analysis and space biology. As such, Genelab will constitute a powerful general public outreach capability of NASA and the Space Biology community at large. Data mining of the GeneLab database by the AWG has already started generating very interesting findings, including reports linking specific spaceflight conditions such as radiation, microgravity or carbon dioxide levels to molecular changes seen across various species. In this presentation, we will report on the current and future objectives for GeneLab, and review recent studies reported by the various AWGs relating molecular changes observed in various animal models and tissue with microgravity, radiation, circadian rhythm, hydration and carbon dioxide conditions.

Omics↗

Improve Data Mining and Knowledge Discovery Through the Use of MatLab

Data mining is widely used to mine business, engineering, and scientific data. Data mining uses pattern based queries, searches, or other analyses of one or more electronic databases/datasets in order to discover or locate a predictive pattern or anomaly indicative of system failure, criminal or terrorist activity, etc. There are various algorithms, techniques and methods used to mine data; including neural networks, genetic algorithms, decision trees, nearest neighbor method, rule induction association analysis, slice and dice, segmentation, and clustering. These algorithms, techniques and methods used to detect patterns in a dataset, have been used in the development of numerous open source and commercially available products and technology for data mining. Data mining is best realized when latent information in a large quantity of data stored is discovered. No one technique solves all data mining problems; challenges are to select algorithms or methods appropriate to strengthen data/text mining and trending within given datasets. In recent years, throughout industry, academia and government agencies, thousands of data systems have been designed and tailored to serve specific engineering and business needs. Many of these systems use databases with relational algebra and structured query language to categorize and retrieve data. In these systems, data analyses are limited and require prior explicit knowledge of metadata and database relations; lacking exploratory data mining and discoveries of latent information. This presentation introduces MatLab(R) (MATrix LABoratory), an engineering and scientific data analyses tool to perform data mining. MatLab was originally intended to perform purely numerical calculations (a glorified calculator). Now, in addition to having hundreds of mathematical functions, it is a programming language with hundreds built in standard functions and numerous available toolboxes. MatLab's ease of data processing, visualization and its enormous availability of built in functionalities and toolboxes make it suitable to perform numerical computations and simulations as well as a data mining tool. Engineers and scientists can take advantage of the readily available functions/toolboxes to gain wider insight in their perspective data mining experiments.

Shaykhian, Gholam Ali↗

Improve Data Mining and Knowledge Discovery through the use of MatLab

Data mining is widely used to mine business, engineering, and scientific data. Data mining uses pattern based queries, searches, or other analyses of one or more electronic databases/datasets in order to discover or locate a predictive pattern or anomaly indicative of system failure, criminal or terrorist activity, etc. There are various algorithms, techniques and methods used to mine data; including neural networks, genetic algorithms, decision trees, nearest neighbor method, rule induction association analysis, slice and dice, segmentation, and clustering. These algorithms, techniques and methods used to detect patterns in a dataset, have been used in the development of numerous open source and commercially available products and technology for data mining. Data mining is best realized when latent information in a large quantity of data stored is discovered. No one technique solves all data mining problems; challenges are to select algorithms or methods appropriate to strengthen data/text mining and trending within given datasets. In recent years, throughout industry, academia and government agencies, thousands of data systems have been designed and tailored to serve specific engineering and business needs. Many of these systems use databases with relational algebra and structured query language to categorize and retrieve data. In these systems, data analyses are limited and require prior explicit knowledge of metadata and database relations; lacking exploratory data mining and discoveries of latent information. This presentation introduces MatLab(TradeMark)(MATrix LABoratory), an engineering and scientific data analyses tool to perform data mining. MatLab was originally intended to perform purely numerical calculations (a glorified calculator). Now, in addition to having hundreds of mathematical functions, it is a programming language with hundreds built in standard functions and numerous available toolboxes. MatLab's ease of data processing, visualization and its enormous availability of built in functionalities and toolboxes make it suitable to perform numerical computations and simulations as well as a data mining tool. Engineers and scientists can take advantage of the readily available functions/toolboxes to gain wider insight in their perspective data mining experiments.

Shaykahian, Gholan Ali↗

High Performance EVA Glove Collaboration: Glove Injury Data Mining Effort

Human hands play a significant role during extravehicular activity (EVA) missions and Neutral Buoyancy Lab (NBL) training events, as they are needed for translating and performing tasks in the weightless environment. It is because of this high frequency usage that hand- and arm-related injuries and discomfort are known to occur during training in the NBL and while conducting EVAs. Hand-related injuries and discomforts have been occurring to crewmembers since the days of Apollo. While there have been numerous engineering changes to the glove design, hand-related issues still persist. The primary objectives of this study are therefore to: 1) document all known EVA glove-related injuries and the circumstances of these incidents, 2) determine likely risk factors, and 3) recommend ergonomic mitigations or design strategies that can be implemented in the current and future glove designs. METHODS: The investigator team conducted an initial set of literature reviews, data mining of Lifetime Surveillance of Astronaut Health (LSAH) databases, and data distribution analyses to understand the ergonomic issues related to glove-related injuries and discomforts. The investigation focused on the injuries and discomforts of U.S. crewmembers who had worn pressurized suits and experienced glove-related incidents during the 1980 to 2010 time frame, either during training or on-orbit EVA. In addition to data mining of the LSAH database, the other objective of the study was to find complimentary sources of information such as training experience, EVA experience, suit-related sizing data, and hand-arm anthropometric data to be tied to the injury data from LSAH. RESULTS: Past studies indicated that the hand was the most frequently injured part of the body during both EVA and NBL training. This study effort thus focused primarily on crew training data in the NBL between 2002 and 2010. Of the 87 recorded training incidents, 19 occurred to women and 68 to men. While crew ages ranged from thirties to fifties, the age category most affected was in the forties range. Incident rate calculations (incidents per 100 training runs) revealed that the 2002, 2003, and 2004 time periods registered the highest reported incident rate levels (3.4, 6.1, and 4.1 respectively) when compared to the following years (all ≤ 1.0). In addition to general hand-arm discomfort being the highest reported result from training, specific types of hand injuries or symptoms included erythema, fingernail delamination, abrasions, muscle soreness/fatigue, paresthesia, bruising, blanching, and edema. Specific body locations most affected by hand injuries included the metacarpophalangeal joints, fingernails, finger crotches, fingers in general, interphalangeal joints, and fingertips. Causes of injuries reported in the LSAH data were primarily attributed to the forces that the gloved hands were exposed to due to hand intensive tasks and/or poor glove sizing. DISCUSSION: Although the age data indicate that most injuries are reported by male crewmembers in their forties, that is also the dominant gender and age range of most EVA crew therefore it is not an unexpected finding. Age and gender analysis will continue as more details on the uninjured population is accrued. While there is a reasonable mechanism to link training quantity to injury, the results were inconsistent and point to the need for a consistent method of suit-related injury screening and documentation. For instance, the high-incident rate levels for the years 2002 to 2004 could be attributed to a comprehensive medical review of crewmembers post-NBL EVA training that occurred from July 19, 2002 to January 16, 2004. Furthermore, there could have been increased awareness from an investigation at the NBL. These investigations may have temporarily increased the fidelity of reported injuries and discomforts during these dates as compared to surrounding years, when injury signs and symptom were no longer actively being investigated but rather voluntarily reported. Data mining for possible mechanistic factors continues and includes more detailed training timelines, hand anthropometry, and suit sizing information. The limited published data looking at hand-arm anthropometry correlated hand-anthropometry metrics with injuries stemming from glove design and operation. Future work will include further evaluation of body sizing and fit in relation to hand injury incidents.

Reid, C. R.↗

Formal Provenance Representation of the Data and Information Supporting the National Climate Assessment

The Global Change Information System (GCIS) provides a framework for the formal representation of structured metadata about data and information about global change. The pilot deployment of the system supports the National Climate Assessment (NCA), a major report of the U.S. Global Change Research Program (USGCRP). A consumer of that report can use the system to browse and explore that supporting information. Additionally, capturing that information into a structured data model and presenting it in standard formats through well defined open inter- faces, including query interfaces suitable for data mining and linking with other databases, the information becomes valuable for other analytic uses as well.

Provenance↗

The Large Footprint of Small-scale Artisanal Gold Mining in Ghana

Gold mining has played a significant role in Ghana's economy for centuries. Regulation of this industry has varied over time and while industrial mining is prevalent in the country, the expansion of artisanal mining, or Galamsey has escalated in recent years. Many of these artisanal mines are not only harmful to human health due to the use of Mercury (Hg) in the amalgamation process, but also leave a significant footprint on terrestrial ecosystems, degrading and destroying forested ecosystems in the region. In this study, the Landsat image archive available through Google Earth Engine was used to quantify the total footprint of vegetation loss due to artisanal goldmines in Ghana from 2005 to 2019 and understand how conversion of forested regions to mining has changed over a decadal period from 2007 to 2017. A combination of machine learning and change detection algorithms were used to calculate different land cover conversions and the timing of conversion annually. Within the study area of southwestern Ghana, our results indicate that approximately 47,000 ha (⨦2218 ha) of vegetation were converted to mining at an average rate of ~2600 ha yr−1. The results indicate that a high percentage(~50%) of this mining occurred between 2014 and 2017. Around 700 ha of this mining occurred within protected areas as mapped by the World Database of Protected Areas. In addition to deforestation, increased artisanal mining activity in recent years has the potential to affect human health, access to drinking water resources and food security. This work expands upon limited research into the spatial footprint of Galamseyin Ghana, complements mapping efforts by local geographers, and will support efforts by the government of Ghana to monitor deforestation caused by artisanal mining.

Abigail Barenblitt↗

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, whole organism, behavior; tabular, imagery). Open Science is the concept that the more people have access to scientifically curated data, the more knowledge will be gained. This led NASA to start the development of GeneLab in 2015. GeneLab houses spaceflight and space-analog multi-omics datasets from plant, rodent, small animal, and microbial experiments. The success and knowledge gained from GeneLab led to a new alliance of NASA “Open Science Data Repositories” (OSDR), which include the Ames Life Sciences Data Archive (ALSDA) and the NASA Biological Institutional Scientific Collection (NBISC). Both are adopting the GeneLab data system, so data are more findable, accessible, interoperable, and reusable (FAIR). OSDR systems provide users the ability to upload, download, search, share, analyze, and visualize. Open Science also needs strong confidence in the data, which is gained through building science communities. With ~400 current members, GeneLab and ALSDA formed Analysis Working Groups (AWGs) to provide feedback on processing pipelines, metadata curation standards (for ‘omics and phenotypic-physiological-behavioral assays), and to collaborate in effectively reusing data. The AWG also led to the development of the Radiation Biology Ontology (RBO), ensuring radiation metadata are efficiently captured, connected, and interoperable. Feedback from the AWG provided design input toward the new single point-of-entry data submission portal for all investigators to submit, curate, and share their research data. Space biological data is now maximally open access, collected-curated with rich metadata, and formatted for interoperability to enable systems biology, meta-analysis, knowledge graphs, machine learning, modeling, and other reuse approaches. With potential for further federation of OSDR for data mining with traditional biological and medical databases (NIH, NCI, EBI, etc.), a new era for space biology has begun to support the knowledge discovery necessary for Lunar and Martian missions.

Ryan T Scott↗

Examination of Lunar Regolith Simulants By SEM-EDS and Imaging Raman Spectroscopy

The Artemis series lunar missions will include sample returns from the lunar south pole. Lunar regolith simulants (RS) generated in the lab provide opportunities to compare two surface science techniques: scanning electron microscopy (SEM-EDS) and Raman induced surface spectroscopy. The results will also be applicable for supporting future commercial lunar payload services (CLPS) and Artemis surface. Surface characterization contributes to continued development of lunar regolith studies and adds to various regolith databases. In characterizing various lunar regolith simulants, part of the aim should be to standardize techniques and utilize anticipated methods available for astromaterials studied during, or returned from, upcoming missions, particularly samples collected from the Moon’s south pole and permanently shadowed regions (PSRs).. An initial survey of available simulants has been started with Raman scanning process for particle counting the results of which will enrich the NASA-JSC Simulant Development Lab (SDL) simulant properties database and the Colorado School of Mines Planetary Simulant Data Base. An SEM-EDS dataset of raw regolith simulant materials are collected.

Raman Microscopy↗