Search NASASearch

SEARCH · Search NASA

Results for “knowledge discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Ask-the-expert: Active Learning Based Knowledge Discovery Using the Expert

Often the manual review of large data sets, either for purposes of labeling unlabeled instances or for classifying meaningful results from uninteresting (but statistically significant) ones is extremely resource intensive, especially in terms of subject matter expert (SME) time. Use of active learning has been shown to diminish this review time significantly. However, since active learning is an iterative process of learning a classifier based on a small number of SME-provided labels at each iteration, the lack of an enabling tool can hinder the process of adoption of these technologies in real-life, in spite of their labor-saving potential. In this demo we present ASK-the-Expert, an interactive tool that allows SMEs to review instances from a data set and provide labels within a single framework. ASK-the-Expert is powered by an active learning algorithm for training a classifier in the backend. We demonstrate this system in the context of an aviation safety application, but the tool can be adopted to work as a simple review and labeling tool as well, without the use of active learning.

software

Ask-the-Expert: Active Learning Based Knowledge Discovery Using the Expert

Often the manual review of large data sets, either for purposes of labeling unlabeled instances or for classifying meaningful results from uninteresting (but statistically significant) ones is extremely resource intensive, especially in terms of subject matter expert (SME) time. Use of active learning has been shown to diminish this review time significantly. However, since active learning is an iterative process of learning a classifier based on a small number of SME-provided labels at each iteration, the lack of an enabling tool can hinder the process of adoption of these technologies in real-life, in spite of their labor-saving potential. In this demo we present ASK-the-Expert, an interactive tool that allows SMEs to review instances from a data set and provide labels within a single framework. ASK-the-Expert is powered by an active learning algorithm for training a classifier in the back end. We demonstrate this system in the context of an aviation safety application, but the tool can be adopted to work as a simple review and labeling tool as well, without the use of active learning.

GUI

Systems Development, Data Mining, and Knowledge Discovery

The primary role of the Technical Integration Office is to provide technical solutions and services to different branches at KSC (Kennedy Space Center) and NASA program customers. The Technical Integration Office helps support KSC's operational needs by providing services such as digital connectivity, data center services, modelling and simulation tools, and communication video services. To learn the necessary technology and processes for my internship, I am working on two projects: learning C# (C Sharp programming language) with SQL and developing requirements for a PX (Communication and Public Engagement) inventory management system. To learn how to efficiently program with C#, my mentor assigned me to complete a sports informatics application that would let users discover facts and rules about various sports. The sports informatics application comes with search capabilities, report generating features, rule lists that users can modify, and diagrams for various sport strategies. To further build upon this project, I also developed a sport simulation game with the application. Once I begin more SQL-based projects, I will have the opportunity to learn how to manage databases and link SQL servers with C# programs. To develop requirements for the inventory management system, I have met with PX representatives and toured their storage facilities to see how they organize and store their items and equipment. I will also be meeting with representatives from the budget office to find out what information must be in a system budget report. The main components the system must have are customer request management, a search feature for items and equipment, report generation capabilities, and automated system warnings when item quantities reach or go below administrator-specified threshold levels. I have drafted questions and shall statements that will ultimately become part of the inventory management system requirements document.

Espinosa, Gabriel

Systems Development, Datamining and Knowledge Discovery Abstract

This summer, 2020, during my NASA internship I worked with my mentor, Ali Shaykhian, as well as a group of four other interns: Javel Gramling, Janelisse Morales, Tristian Running Crane, and Zulmarie Jiménez. Our research and projects all differ but work together to solve datamining unstructured data into an easy to read format. Turning unstructured data into something easier to follow is important for helping quickly pull data out of larger documents, so that one doesn’t have to go through multiple pages to find certain data points. By being able to structure data pulled from a document, it can be usedto collect data from mass amounts of forms and arrange it in an easy to glance at table instead of multiple forms. I chose to focus mainly on creating form templates with both Microsoft Word and Excel and getting used to the types of data that can be collected; as well as learning where both programs differed. After I was familiar with what could be gathered, I worked towards taking data collected by a Word form and importing it into an Excel spreadsheet. By being able to transfer data from a Word document to an Excel document, there is an added layer of functionality to the datamining. Moving data around between Excel sheets isn’t that complex of a process, but when you try to import from a Word document a lot of formatting and readability can be lost. The purpose of my research is to reduce that loss by using Visual Basic scripts to clean and arrange imported data.

Makayla Amber Renfro

Systems Development, Data Mining, and Knowledge Discovery

I worked as a NASA OSTEM virtually during the Fall 2020 term. Working in the IT division under my mentor Dr. Ali Shaykhian, our overall goal for the duration of this internship is to get a better understanding of the Visual Basic Language and how it can be used to make forms and collaboration with other workers to be more dynamics. I also worked with 3 other interns throughout this internship, using Microsoft Office to help each other to get a better understanding of how we approached our projects individually. Every week, my mentor Dr. Ali assigned me a task and a goal to finish by the end of each week. My overall project was figuring out a way to make email submissions and emails in general more dynamic for the NASA database. For example, instead of only using the same generic email for hundreds of different workers, a code can be used to send an individual email with more personalization such as each recipient’s name and personal info. I have also been assigned to create an email graphical user interface by the end of this internship. For these tasks to be made possible, I had to learn more about the capabilities of Microsoft Office, such as Macros in Microsoft Excel and Visual Basic in both Microsoft Excel and Word. Although it was challenging at first, I developed new programming skilled in a new language and actually made it possible to create this code along the way. Alongside doing the individual assignments, Dr. Ali also assigned up to replicate other intern’s work to get a better understanding of their approach and learn how to do it ourselves.

Janelisse Morales Gonzalez

Enabling Space Biology Knowledge Discovery Through Biospecimen Sharing: The NASA Biological Institutional Scientific Collection and Space Microbial Culture Collection

NASA and international partners have conducted experiments in space to understand the biological impacts and address hazards to health. The resulting basic and applied science is imperative to enabling humanity to venture back to the Moon and then to Mars and beyond. Sending organisms into space is a costly endeavor. All biospecimens not required by spaceflight-relevant Principal Investigators are harvested, preserved, and archived in the NASA Biological Institutional Scientific Collection (NBISC) to maximize the scientific return. The NASA Biological and Physical Sciences (BPS) Division ‘Open Science’ endeavor includes NASA Genelab, the Space Biology Program’s Biospecimen Sharing Program, Physical Sciences Informatics, the Ames Life Sciences Data Archive, and NBISC to integrate extensive data and biospecimen resources from spaceflight and/or ground-based analog experiments. NBISC biospecimens are collected and preserved according to well-established standard operating procedures to maintain scientific quality and are available on-request by the international scientific community. NBISC currently stores over 32,000 biospecimens from Shuttle, International Space Station, and ground-based space analog investigations. Tissue sharing has resulted in at least 33 publications since 2011 and 48 requests since 2016. Many requests for NBISC biospecimen come from first-time investigators who subsequently submit grants as the port-of-entry into the field of space biology. Some NBISC biospecimens have been awarded to NASA Genelab, who then generate various ‘Open Science’ -omics data sets on their platform for bioinformatics. Other NBISC biospecimen awards have led to multiple studies such as fecal microbiome analysis, DNA damage analysis using single-cell DNA sequencing, enzymatic-pathway identification involved in spaceflight muscle atrophy, and characterization of ocular morphological changes. Of note, NBISC has expanded to include a new Space Microbial Culture Collection (SMCC) for the collection, identification, documentation, long-term preservation, and distribution of space-related microbial isolates.

biospecimens

Enabling Space Biology Knowledge Discovery Through Biospecimen Sharing: The NASA Biological Institutional Scientific Collection

NASA and international partners have conducted experiments in space to understand the biological impacts and address hazards to health. The resulting basic and applied science is imperative to enabling humanity to venture back to the Moon and then to Mars and beyond. Sending organisms into space is a costly endeavor. All biospecimens not required by spaceflight-relevant Principal Investigators are harvested, preserved, and archived in the NASA Biological Institutional Scientific Collection (NBISC) to maximize the scientific return. The NASA Biological and Physical Sciences (BPS) Division has an ‘Open Science’ endeavor which includes NASA Genelab, the Space Biology Program’s Biospecimen Sharing Program, Physical Sciences Informatics, the Ames Life Sciences Data Archive, and NBISC. Its purpose is to integrate extensive data and biospecimen resources from spaceflight and/or ground-based analog experiments. NBISC biospecimens are collected and preserved according to well-established standard operating procedures to maintain scientific quality and are available on-request by the international scientific community. NBISC currently stores over 32,000 biospecimens from Shuttle, International Space Station, and ground-based space analog investigations. Tissue sharing has resulted in at least 33 publications since 2011 and 48 requests since 2016. Many requests for NBISC biospecimen come from first-time investigators who subsequently submit grants as the port-of-entry into the field of space biology. Some NBISC biospecimens have been awarded to NASA Genelab, who then generate various ‘Open Science’ -omics data sets on their platform for bioinformatics. Other NBISC biospecimen awards have led to multiple studies such as fecal microbiome analysis, DNA damage analysis using single-cell DNA sequencing, enzymatic-pathway identification involved in spaceflight muscle atrophy, and characterization of ocular morphological changes. Of note, NBISC has expanded to include a new Space Microbial Culture Collection (SMCC) for the collection, identification, documentation, long-term preservation, and distribution of space-related microbial isolates.

Ryan T. Scott

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The next era in human space exploration is rapidly approaching. The use of health countermeasures and biomonitoring systems for space missions are required to counteract space health hazards and to support life to thrive in deep space (e.g., humans, animals, plants, crops; entire ecosystems within spacecrafts/habitats/spacesuits). The development of these mission components will be highly dependent on our understanding of basic biological and health responses to myriad space hazards (ionizing radiation, altered gravitational fields, altered day-night cycles, confined isolation, hostile-closed environments, distance-duration from Earth, planetary dust-regolith, and extreme temperatures/atmospheres). The fast-growing array of space biological and mission telemetry data, which in the past was simply archived after minimal analysis, holds great potential once applied to these mission challenges if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its multi-hierarchical, multi-modal, and heterogenous nature (molecular, cellular, tissue, organ, whole organism, behavior, ecosystem, microbiome; tabular, omics, imaging, video, biospecimen, environmental physical-chemical telemetry). This session focuses on current approaches in this domain such as: making space biological data FAIR (findable, accessible, interoperable, reusable), effective data ingestion/dissemination, observational versus experimental data, Open Science collaborations, data analysis techniques, AI/ML/knowledge graph/modeling methods, and data integration/discovery tools.

open science

Open Science for Life in Space: Bioimaging, Data Sharing, and Tools for Knowledge Discovery

Precious space-flown biological experiments have both multi-omic and phenotypic data which NASA strives to make maximally open access for reuse. Currently a number of these space-relevant bioimaging datasets are being reused for AI/ML approaches. NASA Ames Life Science Data Archive and NASA GeneLab are working to make all current and future bioimaging data even more accessible and reusable. Standards for collection and curation are being implemented to enable scientists worldwide access to these data for further discovery and use.

data science

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

Molecular-omics, physiological-phenotypic-behavioral, and environmental-radiation telemetry data from spaceflight biological and health studies are increasingly being made findable, accessible, interoperable, and reusable for the scientific public. These data, as well as space science-relevant biospecimens, are available through NASA’s Open Science Data Repository (OSDR), which is the new umbrella grouping of NASA GeneLab, the Ames Life Sciences Data Archive (ALSDA), and the NASA Biological Institutional Scientific Collection (NBISC). The quality of data is underpinned by datasets having rich metadata (determined through Analysis Working Group members), processing pipelines to enable data reuse standards, and ontologies specifying terminology semantics (e.g., the Radiation Biology Ontology).

space biology

Open Science for Plants in Space: Data Sharing, Standards, and Informatics for Reuse and Knowledge Discovery

Upcoming deep space missions rely on plants and crops for crew and ecosystem health. Access to space plant data enables scientists to gain a deeper understanding of biological responses to ionizing radiation, altered gravity, low atmospheric pressure, elevated CO2, and altered photoperiods. Open Science is the practice of making research available to all, while respecting diverse cultures, fostering collaborations with equity. 2023 is the ‘Year of Open Science’, and NASA has a 5-year Transform to Open Science (TOPS) mission designed to rapidly transform the agency toward an inclusive culture of open science. NASA’s Open Science Data Repository (OSDR) developed by NASA’s Biological and Physical Sciences Division provides access to data from space-relevant biological experiments. OSDR combines two databases, GeneLab and Ames Life Sciences Data Archive (ALSDA) to maximize access to standardized ‘omics (e.g., transcriptomics, proteomics) and phenotypic data (e.g., microscopy, biomass), respectively. OSDR started in 2014 with the creation of the first space-relevant FAIR (Findable, Accessible, Interoperable, Reusable) biological ‘omics repository (GeneLab), providing detailed metadata on investigation, sample, and assay levels. Today, GeneLab hosts 62 plant datasets which have led to 5 published peer-reviewed meta-analysis publications. Most of these publications were collaboration efforts under the OSDR Analysis Working Groups (AWGs). AWGs provide great opportunities for investigators to collaborate and set new standards for space-relevant data and metadata. The AWGs are welcoming any ASPB members interested in providing plant expertise for space biology. The addition of ALSDA to OSDR is also expanding analysis capability beyond ‘omics. Now is the time to get involved as a Subject Matter Expert as we establish the framework for modern plant data archiving through the AWGs. Investigators are invited to submit their space-relevant plant datasets to OSDR and visit the site to learn about the tools OSDR has to offer (osdr.nasa.gov/bio).

FAIR

Open Science for Plants in Space: Data Sharing, Standards, and Informatics for Reuse and Knowledge Discovery

Upcoming deep space missions will rely on plants for crew and ecosystem health. Open access space biology data enables scientists to examine the biological responses of plants to ionizing radiation, altered gravity, low atmospheric pressure, elevated CO2, altered photoperiods and many other abiotic stressors. Open Science is the practice of making research available to all, while respecting diverse cultures, and fostering collaborations with equity. 2023 is the ‘Year of Open Science’, and NASA has a 5-year Transform to Open Science (TOPS) initiative designed to rapidly transform the agency toward an inclusive culture of open science. NASA’s Open Science Data Repository (OSDR) within NASA’s Biological and Physical Sciences Division provides access to data from space-relevant biological experiments. OSDR combines two databases, GeneLab and Ames Life Sciences Data Archive (ALSDA) to maximize access to standardized ‘omics (e.g., transcriptomics, proteomics) and phenotypic data (e.g., microscopy, biomass), respectively. GeneLab started in 2014 with the creation of the first space-relevant FAIR (Findable, Accessible, Interoperable, Reusable) biological ‘omics repository, providing detailed metadata on investigation, sample, and assay levels. The addition of ALSDA to OSDR expands plant data analysis capabilities across both phenotypic and ‘omics data. Today, OSDR hosts 62+ plant datasets and has enabled 58 peer-reviewed publications. Most of these publications were collaboration efforts under the OSDR Analysis Working Groups (AWGs). AWGs provide great opportunities for investigators to collaborate and set new standards for space-relevant data and metadata. The AWGs welcome any ASPB members interested in contributing plant expertise for space biology, and to serve as subject matter experts as we establish the framework for modern plant data archiving. Investigators are invited to submit their space-relevant plant datasets to OSDR and visit the site to learn about the tools OSDR has to offer (osdr.nasa.gov/bio).

FAIR

Open Science for Plants in Space: Data Sharing, Standards, and Informatics for Reuse and Knowledge Discovery

Upcoming deep space missions will rely on plants for crew and ecosystem health. Open access space biology data enables scientists to examine the biological responses of plants to ionizing radiation, altered gravity, low atmospheric pressure, elevated CO2, altered photoperiods and many other abiotic stressors. Open Science is the practice of making research available to all, while respecting diverse cultures, to foster collaborations with equity. NASA has declared 2023 as the ‘Year of Open Science’ and created a 5-year Transform to Open Science (TOPS) initiative designed to rapidly transform the agency toward an inclusive culture of open science. NASA’s Open Science Data Repository (OSDR) within the Biological and Physical Sciences Division provides access to data from space-relevant biological experiments. OSDR combines two databases, GeneLab and Ames Life Sciences Data Archive (ALSDA) to maximize access to standardized ‘omics (e.g., transcriptomics, proteomics) and phenotypic data (e.g., microscopy, biomass), respectively. GeneLab started in 2014 with the creation of the first space-relevant FAIR (Findable, Accessible, Interoperable, Reusable) biological ‘omics repository, providing detailed metadata on investigation, sample, and assay levels. The addition of ALSDA to OSDR expands plant data analysis capabilities across both phenotypic and ‘omics data. Today, OSDR hosts 62+ plant datasets and has enabled 58 peer-reviewed publications. Most of these publications were collaboration efforts under the OSDR Analysis Working Groups (AWGs). AWGs provide great opportunities for investigators to collaborate with community members and set new standards for space-relevant data and metadata. The AWGs welcome any ASGSR members interested in contributing plant expertise for space biology, and to serve as subject matter experts as we establish the framework for modern plant data archiving. Investigators are encouraged to submit their space-relevant plant datasets to OSDR and visit the site to learn about the tools OSDR has to offer (osdr.nasa.gov/bio).

FAIR

SEED: Semantic Energy Exploration and Discovery

The Bioenergy Knowledge Discovery Framework (KDF) hosts a vast repository of specialized data, yet traditional keyword-based search methods often struggle to provide direct answers, requiring significant domain expertise and manual effort to filter through raw documents. To overcome these barriers, this software introduces a semantic search engine that enables both specialists and non-specialists to query the KDF using natural language. By shifting from rigid keyword matching to intent-based retrieval, the tool automatically identifies and ranks the most relevant sources within the database. The system functions by processing natural language queries to extract the most pertinent information, delivering an AI-generated plain-language summary alongside exact supporting quotes from retrieved documents. This integrated approach provides users with immediate, evidence-based answers while eliminating the need for exhaustive manual review. By surfacing direct insights and contextual evidence, the software enhances the usability of existing KDF resources and democratizes access to complex bioenergy data. Ultimately, this semantic search solution accelerates the discovery process and supports faster, more informed decision-making across the bioenergy sector.

Pan, Meiyu (Melrose) [Oak Ridge National Laborator

Natural Language Processing Techniques for Intelligent Knowledge Management of Safety Reports

Safety, failure, and incident reports are common artifacts across various domains, including aviation and wildfire response. These reports are often mandatory to submit, resulting in the culmination of large repositories of text-based documents. Simultaneously, these reports and corresponding repositories are often only manually analyzed and queried by users via out-of-date search engines. As a consequence, we have been developing the Manager for Intelligent Knowledge Access (MIKA) toolkit, which uses natural language processing to improve information access and reuse. In this presentation, we discuss natural language processing techniques for knowledge discovery and apply these methods to a repository of aerial wildfire mishap reports. Two methods are used for knowledge discovery: topic modeling and named-entity recognition. We use topic modeling to identify hazards and perform a trend analysis to produce a data-driven risk matrix. A custom named-entity recognition model, build from fine tuning a pre-trained language model, is used to identify failure modes, failure causes, failure effects, control processes, and recommendations to aid in failure modes and effects analysis (FMEA). Throughout the presentation, we discuss and apply natural language processing techniques to better leverage the vast amount of information contained in report repositories.

Machine learning

Enhancing Dataset Discovery With Knowledge Graph Link Prediction Techniques

● In the evolving landscape of open science, the ability to navigate and discover pertinent datasets is increasingly significant. This primarily hinges on the presence of detailed metadata, delineating the dataset’s content, and potential spheres of application. ● The GES DISC datasets are characterized by science keywords to enable dataset discovery in web search interfaces. ● A problem may arise where a dataset lacks a science keyword that it otherwise should have. ● Machine learning techniques such as link prediction can be used to detect these missing science keywords by estimating the probability of new links forming between dataset and keyword nodes.

machine learning