Search NASASearch

SEARCH · Search NASA

Results for “knowledge discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

An Approach to Data Center-Based KDD of Remote Sensing Datasets

The data explosion in remote sensing is straining the ability of data centers to deliver the data to the user community, yet many large-volume users actually seek a relatively small information component within the data, which they extract at their sites using Knowledge Discovery in Databases (KDD) techniques. To improve the efficiency of this process, the Goddard Earth Sciences Distributed Active Archive Center (GES DAAC) has implemented a KDD subsystem that supports execution of the user's KDD algorithm at the data center, dramatically reducing the volume that is sent to the user. The data are extracted from the archive in a planned, organized "campaign"; the algorithms are executed, and the output products sent to the users over the network. The first campaign, now complete, has resulted in overall reductions in shipped volume from 3.3 TB to 0.4 TB.

Lynnes, Christopher

Representing Learning With Graphical Models

Probabilistic graphical models are being used widely in artificial intelligence, for instance, in diagnosis and expert systems, as a unified qualitative and quantitative framework for representing and reasoning with probabilities and independencies. Their development and use spans several fields including artificial intelligence, decision theory and statistics, and provides an important bridge between these communities. This paper shows by way of example that these models can be extended to machine learning, neural networks and knowledge discovery by representing the notion of a sample on the graphical model. Not only does this allow a flexible variety of learning problems to be represented, it also provides the means for representing the goal of learning and opens the way for the automatic development of learning algorithms from specifications.

Buntine, Wray L.

Recursive Hierarchical Image Segmentation by Region Growing and Constrained Spectral Clustering

This paper describes an algorithm for hierarchical image segmentation (referred to as HSEG) and its recursive formulation (referred to as RHSEG). The HSEG algorithm is a hybrid of region growing and constrained spectral clustering that produces a hierarchical set of image segmentations based on detected convergence points. In the main, HSEG employs the hierarchical stepwise optimization (HS WO) approach to region growing, which seeks to produce segmentations that are more optimized than those produced by more classic approaches to region growing. In addition, HSEG optionally interjects between HSWO region growing iterations merges between spatially non-adjacent regions (i.e., spectrally based merging or clustering) constrained by a threshold derived from the previous HSWO region growing iteration. While the addition of constrained spectral clustering improves the segmentation results, especially for larger images, it also significantly increases HSEG's computational requirements. To counteract this, a computationally efficient recursive, divide-and-conquer, implementation of HSEG (RHSEG) has been devised and is described herein. Included in this description is special code that is required to avoid processing artifacts caused by RHSEG s recursive subdivision of the image data. Implementations for single processor and for multiple processor computer systems are described. Results with Landsat TM data are included comparing HSEG with classic region growing. Finally, an application to image information mining and knowledge discovery is discussed.

Tilton, James C.

An Artificial Intelligence Classification Tool and Its Application to Gamma-Ray Bursts

Despite being the most energetic phenomenon in the known universe, the astrophysics of gamma-ray bursts (GRBs) has still proven difficult to understand. It has only been within the past five years that the GRB distance scale has been firmly established, on the basis of a few dozen bursts with x-ray, optical, and radio afterglows. The afterglows indicate source redshifts of z=1 to z=5, total energy outputs of roughly 10(exp 52) ergs, and energy confined to the far x-ray to near gamma-ray regime of the electromagnetic spectrum. The multi-wavelength afterglow observations have thus far provided more insight on the nature of the GRB mechanism than the GRB observations; far more papers have been written about the few observed gamma-ray burst afterglows in the past few years than about the thousands of detected gamma-ray bursts. One reason the GRB central engine is still so poorly understood is that GRBs have complex, overlapping characteristics that do not appear to be produced by one homogeneous process. At least two subclasses have been found on the basis of duration, spectral hardness, and fluence (time integrated flux); Class 1 bursts are softer, longer, and brighter than Class 2 bursts (with two second durations indicating a rough division). A third GRB subclass, overlapping the other two, has been identified using statistical clustering techniques; Class 3 bursts are intermediate between Class 1 and Class 2 bursts in brightness and duration, but are softer than Class 1 bursts. We are developing a tool to aid scientists in the study of GRB properties. In the process of developing this tool, we are building a large gamma-ray burst classification database. We are also scientifically analyzing some GRB data as we develop the tool. Tool development thus proceeds in tandem with the dataset for which it is being designed. The tool invokes a modified KDD (Knowledge Discovery in Databases) process, which is described as follows.

Hakkila, Jon

Putting Priors in Mixture Density Mercer Kernels

This paper presents a new methodology for automatic knowledge driven data mining based on the theory of Mercer Kernels, which are highly nonlinear symmetric positive definite mappings from the original image space to a very high, possibly infinite dimensional feature space. We describe a new method called Mixture Density Mercer Kernels to learn kernel function directly from data, rather than using predefined kernels. These data adaptive kernels can en- code prior knowledge in the kernel using a Bayesian formulation, thus allowing for physical information to be encoded in the model. We compare the results with existing algorithms on data from the Sloan Digital Sky Survey (SDSS). The code for these experiments has been generated with the AUTOBAYES tool, which automatically generates efficient and documented C/C++ code from abstract statistical model specifications. The core of the system is a schema library which contains template for learning and knowledge discovery algorithms like different versions of EM, or numeric optimization methods like conjugate gradient methods. The template instantiation is supported by symbolic- algebraic computations, which allows AUTOBAYES to find closed-form solutions and, where possible, to integrate them into the code. The results show that the Mixture Density Mercer-Kernel described here outperforms tree-based classification in distinguishing high-redshift galaxies from low- redshift galaxies by approximately 16% on test data, bagged trees by approximately 7%, and bagged trees built on a much larger sample of data by approximately 2%.

Srivastava, Ashok N.

MIMIC II: a massive temporal ICU patient database to support research in intelligent patient monitoring

Development and evaluation of Intensive Care Unit (ICU) decision-support systems would be greatly facilitated by the availability of a large-scale ICU patient database. Following our previous efforts with the MIMIC (Multi-parameter Intelligent Monitoring for Intensive Care) Database, we have leveraged advances in networking and storage technologies to develop a far more massive temporal database, MIMIC II. MIMIC II is an ongoing effort: data is continuously and prospectively archived from all ICU patients in our hospital. MIMIC II now consists of over 800 ICU patient records including over 120 gigabytes of data and is growing. A customized archiving system was used to store continuously up to four waveforms and 30 different parameters from ICU patient monitors. An integrated user-friendly relational database was developed for browsing of patients' clinical information (lab results, fluid balance, medications, nurses' progress notes). Based upon its unprecedented size and scope, MIMIC II will prove to be an important resource for intelligent patient monitoring research, and will support efforts in medical data mining and knowledge-discovery.

NASA Discipline Cardiopulmonary

Exploiting Recurring Structure in a Semantic Network

With the growing popularity of the Semantic Web, an increasing amount of information is becoming available in machine interpretable, semantically structured networks. Within these semantic networks are recurring structures that could be mined by existing or novel knowledge discovery methods. The mining of these semantic structures represents an interesting area that focuses on mining both for and from the Semantic Web, with surprising applicability to problems confronting the developers of Semantic Web applications. In this paper, we present representative examples of recurring structures and show how these structures could be used to increase the utility of a semantic repository deployed at NASA.

Wolfe, Shawn R.

An Ensemble Approach to Building Mercer Kernels with Prior Information

This paper presents a new methodology for automatic knowledge driven data mining based on the theory of Mercer Kernels, which are highly nonlinear symmetric positive definite mappings from the original image space to a very high, possibly dimensional feature space. we describe a new method called Mixture Density Mercer Kernels to learn kernel function directly from data, rather than using pre-defined kernels. These data adaptive kernels can encode prior knowledge in the kernel using a Bayesian formulation, thus allowing for physical information to be encoded in the model. Specifically, we demonstrate the use of the algorithm in situations with extremely small samples of data. We compare the results with existing algorithms on data from the Sloan Digital Sky Survey (SDSS) and demonstrate the method's superior performance against standard methods. The code for these experiments has been generated with the AUTOBAYES tool, which automatically generates efficient and documented C/C++ code from abstract statistical model specifications. The core of the system is a schema library which contains templates for learning and knowledge discovery algorithms like different versions of EM, or numeric optimization methods like conjugate gradient methods. The template instantiation is supported by symbolic-algebraic computations, which allows AUTOBAYES to find closed-form solutions and, where possible, to integrate them into the code.

Srivastava, Ashok N.

NASA Tech Briefs, July 2007

Topics covered include: Miniature Intelligent Sensor Module; "Smart" Sensor Module; Portable Apparatus for Electrochemical Sensing of Ethylene; Increasing Linear Dynamic Range of a CMOS Image Sensor; Flight Qualified Micro Sun Sensor; Norbornene-Based Polymer Electrolytes for Lithium Cells; Making Single-Source Precursors of Ternary Semiconductors; Water-Free Proton-Conducting Membranes for Fuel Cells; Mo/Ti Diffusion Bonding for Making Thermoelectric Devices; Photodetectors on Coronagraph Mask for Pointing Control; High-Energy-Density, Low-Temperature Li/CFx Primary Cells; G4-FETs as Universal and Programmable Logic Gates; Fabrication of Buried Nanochannels From Nanowire Patterns; Diamond Smoothing Tools; Infrared Imaging System for Studying Brain Function; Rarefying Spectra of Whispering-Gallery-Mode Resonators; Large-Area Permanent-Magnet ECR Plasma Source; Slot-Antenna/Permanent-Magnet Device for Generating Plasma; Fiber-Optic Strain Gauge With High Resolution And Update Rate; Broadband Achromatic Telecentric Lens; Temperature-Corrected Model of Turbulence in Hot Jet Flows; Enhanced Elliptic Grid Generation; Automated Knowledge Discovery From Simulators; Electro-Optical Modulator Bias Control Using Bipolar Pulses; Generative Representations for Automated Design of Robots; Mars-Approach Navigation Using In Situ Orbiters; Efficient Optimization of Low-Thrust Spacecraft Trajectories; Cylindrical Asymmetrical Capacitors for Use in Outer Space; Protecting Against Faults in JPL Spacecraft; Algorithm Optimally Allocates Actuation of a Spacecraft; and Radar Interferometer for Topographic Mapping of Glaciers and Ice Sheets.

Source record

Integration and Cooperation in the Next Golden Age of Human Space Flight Data Repositories: Tools for Retrospective Analysis and Future Planning

As NASA transitions from the Space Shuttle era into the next phase of space exploration, the need to ensure the capture, analysis, and application of its research and medical data is of greater urgency than at any other previous time. In this era of limited resources and challenging schedules, the Human Research Program (HRP) based at NASA s Johnson Space Center (JSC) recognizes the need to extract the greatest possible amount of information from the data already captured, as well as focus current and future research funding on addressing the HRP goal to provide human health and performance countermeasures, knowledge, technologies, and tools to enable safe, reliable, and productive human space exploration. To this end, the Science Management Office and the Medical Informatics and Health Care Systems Branch within the HRP and the Space Medicine Division have been working to make both research data and clinical data more accessible to the user community. The Life Sciences Data Archive (LSDA), the research repository housing data and information regarding the physiologic effects of microgravity, and the Lifetime Surveillance of Astronaut Health Repository (LSAH-R), the clinical repository housing astronaut data, have joined forces to achieve this goal. The task of both repositories is to acquire, preserve, and distribute data and information both within the NASA community and to the science community at large. This is accomplished via the LSDA s public website (http://lsda.jsc.nasa.gov), which allows access to experiment descriptions including hardware, datasets, key personnel, mission descriptions and a mechanism for researchers to request additional data, research and clinical, that is not accessible from the public website. This will result in making the work of NASA and its partners available to the wider sciences community, both domestic and international. The desired outcome is the use of these data for knowledge discovery, retrospective analysis, and planning of future research studies.

Thomas, D.

Data Mining at NASA: From Theory to Applications

This slide presentation demonstrates the data mining/machine learning capabilities of NASA Ames and Intelligent Data Understanding (IDU) group. This will encompass the work done recently in the group by various group members. The IDU group develops novel algorithms to detect, classify, and predict events in large data streams for scientific and engineering systems. This presentation for Knowledge Discovery and Data Mining 2009 is to demonstrate the data mining/machine learning capabilities of NASA Ames and IDU group. This will encompass the work done re cently in the group by various group members.

Srivastava, Ashok N.

NASA Life Sciences Data Repositories: Tools for Retrospective Analysis and Future Planning

As NASA transitions from the Space Shuttle era into the next phase of space exploration, the need to ensure the capture, analysis, and application of its research and medical data is of greater urgency than at any other previous time. In this era of limited resources and challenging schedules, the Human Research Program (HRP) based at NASA s Johnson Space Center (JSC) recognizes the need to extract the greatest possible amount of information from the data already captured, as well as focus current and future research funding on addressing the HRP goal to provide human health and performance countermeasures, knowledge, technologies, and tools to enable safe, reliable, and productive human space exploration. To this end, the Science Management Office and the Medical Informatics and Health Care Systems Branch within the HRP and the Space Medicine Division have been working to make both research data and clinical data more accessible to the user community. The Life Sciences Data Archive (LSDA), the research repository housing data and information regarding the physiologic effects of microgravity, and the Lifetime Surveillance of Astronaut Health (LSAH-R), the clinical repository housing astronaut data, have joined forces to achieve this goal. The task of both repositories is to acquire, preserve, and distribute data and information both within the NASA community and to the science community at large. This is accomplished via the LSDA s public website (http://lsda.jsc.nasa.gov), which allows access to experiment descriptions including hardware, datasets, key personnel, mission descriptions and a mechanism for researchers to request additional data, research and clinical, that is not accessible from the public website. This will result in making the work of NASA and its partners available to the wider sciences community, both domestic and international. The desired outcome is the use of these data for knowledge discovery, retrospective analysis, and planning of future research studies.

Thomas, D.

Prognostics

Knowledge discovery, statistical learning, and more specifically an understanding of the system evolution in time when it undergoes undesirable fault conditions, are critical for an adequate implementation of successful prognostic systems. Prognosis may be understood as the generation of long-term predictions describing the evolution in time of a particular signal of interest or fault indicator, with the purpose of estimating the remaining useful life (RUL) of a failing component/subsystem. Predictions are made using a thorough understanding of the underlying processes and factor in the anticipated future usage.

Systems Health Management

Ask-The-Expert: Minimizing Human Review for Big Data Analytics Through Active Learning

In this CIF project, we worked toward semi-automating knowledge discovery from anomaly detection algorithms through the use of active learning. Active learning is an area of research within machine learning that uses an "expert in the loop" to learn from large data sets that have very few annotations or labels available, and where providing such labels is expensive. In our case, the task can be defined as the identification of safety events from flight operational data. Since traditional anomaly detection algorithms cannot differentiate between operationally relevant and irrelevant statistical anomalies, Subject Matter Experts (SMEs) have a lengthy and expensive burden of investigating every example identified by the detection algorithm, classifying and labeling them as relevant or irrelevant. Active learningidentifies the unlabeled example for which a label would most improve the classifier, asks the domain expert for a label, and repeats this process until there are no more resources (time, budget) available for labeling or a minimum required performance is reached. A positive label indicates an operationally significant safety event whereas a negative label indicates otherwise. Based on these few labels we propose to build an active learning system that utilizes the SME's time in the most effective manner by iteratively asking for labels for as few informative instances as possible. Our work was proposed to be a stepping stone toward implementation and deployment of the system with user interface to be pursued by the Aviation Operations and Safety Program (AOSP) given its interest in safety monitoring and discovery of safety incidents.

aviation safety

GeoNEX: A Cloud Gateway for Near Real-time Processing of Geostationary Satellite Products

The emergence of a new generation of geostationary satellite sensors provides land andatmosphere monitoring capabilities similar to MODIS and VIIRS with far greater temporal resolution (5-15 minutes). However, processing such large volume, highly dynamic datasets requires computing capabilities that (1) better support data access and knowledge discovery for scientists; (2) provide resources to enable real-time processing for emergency response (wildfire, smoke, dust, etc.); and (3) provide reliable and scalable services for the broader user community. This paper presents an implementation of GeoNEX (Geostationary NASA-NOAA Earth Exchange) services that integrate scientific algorithms with Amazon Web Services (AWS) to provide near realtime monitoring (~5 minute latency) capability in a hybrid cloud-computing environment. It offers a user-friendly, manageable and extendable interface and benefits from the scalability provided by Amazon Web Services. Four use cases are presented to illustrate how to (1) search and access geostationary data; (2) configure computing infrastructure to enable near real-time processing; (3) disseminate and utilize research results, visualizations, and animations to concurrent users; and (4) use a Jupyter Notebook-like interface for data exploration and rapid prototyping. As an example of (3), the Wildfire Automated Biomass Burning Algorithm (WF_ABBA) was implemented on GOES-16 and -17 data to produce an active fire map every 5 minutes over the conterminous US. Details of the implementation strategies, architectures, and challenges of the use cases are discussed.

GeoNEX

NASA’s Human Data Repositories: An In Depth Look at the New Data Request Process

As NASA transitions its focus to travel back to the moon and on to new destinations, the need to ensure the capture, analysis, and application of research and medical data is of greater urgency than at any other previous time. In this era of limited resources and challenging schedules, the Human Research Program (HRP), based at NASA’s Johnson Space Center (JSC), recognizes the need to extract the greatest possible amount of information from the data already captured. To this end, the HRP Chief Scientist Office (CSO), HRP Program Planning and Control (PP&C) Office, and the Space Medicine Operations Division have been working together to make reuse of both research data and medical monitoring data more accessible to the user community through the Life Science Data Archive (LSDA) and the Lifetime Surveillance of Astronaut Health (LSAH) Repositories. The task of both LSDA and LSAH repositories is to acquire, preserve, and distribute retrospective research (LSDA) and medical (LSAH) data and information both within the NASA community and to the science community at large, for knowledge discovery, retrospective analysis, and planning of future research studies. An additional goal is to encourage collaboration with non-NASA institutions also faced with enhancing human performance in extreme environments. In September 2022, the LSDA website and its contents transitioned to a new NASA Life Sciences Portal (https://nlsp.nasa.gov/explore/lsdahome). This site continues to feature publicly releasable information such as non-attributable datasets, experiment descriptions (from Project Mercury to ISS, as well as from multiple flight analog missions), descriptions of medical monitoring data, and LSAH newsletters (1992 - 2022). The website also provides an updated portal to request additional research and medical data not accessible from the public website. This presentation will provide an in-depth look at the new system as it relates to finding and requesting retrospective data. We will also detail processes from making a request to delivering data for different types of data requests (i.e., attributable, or non-attributable). This includes descriptions of various approval boards, what information and actions the requestor is responsible for, and key milestones in making data available for reuse.

D. M. Thomas

Spaceflight Environmental-Telemetry Data for Biological Science

There is a critical need for better access and visualization of spaceflight environmental telemetry and mission hardware data from sensors including relative humidity, carbon dioxide, oxygen, radiation, airflow, temperature, acceleration, and acoustics. Under the stewardship of the Ames Life Sciences Data Archive (ALSDA) and GeneLab, an effort is underway to consolidate, normalize and provide accessibility of archived mission environmental data and hardware information, with the purpose of providing important context to biological data. This effort is necessary to provide scientific context of its impact upon biological and biomedical data from spaceflight missions and experiments (genomic, metagenomic, gene expression, proteomic, metabolomic, physiological, phenomics, behavioral; tabular, imaging, video). Environmental spaceflight data is derived from dozens of sources, with various formats, and in the past year a pipeline is in development to collect, curate and present this data efficiently. In the upcoming year, a new Data Visualization Portal will utilize the standardized pipeline data to provide easy user access to compare parameters and environmental conditions between missions, locations, subjects, and durations. Environmental and hardware data enables broad accessibility and analytics, without the need for advanced data informatic expertise. Familiarity with the capabilities and limitations of a variety of existing hardware/tools is a strength that could be applied to creation of improved hardware for future ecosystems on the Moon and Mars. The intention is to make biological and environmental telemetry data maximally open-access and FAIR (findable, accessible, interoperable, reusable) for data mining-informatic approaches to support knowledge discovery necessary for low Earth orbit, cis-Lunar, Mars transit, and Mars surface missions.

Danielle K. Lopez

The NASA Open Science Data Repository: Biomedical Data, Analysis Tools, and Informatic Collaborations

Increased biomedical risks and challenges associated with deep space missions require knowledge discovery, health countermeasures, and biomedical support capabilities. Maximally open-access and reusable data is needed by developers, scientists, and engineers to develop these systems. The NASA Open Science Data Repository (OSDR) is a maximally open access and FAIR database (ie., findable, accessible, interoperable, and reusable), and meets various scientific, technical, and operational needs. It offers users and submitters the ability to upload, download, search, share, analyze, cite, and visualize data across ‘omics, physiological, phenotypic, payload, hardware, behavioral, bioimaging, video, and environmental monitoring telemetry datasets. OSDR is an expanded database, based upon the successes of NASA GeneLab. OSDR has >460 studies with datasets covering model organisms to non-NASA human astronauts. There are ~12 datasets from the Inspiration 4 (I4) mission, spanning metagenomics, comprehensive metabolic panels, clonal hematopoiesis, spatial transcriptomics, proteomics, and cytokine panels. In the interest of data privacy, two I4 datasets with raw files relating to the epitranscriptome, and a new request feature is live in OSDR (with a backend review process established) developed from industry norms. OSDR is collecting and curating biomedical human data from a new sub-orbital research flight and is open to more space life science/biomedical submissions from the international and commercial sectors. OSDR also recently began a collaboration with the European Space Agency (ESA) to collect and curate >200 terabytes of human and model organism data. The OSDR submission portal is designed to ingest and curate ~25 ‘omics and ~50 physiological-phenotypic-imaging assay data types. Tools available for OSDR users include: 1) an Environmental Data Application to compare radiation, CO2, relative humidity, temperature, and other telemetry across missions and subjects, 2) the RadLab database, a collaboration between NASA, ESA, the German and Italian Space Agencies, and the Bulgarian Academy of Sciences, and 3) a Multi-study visualization tool which enables users to look across and combine ‘omics datasets. There are ~600 volunteer OSDR Analysis Working Group (AWG) members providing feedback on scientific data/metadata standards and collaborating to mine-reuse OSDR in research. OSDR/GeneLab has enabled ~60 publications reusing data as of October 2023.

space biology