Search NASASearch

SEARCH · Search NASA

Results for “mining techniques”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Discovering System Health Anomalies Using Data Mining Techniques

We present a data mining framework for the analysis and discovery of anomalies in high-dimensional time series of sensor measurements that would be found in an Integrated System Health Monitoring system. We specifically treat the problem of discovering anomalous features in the time series that may be indicative of a system anomaly, or in the case of a manned system, an anomaly due to the human. Identification of these anomalies is crucial to building stable, reusable, and cost-efficient systems. The framework consists of an analysis platform and new algorithms that can scale to thousands of sensor streams to discovers temporal anomalies. We discuss the mathematical framework that underlies the system and also describe in detail how this framework is general enough to encompass both discrete and continuous sensor measurements. We also describe a new set of data mining algorithms based on kernel methods and hidden Markov models that allow for the rapid assimilation, analysis, and discovery of system anomalies. We then describe the performance of the system on a real-world problem in the aircraft domain where we analyze the cockpit data from aircraft as well as data from the aircraft propulsion, control, and guidance systems. These data are discrete and continuous sensor measurements and are dealt with seamlessly in order to discover anomalous flights. We conclude with recommendations that describe the tradeoffs in building an integrated scalable platform for robust anomaly detection in ISHM applications.

Sriastava, Ashok, N.

Analysis of Airport Ground Delay Program Decisions Using Data Mining Techniques

Air traffic service providers have to make decisions regarding changes to air traffic flow in the event of major weather disturbances and traffic congestions to maintain safety of the system. The behavior of the air traffic management system will be more predictable if consistent decisions are made under similar traffic and weather conditions. Consistency of deciding on control action depends on the weather and traffic conditions as well as accuracy in predicting these conditions. Weather parameters (defined in terms of forecast and actual weather and traffic conditions) on different days can be used to categorize days into days with little decision consistency, days with moderate decision consistency and days with high decision consistency. Four years of traffic, weather and ground delay program decisions data at major airports in the United States are used in the analysis. This paper examines performance of different data mining methods in the three regions of decision consistency. Not surprisingly, data mining methods have the best performance in the region of most decision consistency and have the poorest performance in the region of little decision consistency. In applications where data mining methods have differing performance in differing regions, it would be more useful to characterize the region specific performance instead of characterizing performance by a single parameter. Finally, the results show no significant variation in the performance of different data mining methods for this particular problem. The fact that different mining methods show no significant variation also provides further confidence in the results of data mining methods. Work in this abstract discusses initial results. This paper describes the results in terms of both forecast and actual environmental conditions and discusses how prediction errors impact decision consistency.

Kulkarni, Deepak

Data Fusion and Mining Techniques to Map Water Use and Drought across Spatial and Temporal Scales

As the world’s water resources come under increasing tension due to dual stressors of climate change and population growth, accurate knowledge of water consumption through evapotranspiration (ET) over a range in spatial scales will be critical in developing adaptation strategies. Remote sensing methods for monitoring consumptive water use (e.g, ET) are becoming increasingly important, especially in areas of significant water and food insecurity. One method to estimate ET from satellite-based methods, the Atmosphere Land Exchange Inverse (ALEXI) model uses the change in mid-morning land surface temperature to estimate the partitioning of sensible and latent heat fluxes which are then used to estimate daily ET. This presentation will outline several recent enhancements to the ALEXI modeling system, with a focus on global ET and drought monitoring. Until recently, ALEXI has been limited to areas with high resolution temporal sampling of geostationary sensors. The use of geostationary sensors makes global mapping a complicated process, especially for real-time applications, as data from as many as five different sensors are required to be ingested and harmonized to create a global mosaic. However, our research team has developed a new and novel method of using twice-daily observations from polar-orbiting sensors such as MODIS and VIIRS to estimate the mid-morning rise in LST that is used to drive the energy balance estimations within ALEXI. This allows the method to be applied globally using a single sensor (in this case, initially MODIS with a planned transition to VIIRS) rather than a global compositing of all available geostationary data. Other advantages of this new method include the higher spatial resolution provided by MODIS and VIIRS and the increased sampling at high latitudes where oblique view angles limit the utility of geostationary sensors. This presentation will focus on global applications for mapping water use and drought using data mining and data fusion across spatial scales extending from 5-km to 30-m “field-scale” estimates.

Christopher Hain

Virtual Sensors: Using Data Mining Techniques to Efficiently Estimate Remote Sensing Spectra

Various instruments are used to create images of the Earth and other objects in the universe in a diverse set of wavelength bands with the aim of understanding natural phenomena. These instruments are sometimes built in a phased approach, with some measurement capabilities being added in later phases. In other cases, there may not be a planned increase in measurement capability, but technology may mature to the point that it offers new measurement capabilities that were not available before. In still other cases, detailed spectral measurements may be too costly to perform on a large sample. Thus, lower resolution instruments with lower associated cost may be used to take the majority of measurements. Higher resolution instruments, with a higher associated cost may be used to take only a small fraction of the measurements in a given area. Many applied science questions that are relevant to the remote sensing community need to be addressed by analyzing enormous amounts of data that were generated from instruments with disparate measurement capability. This paper addresses this problem by demonstrating methods to produce high accuracy estimates of spectra with an associated measure of uncertainty from data that is perhaps nonlinearly correlated with the spectra. In particular, we demonstrate multi-layer perceptrons (MLPs), Support Vector Machines (SVMs) with Radial Basis Function (RBF) kernels, and SVMs with Mixture Density Mercer Kernels (MDMK). We call this type of an estimator a Virtual Sensor because it predicts, with a measure of uncertainty, unmeasured spectral phenomena.

Srivastava, Ashok N.

Assessment of practicality of remote sensing techniques for a study of the effects of strip mining in Alabama

Because of the volume of coal produced by strip mining, the proximity of mining operations, and the diversity of mining methods (e.g. contour stripping, area stripping, multiple seam stripping, and augering, as well as underground mining), the Warrior Coal Basin seemed best suited for initial studies on the physical impact of strip mining in Alabama. Two test sites, (Cordova and Searles) representative of the various strip mining techniques and environmental problems, were chosen for intensive studies of the correlation between remote sensing and ground truth data. Efforts were eventually concentrated in the Searles Area, since it is more accessible and offers a better opportunity for study of erosional and depositional processes than the Cordova Area.

Hughes, T. H.

General Purpose Data-Driven System Monitoring for Space Operations

Modern space propulsion and exploration system designs are becoming increasingly sophisticated and complex. Determining the health state of these systems using traditional methods is becoming more difficult as the number of sensors and component interactions grows. Data-driven monitoring techniques have been developed to address these issues by analyzing system operations data to automatically characterize normal system behavior. The Inductive Monitoring System (IMS) is a data-driven system health monitoring software tool that has been successfully applied to several aerospace applications. IMS uses a data mining technique called clustering to analyze archived system data and characterize normal interactions between parameters. This characterization, or model, of nominal operation is stored in a knowledge base that can be used for real-time system monitoring or for analysis of archived events. Ongoing and developing IMS space operations applications include International Space Station flight control, spacecraft vehicle system health management, launch vehicle ground operations, and fleet supportability. As a common thread of discussion this paper will employ the evolution of the IMS data-driven technique as related to several Integrated Systems Health Management (ISHM) elements. Thematically, the projects listed will be used as case studies. The maturation of IMS via projects where it has been deployed or is currently being integrated to aid in fault detection will be described. The paper will also explain how IMS can be used to complement a suite of other ISHM tools, providing initial fault detection support for diagnosis and recovery

Space Propulsion

General Purpose Data-Driven System Monitoring for Space Operations

Modern space propulsion and exploration system designs are becoming increasingly sophisticated and complex. Determining the health state of these systems using traditional methods is becoming more difficult as the number of sensors and component interactions grows. Data-driven monitoring techniques have been developed to address these issues by analyzing system operations data to automatically characterize normal system behavior. The Inductive Monitoring System (IMS) is a data-driven system health monitoring software tool that has been successfully applied to several aerospace applications. IMS uses a data mining technique called clustering to analyze archived system data and characterize normal interactions between parameters. This characterization, or model, of nominal operation is stored in a knowledge base that can be used for real-time system monitoring or for analysis of archived events. Ongoing and developing IMS space operations applications include International Space Station flight control, satellite vehicle system health management, launch vehicle ground operations, and fleet supportability. As a common thread of discussion this paper will employ the evolution of the IMS data-driven technique as related to several Integrated Systems Health Management (ISHM) elements. Thematically, the projects listed will be used as case studies. The maturation of IMS via projects where it has been deployed, or is currently being integrated to aid in fault detection will be described. The paper will also explain how IMS can be used to complement a suite of other ISHM tools, providing initial fault detection support for diagnosis and recovery.

Satellites

Application of LANDSAT data to monitor land reclamation progress in Belmont County, Ohio

Strip and contour mining techniques are reviewed as well as some studies conducted to determine the applicability of LANDSAT and associated digital image processing techniques to the surficial problems associated with mining operations. A nontraditional unsupervised classification approach to multispectral data is considered which renders increased classification separability in land cover analysis of surface mined areas. The approach also reduces the dimensionality of the data and requires only minimal analytical skills in digital data processing.

Bloemer, H. H. L.

Enabling the Discovery of Recurring Anomalies in Aerospace System Problem Reports using High-Dimensional Clustering Techniques

This paper describes the results of a significant research and development effort conducted at NASA Ames Research Center to develop new text mining techniques to discover anomalies in free-text reports regarding system health and safety of two aerospace systems. We discuss two problems of significant importance in the aviation industry. The first problem is that of automatic anomaly discovery about an aerospace system through the analysis of tens of thousands of free-text problem reports that are written about the system. The second problem that we address is that of automatic discovery of recurring anomalies, i.e., anomalies that may be described m different ways by different authors, at varying times and under varying conditions, but that are truly about the same part of the system. The intent of recurring anomaly identification is to determine project or system weakness or high-risk issues. The discovery of recurring anomalies is a key goal in building safe, reliable, and cost-effective aerospace systems. We address the anomaly discovery problem on thousands of free-text reports using two strategies: (1) as an unsupervised learning problem where an algorithm takes free-text reports as input and automatically groups them into different bins, where each bin corresponds to a different unknown anomaly category; and (2) as a supervised learning problem where the algorithm classifies the free-text reports into one of a number of known anomaly categories. We then discuss the application of these methods to the problem of discovering recurring anomalies. In fact the special nature of recurring anomalies (very small cluster sizes) requires incorporating new methods and measures to enhance the original approach for anomaly detection. ?& pant 0-

Srivastava, Ashok, N.

Comparing digital data processing techniques for surface mine and reclamation monitoring

The results of three techniques used for processing Landsat digital data are compared for their utility in delineating areas of surface mining and subsequent reclamation. An unsupervised clustering algorithm (ISOCLS), a maximum-likelihood classifier (CLASFY), and a hybrid approach utilizing canonical analysis (ISOCLS/KLTRANS/ISOCLS) were compared by means of a detailed accuracy assessment with aerial photography at NASA's Goddard Space Flight Center. Results show that the hybrid approach was superior to the traditional techniques in distinguishing strip mined and reclaimed areas.

Witt, R. G.

Utilization of lunar materials in space

Reasons for conducting commercial mining operations on the moon are discussed with attention to physical parameters, material abundances, and economics. Adaptations of currently used mining techniques are considered, and space applications of moon-derived materials are suggested. Possible organization of the mining project is examined, and it is suggested that the transition from concept phase to implementation could proceed rapidly. Characteristics of maturing space industries and the roles of the public and the private sectors are considered.

Criswell, D. R.

Vibration Based Crack Detection in a Rotating Disk: Experimental Results - Part 2

This paper describes the experimental results concerning the detection of a crack in a rotating disk. The goal was to utilize blade tip clearance and shaft vibration measurements to monitor changes in the system's center of mass and/or blade deformation behaviors. The concept of the approach is based on the fact that the development of a disk crack results in a distorted strain field within the component. As a result, a minute deformation in the disk's geometry as well as a change in the system's center of mass occurs. Here, a notch was used to simulate an actual crack. The vibration based experimental results failed to identify the existence of a notch when utilizing the approach described above, even with a rather large, circumferential notch (l.2 in.) located approximately mid-span on the disk (disk radius = 4.63 in. with notch at r = 2.12 in.). This was somewhat expected, since the finite element based results in Part 1 of this study predicted changes in blade tip clearance as well as center of mass shifts due to a notch to be less than 0.001 in. Therefore, the small changes incurred by the notch could not be differentiated from the mechanical and electrical noise of the rotor system. Although the crack detection technique of interest failed to identify the existence ofthe notch, the vibration data produced and captured here will be utilized in upcoming studies that will focus on different data mining techniques concerning damage detection in a disk.

Gyekenyesi, Andrew L.

The Ocean Carbon States Database: A Proof-of-Concept Application of Cluster Analysis in the Ocean Carbon Cycle

In this paper, we present a database of the basic regimes of the carbon cycle in the ocean, the 'ocean carbon states', as obtained using a data mining/pattern recognition technique in observation-based as well as model data. The goal of this study is to establish a new data analysis methodology, test it and assess its utility in providing more insights into the regional and temporal variability of the marine carbon cycle. This is important as advanced data mining techniques are becoming widely used in climate and Earth sciences and in particular in studies of the global carbon cycle, where the interaction of physical and biogeochemical drivers confounds our ability to accurately describe, understand, and predict CO2 concentrations and their changes in the major planetary carbon reservoirs. In this proof-of-concept study, we focus on using well-understood data that are based on observations, as well as model results from the NASA Goddard Institute for Space Studies (GISS) climate model. Our analysis shows that ocean carbon states are associated with the subtropical-subpolar gyre during the colder months of the year and the tropics during the warmer season in the North Atlantic basin. Conversely, in the Southern Ocean, the ocean carbon states can be associated with the subtropical and Antarctic convergence zones in the warmer season and the coastal Antarctic divergence zone in the colder season. With respect to model evaluation, we find that the GISS model reproduces the cold and warm season regimes more skillfully in the North Atlantic than in the Southern Ocean and matches the observed seasonality better than the spatial distribution of the regimes. Finally, the ocean carbon states provide useful information in the model error attribution. Model air-sea CO2 flux biases in the North Atlantic stem from wind speed and salinity biases in the subpolar region and nutrient and wind speed biases in the subtropics and tropics. Nutrient biases are shown to be most important in the Southern Ocean flux bias.

carbon cycle

KDD Services at the Goddard Earth Sciences Distributed Active Archive Center

NASA's Goddard Earth Sciences Distributed Active Archive Center (GES DAAC) processes, stores and distributes earth science data from a variety of remote sensing satellites. End users of the data range from instrument scientists to global change and climate researchers to federal agencies and foreign governments. Many of these users apply data mining techniques to large volumes of data (up to 1 TB) received from the GES DAAC. However, rapid advances in processing power are enabling increases in data processing that are outpacing tape drive performance and network capacity. As a result, the proportion of data that can be distributed to users continues to decrease. As mitigation, we are migrating more data mining and mining preparation activities into the data center in order to reduce the data volume that needs to be distributed and to offer the users a more useful and manageable product. This migration of activities faces a number of technical and human-factor challenges. As data reduction and mining algorithms are normally quite specific to the user's research needs, the user's algorithm must be integrated virtually unchanged into the archive environment. Also, the archive itself is busy with everyday data archive and distribution activities and cannot be dedicated to, or even impacted by, the mining activities. Therefore, we schedule KDD 'campaigns' (similar to reprocessing campaigns), during which we schedule a wholesale retrieval of specific data products, offering users the opportunity to extract information from the data being retrieved during the campaign.

Lynnes, Christopher

Inductive System Health Monitoring

The Inductive Monitoring System (IMS) software was developed to provide a technique to automatically produce health monitoring knowledge bases for systems that are either difficult to model (simulate) with a computer or which require computer models that are too complex to use for real time monitoring. IMS uses nominal data sets collected either directly from the system or from simulations to build a knowledge base that can be used to detect anomalous behavior in the system. Machine learning and data mining techniques are used to characterize typical system behavior by extracting general classes of nominal data from archived data sets. IMS is able to monitor the system by comparing real time operational data with these classes. We present a description of learning and monitoring method used by IMS and summarize some recent IMS results.

Iverson, David L.

INDUCTIVE SYSTEM HEALTH MONITORING WITH STATISTICAL METRICS

Model-based reasoning is a powerful method for performing system monitoring and diagnosis. Building models for model-based reasoning is often a difficult and time consuming process. The Inductive Monitoring System (IMS) software was developed to provide a technique to automatically produce health monitoring knowledge bases for systems that are either difficult to model (simulate) with a computer or which require computer models that are too complex to use for real time monitoring. IMS processes nominal data sets collected either directly from the system or from simulations to build a knowledge base that can be used to detect anomalous behavior in the system. Machine learning and data mining techniques are used to characterize typical system behavior by extracting general classes of nominal data from archived data sets. In particular, a clustering algorithm forms groups of nominal values for sets of related parameters. This establishes constraints on those parameter values that should hold during nominal operation. During monitoring, IMS provides a statistically weighted measure of the deviation of current system behavior from the established normal baseline. If the deviation increases beyond the expected level, an anomaly is suspected, prompting further investigation by an operator or automated system. IMS has shown potential to be an effective, low cost technique to produce system monitoring capability for a variety of applications. We describe the training and system health monitoring techniques of IMS. We also present the application of IMS to a data set from the Space Shuttle Columbia STS-107 flight. IMS was able to detect an anomaly in the launch telemetry shortly after a foam impact damaged Columbia's thermal protection system.

Iverson, David L.