Search NASASearch

SEARCH · Search NASA

Results for “data mining”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Data Mining for Understanding and Impriving Decision-Making Affecting Ground Delay Programs

The continuous growth in the demand for air transportation results in an imbalance between airspace capacity and traffic demand. The airspace capacity of a region depends on the ability of the system to maintain safe separation between aircraft in the region. In addition to growing demand, the airspace capacity is severely limited by convective weather. During such conditions, traffic managers at the FAA's Air Traffic Control System Command Center (ATCSCC) and dispatchers at various Airlines' Operations Center (AOC) collaborate to mitigate the demand-capacity imbalance caused by weather. The end result is the implementation of a set of Traffic Flow Management (TFM) initiatives such as ground delay programs, reroute advisories, flow metering, and ground stops. Data Mining is the automated process of analyzing large sets of data and then extracting patterns in the data. Data mining tools are capable of predicting behaviors and future trends, allowing an organization to benefit from past experience in making knowledge-driven decisions. The work reported in this paper is focused on ground delay programs. Data mining algorithms have the potential to develop associations between weather patterns and the corresponding ground delay program responses. If successful, they can be used to improve and standardize TFM decision resulting in better predictability of traffic flows on days with reliable weather forecasts. The approach here seeks to develop a set of data mining and machine learning models and apply them to historical archives of weather observations and forecasts and TFM initiatives to determine the extent to which the theory can predict and explain the observed traffic flow behaviors.

data mining

Data Mining and Optimization Tools for Developing Engine Parameters Tools

This project was awarded for understanding the problem and developing a plan for Data Mining tools for use in designing and implementing an Engine Condition Monitoring System. Tricia Erhardt and I studied the problem domain for developing an Engine Condition Monitoring system using the sparse and non-standardized datasets to be available through a consortium at NASA Lewis Research Center. We visited NASA three times to discuss additional issues related to dataset which was not made available to us. We discussed and developed a general framework of data mining and optimization tools to extract useful information from sparse and non-standard datasets. These discussions lead to the training of Tricia Erhardt to develop Genetic Algorithm based search programs which were written in C++ and used to demonstrate the capability of GA algorithm in searching an optimal solution in noisy, datasets. From the study and discussion with NASA LeRC personnel, we then prepared a proposal, which is being submitted to NASA for future work for the development of data mining algorithms for engine conditional monitoring. The proposed set of algorithm uses wavelet processing for creating multi-resolution pyramid of tile data for GA based multi-resolution optimal search.

Dhawan, Atam P.

Data Mining and Optimization Tools for Developing Engine Parameters Tools

This project was awarded for understanding the problem and developing a plan for Data Mining tools for use in designing and implementing an Engine Condition Monitoring System. From the total budget of $5,000, Tricia and I studied the problem domain for developing ail Engine Condition Monitoring system using the sparse and non-standardized datasets to be available through a consortium at NASA Lewis Research Center. We visited NASA three times to discuss additional issues related to dataset which was not made available to us. We discussed and developed a general framework of data mining and optimization tools to extract useful information from sparse and non-standard datasets. These discussions lead to the training of Tricia Erhardt to develop Genetic Algorithm based search programs which were written in C++ and used to demonstrate the capability of GA algorithm in searching an optimal solution in noisy datasets. From the study and discussion with NASA LERC personnel, we then prepared a proposal, which is being submitted to NASA for future work for the development of data mining algorithms for engine conditional monitoring. The proposed set of algorithm uses wavelet processing for creating multi-resolution pyramid of the data for GA based multi-resolution optimal search. Wavelet processing is proposed to create a coarse resolution representation of data providing two advantages in GA based search: 1. We will have less data to begin with to make search sub-spaces. 2. It will have robustness against the noise because at every level of wavelet based decomposition, we will be decomposing the signal into low pass and high pass filters.

Dhawan, Atam P.

Knowledge Discovery and Data Mining: An Overview

The process of knowledge discovery and data mining is the process of information extraction from very large databases. Its importance is described along with several techniques and considerations for selecting the most appropriate technique for extracting information from a particular data set.

data mining knowledge discovery data search

Data Mine and Forget It?: A Cautionary Tale

With the development of new technologies, data mining has become increasingly popular. However, caution should be exercised in choosing the variables to include in data mining. A series of regression trees was created to demonstrate the change in the selection by the program of significant predictors based on the nature of variables.

Tada, Yuri

Data Mining for Science of the Sun-Earth Connection as a Single System

Establishing the Sun-Earth connection requires overcoming the challenges of exploring the data from past and current missions and leveraging tools and models (data mining) to create an efficient system treatment of the Sun and heliosphere. However, solar and heliospheric environment data constitute a vast source of information whose potential is far from being optimally exploited. In the next decade, the solar and heliospheric community will have to manage the increasing amount of information coming from new missions, improve reanalysis of data from past and current missions, and create new data products from the application of new methodologies. This complex task is further complicated by practical challenges such as different datasets and catalogs in different formats that may require different pre-processing and analysis tools, and the need for numerous analysis approaches that are not all fully optimized for large volumes of data. While several ongoing efforts aim at addressing these problems, the available datasets and tools are not always used to their full potential often due to lack of awareness of available resources. In this paper, we summarize the issues raised and goals discussed by members of the community during recent conference sessions focused on data mining for science.

Sun-Earth connection

High Performance EVA Glove Collaboration: Glove Injury Data Mining Effort

Human hands play a significant role during extravehicular activity (EVA) missions and Neutral Buoyancy Lab (NBL) training events, as they are needed for translating and performing tasks in the weightless environment. It is because of this high frequency usage that hand- and arm-related injuries and discomfort are known to occur during training in the NBL and while conducting EVAs. Hand-related injuries and discomforts have been occurring to crewmembers since the days of Apollo. While there have been numerous engineering changes to the glove design, hand-related issues still persist. The primary objectives of this study are therefore to: 1) document all known EVA glove-related injuries and the circumstances of these incidents, 2) determine likely risk factors, and 3) recommend ergonomic mitigations or design strategies that can be implemented in the current and future glove designs. METHODS: The investigator team conducted an initial set of literature reviews, data mining of Lifetime Surveillance of Astronaut Health (LSAH) databases, and data distribution analyses to understand the ergonomic issues related to glove-related injuries and discomforts. The investigation focused on the injuries and discomforts of U.S. crewmembers who had worn pressurized suits and experienced glove-related incidents during the 1980 to 2010 time frame, either during training or on-orbit EVA. In addition to data mining of the LSAH database, the other objective of the study was to find complimentary sources of information such as training experience, EVA experience, suit-related sizing data, and hand-arm anthropometric data to be tied to the injury data from LSAH. RESULTS: Past studies indicated that the hand was the most frequently injured part of the body during both EVA and NBL training. This study effort thus focused primarily on crew training data in the NBL between 2002 and 2010. Of the 87 recorded training incidents, 19 occurred to women and 68 to men. While crew ages ranged from thirties to fifties, the age category most affected was in the forties range. Incident rate calculations (incidents per 100 training runs) revealed that the 2002, 2003, and 2004 time periods registered the highest reported incident rate levels (3.4, 6.1, and 4.1 respectively) when compared to the following years (all ≤ 1.0). In addition to general hand-arm discomfort being the highest reported result from training, specific types of hand injuries or symptoms included erythema, fingernail delamination, abrasions, muscle soreness/fatigue, paresthesia, bruising, blanching, and edema. Specific body locations most affected by hand injuries included the metacarpophalangeal joints, fingernails, finger crotches, fingers in general, interphalangeal joints, and fingertips. Causes of injuries reported in the LSAH data were primarily attributed to the forces that the gloved hands were exposed to due to hand intensive tasks and/or poor glove sizing. DISCUSSION: Although the age data indicate that most injuries are reported by male crewmembers in their forties, that is also the dominant gender and age range of most EVA crew therefore it is not an unexpected finding. Age and gender analysis will continue as more details on the uninjured population is accrued. While there is a reasonable mechanism to link training quantity to injury, the results were inconsistent and point to the need for a consistent method of suit-related injury screening and documentation. For instance, the high-incident rate levels for the years 2002 to 2004 could be attributed to a comprehensive medical review of crewmembers post-NBL EVA training that occurred from July 19, 2002 to January 16, 2004. Furthermore, there could have been increased awareness from an investigation at the NBL. These investigations may have temporarily increased the fidelity of reported injuries and discomforts during these dates as compared to surrounding years, when injury signs and symptom were no longer actively being investigated but rather voluntarily reported. Data mining for possible mechanistic factors continues and includes more detailed training timelines, hand anthropometry, and suit sizing information. The limited published data looking at hand-arm anthropometry correlated hand-anthropometry metrics with injuries stemming from glove design and operation. Future work will include further evaluation of body sizing and fit in relation to hand injury incidents.

Reid, C. R.

Data Mining SIAM Presentation

This viewgraph document describes the data mining system developed at NASA Ames. Many NASA programs have large numbers (and types) of problem reports.These free text reports are written by a number of different people, thus the emphasis and wording vary considerably With so much data to sift through, analysts (subject experts) need help identifying any possible safety issues or concerns and help them confirm that they haven't missed important problems. Unsupervised clustering is the initial step to accomplish this; We think we can go much farther, specifically, identify possible recurring anomalies. Recurring anomalies may be indicators of larger systemic problems. The requirement to identify these anomalies has led to the development of Recurring Anomaly Discovery System (ReADS).

Srivastava, Ashok

Data Mining for Understanding and Improving Decision-making Affecting Ground Delay Programs

The continuous growth in the demand for air transportation results in an imbalance between airspace capacity and traffic demand. The airspace capacity of a region depends on the ability of the system to maintain safe separation between aircraft in the region. In addition to growing demand, the airspace capacity is severely limited by convective weather. During such conditions, traffic managers at the FAA's Air Traffic Control System Command Center (ATCSCC) and dispatchers at various Airlines' Operations Center (AOC) collaborate to mitigate the demand-capacity imbalance caused by weather. The end result is the implementation of a set of Traffic Flow Management (TFM) initiatives such as ground delay programs, reroute advisories, flow metering, and ground stops. Data Mining is the automated process of analyzing large sets of data and then extracting patterns in the data. Data mining tools are capable of predicting behaviors and future trends, allowing an organization to benefit from past experience in making knowledge-driven decisions.

Weather

Advanced Query and Data Mining Capabilities for MaROS

The Mars Relay Operational Service (MaROS) comprises a number of tools to coordinate, plan, and visualize various aspects of the Mars Relay network. These levels include a Web-based user interface, a back-end "ReSTlet" built in Java, and databases that store the data as it is received from the network. As part of MaROS, the innovators have developed and implemented a feature set that operates on several levels of the software architecture. This new feature is an advanced querying capability through either the Web-based user interface, or through a back-end REST interface to access all of the data gathered from the network. This software is not meant to replace the REST interface, but to augment and expand the range of available data. The current REST interface provides specific data that is used by the MaROS Web application to display and visualize the information; however, the returned information from the REST interface has typically been pre-processed to return only a subset of the entire information within the repository, particularly only the information that is of interest to the GUI (graphical user interface). The new, advanced query and data mining capabilities allow users to retrieve the raw data and/or to perform their own data processing. The query language used to access the repository is a restricted subset of the structured query language (SQL) that can be built safely from the Web user interface, or entered as freeform SQL by a user. The results are returned in a CSV (Comma Separated Values) format for easy exporting to third party tools and applications that can be used for data mining or user-defined visualization and interpretation. This is the first time that a service is capable of providing access to all cross-project relay data from a single Web resource. Because MaROS contains the data for a variety of missions from the Mars network, which span both NASA and ESA, the software also establishes an access control list (ACL) on each data record in the database repository to enforce user access permissions through a multilayered approach.

Wang, Paul

Data Mining Tools Make Flights Safer, More Efficient

A small data mining team at Ames Research Center developed a set of algorithms ideal for combing through flight data to find anomalies. Dallas-based Southwest Airlines Co. signed a Space Act Agreement with Ames in 2011 to access the tools, helping the company refine its safety practices, improve its safety reviews, and increase flight efficiencies.

Source record

Traffic Flow Management: Data Mining Update

This presentation provides an update on recent data mining efforts that have been designed to (1) identify like/similar days in the national airspace system, (2) cluster/aggregate national-level rerouting data and (3) apply machine learning techniques to predict when Ground Delay Programs are required at a weather-impacted airport

Grabbe, Shon R.

Data Mining of NASA Boeing 737 Flight Data: Frequency Analysis of In-Flight Recorded Data

Data recorded during flights of the NASA Trailblazer Boeing 737 have been analyzed to ascertain the presence of aircraft structural responses from various excitations such as the engine, aerodynamic effects, wind gusts, and control system operations. The NASA Trailblazer Boeing 737 was chosen as a focus of the study because of a large quantity of its flight data records. The goal of this study was to determine if any aircraft structural characteristics could be identified from flight data collected for measuring non-structural phenomena. A number of such data were examined for spatial and frequency correlation as a means of discovering hidden knowledge of the dynamic behavior of the aircraft. Data recorded from on-board dynamic sensors over a range of flight conditions showed consistently appearing frequencies. Those frequencies were attributed to aircraft structural vibrations.

Butterfield, Ansel J.

Data Mining at NASA: From Theory to Applications

This slide presentation demonstrates the data mining/machine learning capabilities of NASA Ames and Intelligent Data Understanding (IDU) group. This will encompass the work done recently in the group by various group members. The IDU group develops novel algorithms to detect, classify, and predict events in large data streams for scientific and engineering systems. This presentation for Knowledge Discovery and Data Mining 2009 is to demonstrate the data mining/machine learning capabilities of NASA Ames and IDU group. This will encompass the work done re cently in the group by various group members.

Srivastava, Ashok N.

Virtual Sensors: Using Data Mining to Efficiently Estimate Spectra

Detecting clouds within a satellite image is essential for retrieving surface geophysical parameters, such as albedo and temperature, from optical and thermal imagery because the retrieval methods tend to be valid for clear skies only. Thus, routine satellite data processing requires reliable automated cloud detection algorithms that are applicable to many surface types. Unfortunately, cloud detection over snow and ice is difficult due to the lack of spectral contrast between clouds and snow. Snow and clouds are both highly reflective in the visible wavelen,ats and often show little contrast in the thermal Infrared. However, at 1.6 microns, the spectral signatures of snow and clouds differ enough to allow improved snow/ice/cloud discrimination. The recent Terra and Aqua Moderate Resolution Imaging Spectro-Radiometer (MODIS) sensors have a channel (channel 6) at 1.6 microns. Presently the most comprehensive, long-term information on surface albedo and temperature over snow- and ice-covered surfaces comes from the Advanced Very High Resolution Radiometer ( AVHRR) sensor that has been providing imagery since July 1981. The earlier AVHRR sensors (e.g. AVHRR/2) did not however have a channel designed for discriminating clouds from snow, such as the 1.6 micron channel available on the more recent AVHRR/3 or the MODIS sensors. In the absence of the 1.6 micron channel, the AVHRR Polar Pathfinder (APP) product performs cloud detection using a combination of time-series analysis and multispectral threshold tests based on the satellite's measuring channels to produce a cloud mask. The method has been found to work reasonably well over sea ice, but not so well over the ice sheets. Thus, improving the cloud mask in the APP dataset would be extremely helpful toward increasing the accuracy of the albedo and temperature retrievals, as well as extending the time-series of albedo and temperature retrievals from the more recent sensors to the historical ones. In this work, we use data mining methods to construct a model of MODIS channel 6 as a function of other channels that are common to both MODIS and AVHRR. The idea is to use the model to generate the equivalent of MODIS channel 6 for AVHRR as a function of the AVHRR equivalents to MODIS channels. We call this a Virtual Sensor because it predicts unmeasured spectra. The goal is to use this virtual channel 6. to yield a cloud mask superior to what is currently used in APP . Our results show that several data mining methods such as multilayer perceptrons (MLPs), ensemble methods (e.g., bagging), and kernel methods (e.g., support vector machines) generate channel 6 for unseen MODIS images with high accuracy. Because the true channel 6 is not available for AVHRR images, we qualitatively assess the virtual channel 6 for several AVHRR images.

Srivastava, Ashok

Data-Mining Toolset Developed for Determining Turbine Engine Part Life Consumption

The current practice in aerospace turbine engine maintenance is to remove components defined as life-limited parts after a fixed time, on the basis of a predetermined number of flight cycles. Under this schedule-based maintenance practice, the worst-case usage scenario is used to determine the usable life of the component. As shown, this practice often requires removing a part before its useful life is fully consumed, thus leading to higher maintenance cost. To address this issue, the NASA Glenn Research Center, in a collaborative effort with Pratt & Whitney, has developed a generic modular toolset that uses data-mining technology to parameterize life usage models for maintenance purposes. The toolset enables a "condition-based" maintenance approach, where parts are removed on the basis of the cumulative history of the severity of operation they have experienced. The toolset uses data-mining technology to tune life-consumption models on the basis of operating and maintenance histories. The flight operating conditions, represented by measured variables within the engine, are correlated with repair records for the engines, generating a relationship between the operating condition of the part and its service life. As shown, with the condition-based maintenance approach, the lifelimited part is in service until its usable life is fully consumed. This approach will lower maintenance costs while maintaining the safety of the propulsion system. The toolset is a modular program that is easily customizable by users. First, appropriate parametric damage accumulation models, which will be functions of engine variables, must be defined. The tool then optimizes the models to match the historical data by computing an effective-cycle metric that reduces the unexplained variability in component life due to each damage mode by accounting for the variability in operational severity. The damage increment due to operating conditions experienced during each flight is used to compute the effective cycles and ultimately the replacement time. Utilities to handle data problems, such as gaps in the flight data records, are included in the toolset. The tool was demonstrated using the first stage, high-pressure turbine blade of the PW4077 engine (Pratt & Whitney, East Hartford, CT). The damage modes considered were thermomechanical fatigue and oxidation/erosion. Each PW4077 engine contains 82 first-stage, high-pressure turbine blades, and data from a fleet of engines were used to tune the life-consumption models. The models took into account not only measured variables within the engine, but also unmeasured variables such as engine health parameters that are affected by degradation of the engine due to aging. The tool proved effective at predicting the average number of blades scrapped over time due to each damage mode, per engine, given the operating history of the engine. The customizable tools are available to interested parties within the aerospace community.

Litt, Jonathan S.

Email-Based Informed Consent: Innovative Method for Reaching Large Numbers of Subjects for Data Mining Research

Since the 2010 NASA authorization to make the Life Sciences Data Archive (LSDA) and Lifetime Surveillance of Astronaut Health (LSAH) data archives more accessible by the research and operational communities, demand for data has greatly increased. Correspondingly, both the number and scope of requests have increased, from 142 requests fulfilled in 2011 to 224 in 2014, and with some datasets comprising up to 1 million data points. To meet the demand, the LSAH and LSDA Repositories project was launched, which allows active and retired astronauts to authorize full, partial, or no access to their data for research without individual, study-specific informed consent. A one-on-one personal informed consent briefing is required to fully communicate the implications of the several tiers of consent. Due to the need for personal contact to conduct Repositories consent meetings, the rate of consenting has not kept up with demand for individualized, possibly attributable data. As a result, other methods had to be implemented to allow the release of large datasets, such as release of only de-identified data. However the compilation of large, de-identified data sets places a significant resource burden on LSAH and LSDA and may result in diminished scientific usefulness of the dataset. As a result, LSAH and LSDA worked with the JSC Institutional Review Board Chair, Astronaut Office physicians, and NASA Office of General Counsel personnel to develop a "Remote Consenting" process for retrospective data mining studies. This is particularly useful since the majority of the astronaut cohort is retired from the agency and living outside the Houston area. Originally planned as a method to send informed consent briefing slides and consent forms only by mail, Remote Consenting has evolved into a means to accept crewmember decisions on individual studies via their method of choice: email or paper copy by mail. To date, 100 emails have been sent to request participation in eight HRP-funded studies. The development of the Remote Consent process, the laws allowing transmission of consent via electronic means, total metrics to date, and remaining challenges (e.g., response issues, use of International Partner data, biospecimens/genetic data) for the research use of LSAH/LSDA data will be described.

Lee, Lesley R.