Search NASASearch

SEARCH · Search NASA

Results for “Feature selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

The role of eigenvalues in linear feature selection theory

A particular measure of pattern class distinction called the average interclass divergence, or more simply, divergence, is considered. Here divergence will be the pairwise average of the expected interclass divergence derived from Hajek's two-class divergence.

Brown, D. R.

The role of eigenvalues in linear feature selection theory

The analysis concerns the role of eigenvalues in determining a particular measure of pattern class distinction called the divergence, which is the pairwise average of the expected interclass divergence derived from Hajek's two-class divergence. Decel and Quirein (1973) showed that there always exists a k x n real matrix B such that the transformation determined by B maximizes divergence in k-dimensional space, and that B can be written as a product involving an orthogonal n x n matrix U. In the present paper it is shown that divergence measure of pattern class distinction does not depend on the eigenvalues of U.

Brown, D. R.

Feature selection via entropy minimization: An example using LANDSAT satellite data

The author has identified the following significant results. The minimum entropy model may provide several useful advantages over traditional techniques for processing LANDSAT data. Total computer time to conduct a complete pattern recognition process is reduced. Subjective (transformed image), as well as statistically derived information is made available to the analyst/user much earlier in the analysis process. A rapid feedback loop in which numerous training set combinations can be tested for difference and representativeness is available. Additional tests of LANDSAT data processing using the minimum entropy model are clearly justified.

Zandonella, A.

Feature selection methodologies using simulated Thematic Mapper data

The present investigation is concerned with the determination of the intrinsic dimensionality of a simulated Thematic Mapper data set. In addition, the effectiveness and sensitivity of 'standard' statistics separability measures (i.e., transformed divergence) is evaluated in comparison to eigenvectors for identifying the optimum subset of the original Thematic Mapper Simulator (TMS) bands for classifying the various cover types. TMS data were collected on May 2, 1979 by NASA's NS001 aircraft multispectral scanner over a bottomland forested area in South Carolina near the city of Camden. It is found that the eigenvectors and eigenvalues of a covariance matrix from a multispectral scanner system (MSS) data set can be obtained without having to actually transform the data.

Dean, M. E.

Interpretable Machine Learning for Molecular Biosignatures: a Novel Single-Sample Feature Importance Method That Is Sensitive To Statistical Interactions

Isotope ratio mass spectrometry (IRMS) of volatiles (e.g., CO 2 ) promises to be a powerful tool for potential biosignature detection for future missions to ocean worlds (OW) such as Europa and Enceladus. Machine learning (ML) methods for IRMS data could enable science autonomy by onboard prediction of seawater chemistry and biosignature presence. However, ML models are likely to be complex and involve statistical interactions between features (variables), which can make predictions seem opaque and enigmatic. For ML predictions as significant as extraterrestrial biosignatures, we must place extraordinary confidence in models. It is therefore essential that these models make interpretable predictions (i.e., human-understandable) and include false-prediction diagnostics. We achieve high accuracy and interpretability in ML biosignature and seawater chemistry models for OW through a nearest-neighbors feature selection tool that detects statistical interactions between predictors, constructs interaction networks for visualization of selected features working together to make a prediction, and reports single-sample feature importance scores for false-detection diagnostics. Here we develop a novel single-sample nearest-neighbors projected distance regression(ssNPDR) feature selection method that improves upon existing single-sample algorithms through the inclusion of statistical interactions while providing false-prediction diagnostics for ML models.

geochemistry

Input Decimated Ensembles

Using an ensemble of classifiers instead of a single classifier has been shown to improve generalization performance in many pattern recognition problems. However, the extent of such improvement depends greatly on the amount of correlation among the errors of the base classifiers. Therefore, reducing those correlations while keeping the classifiers' performance levels high is an important area of research. In this article, we explore input decimation (ID), a method which selects feature subsets for their ability to discriminate among the classes and uses them to decouple the base classifiers. We provide a summary of the theoretical benefits of correlation reduction, along with results of our method on two underwater sonar data sets, three benchmarks from the Probenl/UCI repositories, and two synthetic data sets. The results indicate that input decimated ensembles (IDEs) outperform ensembles whose base classifiers use all the input features; randomly selected subsets of features; and features created using principal components analysis, on a wide range of domains.

Tumer, Kagan

Efficient feature subset selection with probabilistic distance criteria

Recursive expressions are derived for efficiently computing the commonly used probabilistic distance measures as a change in the criteria both when a feature is added to and when a feature is deleted from the current feature subset. A combinatorial algorithm for generating all possible r feature combinations from a given set of s features in (s/r) steps with a change of a single feature at each step is presented. These expressions can also be used for both forward and backward sequential feature selection.

Chittineni, C. B.

Optimal selection of passes

Preliminary numerical results obtained from the application of a linear feature selection technique to the determination of combinations of passes which best discriminate between a given set of crops in a given area of interest, are reported. The results obtained are not purported to hold in a general situation, but only for the given set of crops and the given, but unknown, levels of several factors-such as soil type, and fertilizer practice, holding in the area of interest. However, by identifying the various factors affecting the spectral signatures, and by formulating a regression model one could use the feature selection technique to determine the regression coefficients for predicting optimal passes for a given set of crops. Another use of the feature selection technique as applied to multiple pass registered data is the generation of enhanced grey scale displays by using a single linear combination of all channels of all designated passes as opposed to a single channel within a single pass.

Guseman, L. F., Jr.

Spectral feature design in high dimensional multispectral data

The High resolution Imaging Spectrometer (HIRIS) is designed to acquire images simultaneously in 192 spectral bands in the 0.4 to 2.5 micrometers wavelength region. It will make possible the collection of essentially continuous reflectance spectra at a spectral resolution sufficient to extract significantly enhanced amounts of information from return signals as compared to existing systems. The advantages of such high dimensional data come at a cost of increased system and data complexity. For example, since the finer the spectral resolution, the higher the data rate, it becomes impractical to design the sensor to be operated continuously. It is essential to find new ways to preprocess the data which reduce the data rate while at the same time maintaining the information content of the high dimensional signal produced. Four spectral feature design techniques are developed from the Weighted Karhunen-Loeve Transforms: (1) non-overlapping band feature selection algorithm; (2) overlapping band feature selection algorithm; (3) Walsh function approach; and (4) infinite clipped optimal function approach. The infinite clipped optimal function approach is chosen since the features are easiest to find and their classification performance is the best. After the preprocessed data has been received at the ground station, canonical analysis is further used to find the best set of features under the criterion that maximal class separability is achieved. Both 100 dimensional vegetation data and 200 dimensional soil data were used to test the spectral feature design system. It was shown that the infinite clipped versions of the first 16 optimal features had excellent classification performance. The overall probability of correct classification is over 90 percent while providing for a reduced downlink data rate by a factor of 10.

Chen, Chih-Chien Thomas

Phase 1 of the earth resources data analysis program

Research completed in the Earth Resources Data Analysis Program is discussed along with recommendations for future study. Projects discussed include use of the Cholesky decomposition in feature selection and classification algorithms; optimal feature selection and extraction, probability density estimation and nonparametric classifiers; use of spatial information in classification; and model for crop row reflectance. The installation of LARSYS on the ICSA's IBM 370/155 is discussed, and a list of technical reports is included.

Source record

Multivariate Methods for Prediction of Geologic Sample Composition with Laser-Induced Breakdown Spectroscopy

Laser-induced breakdown spectroscopy (LIBS) uses pulses of laser light to ablate a material from the surface of a sample and produce an expanding plasma. The optical emission from the plasma produces a spectrum which can be used to classify target materials and estimate their composition. The ChemCam instrument on the Mars Science Laboratory (MSL) mission will use LIBS to rapidly analyze targets remotely, allowing more resource- and time-intensive in-situ analyses to be reserved for targets of particular interest. ChemCam will also be used to analyze samples that are not reachable by the rover's in-situ instruments. Due to these tactical and scientific roles, it is important that ChemCam-derived sample compositions are as accurate as possible. We have compared the results of partial least squares (PLS), multilayer perceptron (MLP) artificial neural networks (ANNs), and cascade correlation (CC) ANNs to determine which technique yields better estimates of quantitative element abundances in rock and mineral samples. The number of hidden nodes in the MLP ANNs was optimized using a genetic algorithm. The influence of two data preprocessing techniques were also investigated: genetic algorithm feature selection and averaging the spectra for each training sample prior to training the PLS and ANN algorithms. We used a ChemCam-like laboratory stand-off LIBS system to collect spectra of 30 pressed powder geostandards and a diverse suite of 196 geologic slab samples of known bulk composition. We tested the performance of PLS and ANNs on a subset of these samples, choosing to focus on silicate rocks and minerals with a loss on ignition of less than 2 percent. This resulted in a set of 22 pressed powder geostandards and 80 geologic samples. Four of the geostandards were used as a validation set and 18 were used as the training set for the algorithms. We found that PLS typically resulted in the lowest average absolute error in its predictions, but that the optimized MLP ANN and the CC ANN often gave results comparable to PLS. Averaging the spectra for each training sample and/or using feature selection to choose a small subset of wavelengths to use for predictions gave mixed results, with degraded performance in some cases and similar or slightly improved performance in other cases. However, training time was significantly reduced for both PLS and ANN methods by implementing feature selection, making this a potentially appealing method for initial, rapid-turn-around analyses necessary for Chemcam's tactical role on MSL. Choice of training samples has a strong influence on the accuracy of predictions. We are currently investigating the use of clustering algorithms (e.g. k-means, neural gas, etc.) to identify training sets that are spectrally similar to the unknown samples that are being predicted, and therefore result in improved predictions

Morris, Richard

Supervised Learning Applied to Air Traffic Trajectory Classification

Given the recent increase of interest in introducing new vehicle types and missions into the National Airspace System, a transition towards a more autonomous air traffic control system is required in order to enable and handle increased density and complexity. This paper presents an exploratory effort of the needed autonomous capabilities by exploring supervised learning techniques in the context of aircraft trajectories. In particular, it focuses on the application of machine learning algorithms and neural network models to a runway recognition trajectory-classification study. It investigates the applicability and effectiveness of various classifiers using datasets containing trajectory records for a month of air traffic. A feature importance and sensitivity analysis are conducted to challenge the chosen time-based datasets and the ten selected features. The study demonstrates that classification accuracy levels of 90% and above can be reached in less than 40 seconds of training for most machine learning classifiers when one track data point, described by the ten selected features at a particular time step, per trajectory is used as input. It also shows that neural network models can achieve similar accuracy levels but at higher training time costs.

Bosson, Christabelle S.

Supervised Learning Applied to Air Traffic Trajectory Classification

Given the recent increase of interest in introducing new vehicle types and missions into the National Airspace System, a transition towards a more autonomous air traffic control system is required in order to enable and handle increased density and complexity. This paper presents an exploratory effort of the needed autonomous capabilities by exploring supervised learning techniques in the context of aircraft trajectories. In particular, it focuses on the application of machine learning algorithms and neural network models to a runway recognition trajectory-classification study. It investigates the applicability and effectiveness of various classifiers using datasets containing trajectory records for a month of air traffic. A feature importance and sensitivity analysis are conducted to challenge the chosen time-based datasets and the ten selected features. The study demonstrates that classification accuracy levels of 90% and above can be reached in less than 40 seconds of training for most machine learning classifiers when one track data point, described by the ten selected features at a particular time step, per trajectory is used as input. It also shows that neural network models can achieve similar accuracy levels but at higher training time costs.

Bosson, Christabelle

Using Whispering-Gallery-Mode Resonators for Refractometry

A method of determining the refractive and absorptive properties of optically transparent materials involves a combination of theoretical and experimental analysis of electromagnetic responses of whispering-gallery-mode (WGM) resonator disks made of those materials. The method was conceived especially for use in studying transparent photorefractive materials, for which purpose this method affords unprecedented levels of sensitivity and accuracy. The method is expected to be particularly useful for measuring temporally varying refractive and absorptive properties of photorefractive materials at infrared wavelengths. Still more particularly, the method is expected to be useful for measuring drifts in these properties that are so slow that, heretofore, the properties were assumed to be constant. The basic idea of the method is to attempt to infer values of the photorefractive properties of a material by seeking to match (1) theoretical predictions of the spectral responses (or selected features thereof) of a WGM of known dimensions made of the material with (2) the actual spectral responses (or selected features thereof). Spectral features that are useful for this purpose include resonance frequencies, free spectral ranges (differences between resonance frequencies of adjacently numbered modes), and resonance quality factors (Q values). The method has been demonstrated in several experiments, one of which was performed on a WGM resonator made from a disk of LiNbO3 doped with 5 percent of MgO. The free spectral range of the resonator was approximately equal to 3.42 GHz at wavelengths in the vicinity of 780 nm, the smallest full width at half maximum of a mode was approximately equal to 50 MHz, and the thickness of the resonator in the area of mode localization was 30 microns. In the experiment, laser power of 9 mW was coupled into the resonator with an efficiency of 75 percent, and the laser was scanned over a frequency band 9 GHz wide at a nominal wavelength of approximately equal to 780 nm. Resonance frequencies were measured as functions of time during several hours of exposure to the laser light. The results of these measurements, plotted in the figure, show a pronounced collective frequency drift of the resonator modes. The size of the drift has been estimated to correspond to a change of 8.5 x 10(exp -5) in the effective ordinary index of refraction of the resonator material.

Matsko, Andrey

Two dimensional convolute integers for machine vision and image recognition

Machine vision and image recognition require sophisticated image processing prior to the application of Artificial Intelligence. Two Dimensional Convolute Integer Technology is an innovative mathematical approach for addressing machine vision and image recognition. This new technology generates a family of digital operators for addressing optical images and related two dimensional data sets. The operators are regression generated, integer valued, zero phase shifting, convoluting, frequency sensitive, two dimensional low pass, high pass and band pass filters that are mathematically equivalent to surface fitted partial derivatives. These operators are applied non-recursively either as classical convolutions (replacement point values), interstitial point generators (bandwidth broadening or resolution enhancement), or as missing value calculators (compensation for dead array element values). These operators show frequency sensitive feature selection scale invariant properties. Such tasks as boundary/edge enhancement and noise or small size pixel disturbance removal can readily be accomplished. For feature selection tight band pass operators are essential. Results from test cases are given.

Edwards, Thomas R.