Search NASASearch

SEARCH · Search NASA

Results for “Feature selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A multi-backend autotuning study of feature selection on GPUs

Abstract Feature selection is an important step in machine learning that can benefit from GPU acceleration. As the number of GPU vendors increases, it is imperative to adapt algorithms such as the minimum Redundancy Maximum Relevance (mRMR) feature selection method to different backends that support several GPU architectures. This work presents a multi-backend implementation of mRMR across CUDA, HIP, and SYCL, and studies its performance when combined with Bayesian optimization and transfer learning to automatically tune execution parameters for different platforms and datasets. Our experimental results show that when tuned, CUDA and HIP achieve comparable performance on NVIDIA architectures, while SYCL exhibits a moderate performance gap. Overall, this work highlights the impact of backend choice and autotuning on GPU-accelerated feature selection and provides insights into deploying mRMR across heterogeneous environments.

Beceiro, Bieito (ORCID:0000000333014890)

Decision boundary feature selection for non-parametric classifier

Feature selection has been one of the most important topics in pattern recognition. Although many authors have studied feature selection for parametric classifiers, few algorithms are available for feature selection for nonparametric classifiers. In this paper we propose a new feature selection algorithm based on decision boundaries for nonparametric classifiers. We first note that feature selection for pattern recognition is equivalent to retaining 'discriminantly informative features', and a discriminantly informative feature is related to the decision boundary. A procedure to extract discriminantly informative features based on a decision boundary for nonparametric classification is proposed. Experiments show that the proposed algorithm finds effective features for the nonparametric classifier with Parzen density estimation.

Lee, Chulhee

Feature Selection for Classification of Polar Regions Using a Fuzzy Expert System

Labeling, feature selection, and the choice of classifier are critical elements for classification of scenes and for image understanding. This study examines several methods for feature selection in polar regions, including the list, of a fuzzy logic-based expert system for further refinement of a set of selected features. Six Advanced Very High Resolution Radiometer (AVHRR) Local Area Coverage (LAC) arctic scenes are classified into nine classes: water, snow / ice, ice cloud, land, thin stratus, stratus over water, cumulus over water, textured snow over water, and snow-covered mountains. Sixty-seven spectral and textural features are computed and analyzed by the feature selection algorithms. The divergence, histogram analysis, and discriminant analysis approaches are intercompared for their effectiveness in feature selection. The fuzzy expert system method is used not only to determine the effectiveness of each approach in classifying polar scenes, but also to further reduce the features into a more optimal set. For each selection method,features are ranked from best to worst, and the best half of the features are selected. Then, rules using these selected features are defined. The results of running the fuzzy expert system with these rules show that the divergence method produces the best set features, not only does it produce the highest classification accuracy, but also it has the lowest computation requirements. A reduction of the set of features produced by the divergence method using the fuzzy expert system results in an overall classification accuracy of over 95 %. However, this increase of accuracy has a high computation cost.

Penaloza, Mauel A.

On minimizing the probability of misclassification for linear feature selection

The use of techniques for feature selection permits treatment of classification problems in spaces of reduced dimensions. A method is considered of linear feature selection for n-dimensional observation vectors which belong to one of two populations, where each population is described by a known multivariate normal density function. More specifically, the problem of finding a 1xn transformation matrix B for which the probability of misclassification with respect to the one-dimensional transformed density functions was minimized was considered. Theoretical results are presented which give rise to a numerically tractable expression for the variation in the probability of misclassification with respect to B. Using this expression a computational procedure is discussed for obtaining a B which minimizes the probability of misclassification. Preliminary numerical results are discussed.

Guseman, L. F., Jr.

Feature selection for neural networks using Parzen density estimator

A feature selection method for neural networks is proposed using the Parzen density estimator. A new feature set is selected using the decision boundary feature selection algorithm. The selected feature set is then used to train a neural network. Using a reduced feature set, an attempt is made to reduce the training time of the neural network and obtain a simpler neural network, which further reduces the classification time for test data.

Lee, Chulhee

A counter example in linear feature selection theory

The linear feature selection problem in multi-class pattern recognition is described as that of linearly transforming statistical information from n-dimensional (real Euclidean) space into k-dimensional space, while requiring that average interclass divergence in the transformed space decrease as little as possible. Divergence is the expected interclass divergence derived from Hajek two-class divergence; it is known that there always exists a k x n matrix B such that the transformation determined by B maximizes the divergence in k-dimensional space. It is known that, if Q is any k x k invertible matrix, and B is as defined above, then QB again maximizes the divergence in k-space. It is shown that the converse of this result is false: two matrices exist, B sub 1 and B sub 2, each of which maximizes transformed divergence, which are not related in the fashion B sub 2 = QB sub 1 for any k x k matrix Q.

Brown, D. R.

Linear feature selection with applications

Several ways in which feature selection techniques were used in LACIE are discussed. In all cases, the methods require some a priori information and assumptions; in most, the classification procedure (Bayes optimal) was chosen in advance. The transformations used for dimensionality reduction are linear, that is, the variables in feature space are always linear combinations of the original measurements. Several numerically tractable criteria developed for LACIE, which provide information about the probability of misclassification, are discussed. Recent results on linear feature selection techniques are included. Their use in LACIE is discussed. Related open questions are mentioned.

Decell, H. P., Jr.

On differentiating the probability of error in multipopular feature selection

A method of linear feature selection for n dimensional observation vectors which belong to one of m populations is presented. Each population has a known apriori probability and is described by a known multivariate normal density function. Specifically we consider the problem of finding a k x n matrix B of rank k (k n) for which the transformed probability of misclassification is minimized. Providing that the transformed a posterior probabilities are distinct theoretical results are obtained which, for the case k = l, give rise to a numerically tractable formula for the derivative of the probability of misclassification. It is shown that for the two population problem this condition is also necessary. The dependence of the minimum probability of error on the a priori probabilities is investigated. The minimum probability of error satisfies a uniform Lipschitz condition with respect to the a priori probabilities.

Peters, B. C.

Virtual refrigerant charge sensor for variable-speed heat pumps based on feature selection

The refrigerant charge level in heat pump systems significantly impacts their energy efficiency. Virtual refrigerant charge (VRC) sensing technology has been comprehensively investigated and well-established due to its lower cost compared to physical sensors. However, the previous VRC research often relied on expert judgment and physical reasoning for their variable selection, which can potentially select redundant (or highly correlated) or insignificant features, and it is also primarily focused on single-speed systems. To address these challenges, this study proposes a VRC algorithm for variable-speed heat pumps that selects features through a rigorous feature selection method in combination with physical insights. We also propose a piecewise linear model structure segmented by subcooling temperature to accurately predict charge levels, particularly when subcooling temperatures are substantially low. The proposed algorithm was evaluated using experimental data of a residential R410A heat pump, and the performance was compared with two baseline VRC algorithms. The results are: (1) The proposed algorithm outperforms for the case with subcooling temperature less than 1 °C. (2) The proposed algorithm achieves a tested mean absolute percentage error (MAPE) of 4.23%, and improves the overall accuracy for cooling conditions by approximately 60%, compared with the two baseline algorithms. (3) The proposed algorithm uses two fewer features and improves the accuracy for undercharge cooling conditions by 68.0%, compared with baseline algorithm 2. These improvements enhance prediction accuracy and prevent overfitting, providing a more reliable refrigerant charge level prediction and helping improve the heat pump energy efficiency.

Liang, Chenjiyu

Two effective feature selection criteria for multispectral remote sensing

Distance measures which are useful for feature selection are considered, giving attention to the divergence distance measure and the Jeffreys-Matusita (JM) distance. Experimental studies show that the JM-distance yields more reliable results than other distance measures. A number of questions which are not solved by the experiments are discussed. The investigation provides an explanation for previous observations that the JM-distance and a saturating transform of divergence are highly useful for feature selection in the multiclass case.

Swain, P. H.

Evaluation of entropy and JM-distance criterions as features selection methods using spectral and spatial features derived from LANDSAT images

A study area near Ribeirao Preto in Sao Paulo state was selected, with predominance in sugar cane. Eight features were extracted from the 4 original bands of LANDSAT image, using low-pass and high-pass filtering to obtain spatial features. There were 5 training sites in order to acquire the necessary parameters. Two groups of four channels were selected from 12 channels using JM-distance and entropy criterions. The number of selected channels was defined by physical restrictions of the image analyzer and computacional costs. The evaluation was performed by extracting the confusion matrix for training and tests areas, with a maximum likelihood classifier, and by defining performance indexes based on those matrixes for each group of channels. Results show that in spatial features and supervised classification, the entropy criterion is better in the sense that allows a more accurate and generalized definition of class signature. On the other hand, JM-distance criterion strongly reduces the misclassification within training areas.

Parada, N. D. J.

Entropy-based feature selection for capturing impacts in Earth system models with abrupt forcing

This paper presents the development of a new entropy-based feature selection method for identifying and quantifying impacts. Here, impacts are defined as statistically significant differences in spatio-temporal fields when comparing datasets with and without an external forcing in an Earth system model. Temporal feature selection is performed by first computing the cross-fuzzy entropy to quantify similarity of patterns between two datasets and then applying changepoint detection to identify regions of statistically constant entropy. The method is used to capture temperate north surface cooling from a 9-member simulation ensemble of the Mt. Pinatubo volcanic eruption, which injected 10 Tg of SO 2 into the stratosphere. The results estimate a mean difference decrease in near surface air temperature of -0.560 K with a 99% confidence interval between -0.864 K and -0.257 K between April and November of 1992, one year following the eruption. A sensitivity analysis with decreasing SO 2 injection revealed that the impact is statistically significant at 5 Tg but not at 3 Tg. Using identified features, a dependency graph model based on a 9-day lag had significantly fewer nodes than a graph based on monthly means. Furthermore, this demonstrates our method’s ability to perform dimension reduction while still uncovering source-to-impact pathways.

Changepoint detection

An Iterative Approach to the Feature Selection Problem

The problem dealt with concerns feature selection or reducing the dimension of the data to be processed from n to k. By reducing the dimension of the data from n to k, classification time is generally reduced. Yet the dimension reduction should not be so great that classification accuracy is impaired. Thus, the general problem is considered of classifying an n-dimensional observation vector x into one of m-distinct classes where each class is normally distributed with mean and covariance. It is shown that the probability of misclassification is minimized if a maximum likelihood classification procedure is used to classify the data. The dimension of each observation vector to be processed is conveniently reduced by performing the transformation y = Bx, where B is a K by n matrix of rank k. Thus, the n-dimensional classification problem transforms into a k-dimensional classification problem.

Decell, H. P., Jr.

Choice: 36 band feature selection software with applications to multispectral pattern recognition

Feature selection software was developed at the Earth Resources Laboratory that is capable of inputting up to 36 channels and selecting channel subsets according to several criteria based on divergence. One of the criterion used is compatible with the table look-up classifier requirements. The software indicates which channel subset best separates (based on average divergence) each class from all other classes. The software employs an exhaustive search technique, and computer time is not prohibitive. A typical task to select the best 4 of 22 channels for 12 classes takes 9 minutes on a Univac 1108 computer.

Jones, W. C.