Search NASASearch

SEARCH · Search NASA

Results for “Feature selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

AHIMSA - Ad hoc histogram information measure sensing algorithm for feature selection in the context of histogram inspired clustering techniques

An algorithm is proposed for dimensionality reduction in the context of clustering techniques based on histogram analysis. The approach is based on an evaluation of the hills and valleys in the unidimensional histograms along the different features and provides an economical means of assessing the significance of the features in a nonparametric unsupervised data environment. The method has relevance to remote sensing applications.

Dasarathy, B. V.

Applications of feature selection

The use of satellite-acquired (LANDSAT) multispectral scanner (MSS) data to conduct an inventory of some crop of economic interest such as wheat over a large geographical area is considered in relation to the development of accurate and efficient algorithms for data classification. The dimension of the measurement space and the computational load for a classification algorithm is increased by the use of multitemporal measurements. Feature selection/combination techniques used to reduce the dimensionality of the problem are described.

Guseman, L. F., Jr.

High resolution spectrophotometry of selected features in the 1.1 micron spectrum of Comet Kohoutek /1973f/

Fabry-Perot interferometry of Comet Kohoutek (1973f) at 1.1 microns with a resolution of 1.2 A showed emission features identified as OH and CN lines in addition to a strong Fraunhofer continuum. Central intensities have been derived for three cases (uniform, Gaussian, and Gaussian plus inverse-rho law) of brightness profiles in the comet coma. Limits for CH4, H2O, HeI, SiI and CrI are also derived.

Meisel, D. D.

Acreage estimation, feature selection, and signature extension dependent upon the maximum likelihood decision rule

A maximum likelihood estimation technique is used for the analysis of agricultural remote sensor data. The m-class probability of misclassification is estimated using unlabeled test samples and labeled training samples. A bound on the variance of a proposed unbiased estimator of the m-class probability of error is derived. The particular case in which each class density is assumed to be a mixture of multivariate normal densities is considered. The extension of spectral signatures in space and time is discussed.

Quirein, J. A.

The role of eigenvalues in linear feature selection theory

A particular measure of pattern class distinction called the average interclass divergence, or more simply, divergence, is considered. Here divergence will be the pairwise average of the expected interclass divergence derived from Hajek's two-class divergence.

Brown, D. R.

The role of eigenvalues in linear feature selection theory

The analysis concerns the role of eigenvalues in determining a particular measure of pattern class distinction called the divergence, which is the pairwise average of the expected interclass divergence derived from Hajek's two-class divergence. Decel and Quirein (1973) showed that there always exists a k x n real matrix B such that the transformation determined by B maximizes divergence in k-dimensional space, and that B can be written as a product involving an orthogonal n x n matrix U. In the present paper it is shown that divergence measure of pattern class distinction does not depend on the eigenvalues of U.

Brown, D. R.

Feature selection via entropy minimization: An example using LANDSAT satellite data

The author has identified the following significant results. The minimum entropy model may provide several useful advantages over traditional techniques for processing LANDSAT data. Total computer time to conduct a complete pattern recognition process is reduced. Subjective (transformed image), as well as statistically derived information is made available to the analyst/user much earlier in the analysis process. A rapid feedback loop in which numerous training set combinations can be tested for difference and representativeness is available. Additional tests of LANDSAT data processing using the minimum entropy model are clearly justified.

Zandonella, A.

Feature selection methodologies using simulated Thematic Mapper data

The present investigation is concerned with the determination of the intrinsic dimensionality of a simulated Thematic Mapper data set. In addition, the effectiveness and sensitivity of 'standard' statistics separability measures (i.e., transformed divergence) is evaluated in comparison to eigenvectors for identifying the optimum subset of the original Thematic Mapper Simulator (TMS) bands for classifying the various cover types. TMS data were collected on May 2, 1979 by NASA's NS001 aircraft multispectral scanner over a bottomland forested area in South Carolina near the city of Camden. It is found that the eigenvectors and eigenvalues of a covariance matrix from a multispectral scanner system (MSS) data set can be obtained without having to actually transform the data.

Dean, M. E.

Interpretable Machine Learning for Molecular Biosignatures: a Novel Single-Sample Feature Importance Method That Is Sensitive To Statistical Interactions

Isotope ratio mass spectrometry (IRMS) of volatiles (e.g., CO 2 ) promises to be a powerful tool for potential biosignature detection for future missions to ocean worlds (OW) such as Europa and Enceladus. Machine learning (ML) methods for IRMS data could enable science autonomy by onboard prediction of seawater chemistry and biosignature presence. However, ML models are likely to be complex and involve statistical interactions between features (variables), which can make predictions seem opaque and enigmatic. For ML predictions as significant as extraterrestrial biosignatures, we must place extraordinary confidence in models. It is therefore essential that these models make interpretable predictions (i.e., human-understandable) and include false-prediction diagnostics. We achieve high accuracy and interpretability in ML biosignature and seawater chemistry models for OW through a nearest-neighbors feature selection tool that detects statistical interactions between predictors, constructs interaction networks for visualization of selected features working together to make a prediction, and reports single-sample feature importance scores for false-detection diagnostics. Here we develop a novel single-sample nearest-neighbors projected distance regression(ssNPDR) feature selection method that improves upon existing single-sample algorithms through the inclusion of statistical interactions while providing false-prediction diagnostics for ML models.

geochemistry

Geographical Insights into Suicide Mortality Through Spatial Machine Learning

Suicide mortality is a leading cause of death in the United States, with an upward trend that emphasizes its significance as a public health issue. Previous research has employed global models like ordinary least squares (OLS) regression and local models such as geographically weighted regression (GWR). While local models are useful for analyzing spatial variations in suicide mortality, they share limitations with traditional global models, particularly about their inability to handle multi-collinearity and non-linear relationships. Machine learning approaches, like random forests (RF), can address some of these limitations but often fail to account for spatial variability. This gap highlights the need for spatial ML models specifically designed to tackle suicide mortality. This research seeks to fill this void by using a geographically weighted random forest model (GWRF) to examine the associations between county-level suicide mortality in the U.S. from 2010 to 2020 and various social and environmental determinants of health. A key aspect of our methodology is disciplined feature selection, which reduces the pool of explanatory variables by about 90%. This refinement enhances the explanatory power of both global (R2 improved from 0.59 to 0.67) and local (R2 improved from 0.64 to 0.67) RF models while reducing their run times. An analysis of the importance scores for these selected features reveals that the drivers of suicide mortality vary by context. Thus, to effectively address regional disparities and inform targeted public health interventions, a holistic approach that incorporates multiple county-level characteristics is essential.

Lebakula, Viswadeep [ORNL] (ORCID:0000000152935914

Virtual refrigerant charge sensing algorithm for residential CO₂ heat pumps

Natural refrigerants are increasingly adopted in next-generation heat pump systems, among which CO₂ heat pumps have attracted significant attention. However, due to their high operating pressures, the leakage risk is higher, resulting in undercharge conditions and degraded heat pump performance. Thus, developing an accurate refrigerant charge level detection technique is necessary to guarantee safe and efficient operation. Although virtual refrigerant charge (VRC) level calculation algorithms for CO₂ heat pumps exist, they typically rely on empirically selected features without a systematic selection framework, leading to multicollinearity and potential overfitting, which limit their prediction accuracy and generalizability. To address these issues, this study proposes a VRC algorithm framework with a systematic feature selection method that identifies physically meaningful and statistically significant features, and is applied using a residential CO₂ heat pump as a case study. The method is extended from previous work on conventional refrigerants to account for charge behavior in CO₂ gas coolers. The selected features include gas cooler outlet density, evaporator pressure, and superheat temperature. The results demonstrate that the proposed feature selection method significantly improves prediction accuracy compared to existing VRC approaches. A relatively small training dataset (∼30 samples) is sufficient for feature identification and model development. The developed algorithm achieves less than 3% prediction error under both undercharge and overcharge conditions, representing reductions of 46.7% and 35.3% compared to two recent reference VRC algorithms for transcritical CO₂ heat pumps reported in the literature. The proposed algorithm and feature selection method enhance leakage detection capability, facilitate the deployment of CO₂ heat pump systems, and contribute to reduced energy waste and maintenance costs.

Guo, Fangzhou [Lawrence Berkeley National Laborato

Input Decimated Ensembles

Using an ensemble of classifiers instead of a single classifier has been shown to improve generalization performance in many pattern recognition problems. However, the extent of such improvement depends greatly on the amount of correlation among the errors of the base classifiers. Therefore, reducing those correlations while keeping the classifiers' performance levels high is an important area of research. In this article, we explore input decimation (ID), a method which selects feature subsets for their ability to discriminate among the classes and uses them to decouple the base classifiers. We provide a summary of the theoretical benefits of correlation reduction, along with results of our method on two underwater sonar data sets, three benchmarks from the Probenl/UCI repositories, and two synthetic data sets. The results indicate that input decimated ensembles (IDEs) outperform ensembles whose base classifiers use all the input features; randomly selected subsets of features; and features created using principal components analysis, on a wide range of domains.

Tumer, Kagan