Walsh functions in image processing, feature selection and pattern recognition
Walsh functions in image processing, rotational feature selection and pattern recognition, defining set of orthogonal transformations
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Walsh functions in image processing, rotational feature selection and pattern recognition, defining set of orthogonal transformations
Mass spectrometry (MS) promises to be a powerful tool for potential biosignature detection during astrobiological missions on ocean worlds in our solar system. Accurate and generalizable machine learning methods could enhance science return on investment by predicting seawater chemistry and classifying isotopic biosignatures, either as a signature consistent with microbial life (biotic) or as a novelty (unclassified/unique). However, machine learning models are likely to be complex and involve interactions between MS features, making biosignatures difficult to interpret. Feature selection methods provide biological and chemical context that help interpret the mechanisms of machine learning models, but these methods also need the ability to detect complex interactions. Previously, we developed a machine learning feature selection algorithm called nearest-neighbor projected distance regression (NPDR) that has the ability to identify important model features that involve complex interactions and automatically reduce correlation and the dimensionality in a high-dimensional variable space. The standard distance metrics used in NPDR – Manhattan and Euclidean – assume the multivariate data are isotropic, which is often violated in real data due to differences in the covariance between variables. Thus, we extend NPDR to include a random forest distance, and other anisotropic distance metrics, for computing nearest neighbors. We also augment the isotope-ratio MS data with time-series features from the raw MS signal to improve biotic classification. We test NPDR on our novel experimental ocean world seawater analog MS data. We measure isotope fractionations of volatile CO 2 that could be measured in exospheres or plumes. Samples include baseline abiotic conditions using a range of possible seawater chemistry consistent with Europa and Enceladus, and biotic samples that include microbes in these seawaters. We use penalized NPDR with random forest proximity to identify interpretable microbial molecular signatures. We compare features with random forest importance, and we train a classifier that discriminates between biotic and abiotic samples with high accuracy. These ML-trained ocean-world analog MS data could be used to assist in identifying biosignatures during future missions.
Because of the continually changing environment of a space station, visual feedback is a vital element of a telerobotic system. A real time visual servoing system would allow a telerobot to track and manipulate randomly moving objects. Methodologies for the automatic selection of image features to be used to visually control the relative position between an eye-in-hand telerobot and a known object are devised. A weighted criteria function with both image recognition and control components is used to select the combination of image features which provides the best control. Simulation and experimental results of a PUMA robot arm visually tracking a randomly moving carburetor gasket with a visual update time of 70 milliseconds are discussed.
A criterion for linear feature selection is proposed which is based on mean square apporximation of class density functions. It is shown that for the widest possible class of approximants, the criterion reduces to Devijver's Bayesian distance. For linear approximants the criterion is equivalent to well known generalized Fisher criteria.
Feature selection and the information content of Thematic Mapper Simulator (TMS) data are investigated for a forested region in northern Idaho. The optimal TMS channels for forest structural characteristics are determined, and the capability of TMS data to describe the structural variability within a forest stand is evaluated. The comparative performance of TMS and MSS data to discriminate forest structural factors using per-pixel maximum likelihood classification is examined, and four optimal TMS channels are classified in order to ascertain if the full complement of TM channels provide higher accuracies than the four optimal ones.
Accurate estimation of the State of Health (SOH) for second-life batteries (SLBs) is crucial given their increasing use in energy storage applications. Precise SOH prediction is essential for safe operation and robust battery management systems. A major challenge is the limited availability of datasets for building reliable degradation models. To address this, synthetic data generation through linear interpolation is performed to extend the available data, making it more representative of real-world battery operating conditions. By analyzing feature correlation with SOH, the most relevant features are selected for the model. The proposed approach employs a convolutional neural network (CNN) model trained on this interpolated, feature-selected dataset, using time series data of voltage, temperature, and current over a cycle. By focusing on highly correlated features, the model achieves over 95% accuracy, with mean absolute error and root mean squared error up to 2.27% and 2.64%, respectively, in SOH estimation for two battery datasets tested. These results highlight the potential of combining synthetic data generation and feature selection to enhance SOH predictions, showcasing the superior performance of the proposed CNN model for both new batteries and SLBs.
Results that suggest the possibility of using a sequential monotone process for solving the feature selection problem using Householder transformations are applied to the divergence separability criterion and an expression for the gradient of the divergence with respect to the generator of a single Householder transformation will be developed. This expression for the gradient is used in any number of differential correction schemes (iterators) that attempt to extremize the divergence. Data sets provided by the Earth Observations Division-JSC are used to demonstrate selecting the Householder transformations that generate the kxn matrix defining the best (in the sense of extremizing the divergence) k linear combinations of features. The tests allow initial comparisons to be made with results. In particular, this technique does not appear to require initial guesses for the iterator to be generated without replacement, exhaustive search, or other similar schemes.
The computational procedure and associated computer program for a linear feature selection technique are presented. The technique assumes that: a finite number, m, of classes exists; each class is described by an n-dimensional multivariate normal density function of its measurement vectors; the mean vector and covariance matrix for each density function are known (or can be estimated); and the a priori probability for each class is known. The technique produces a single linear combination of the original measurements which minimizes the one-dimensional probability of misclassification defined by the transformed densities.
A new clustering algorithm is presented that is based on dimensional information. The algorithm includes an inherent feature selection criterion, which is discussed. Further, a heuristic method for choosing the proper number of intervals for a frequency distribution histogram, a feature necessary for the algorithm, is presented. The algorithm, although usable as a stand-alone clustering technique, is then utilized as a global approximator. Local clustering techniques and configuration of a global-local scheme are discussed, and finally the complete global-local and feature selector configuration is shown in application to a real-time adaptive classification scheme for the analysis of remote sensed multispectral scanner data.
A condition for the Gateaux differentiability of the probability of misclassification as a function of a feature selection matrix B, assuming a maximum likelihood classifier and normally distributed populations, is given. It is also shown that if the probability of error has a local minimum at B then it is differentiable at B.
The B-average divergence for m-distinct classes, resulting from the linear transformation y = Bx, is proposed as a feature selection criterion, where B is a k by n matrix of rank k not greater than n. It is shown that if the B-average divergence resulting from B is large enough, then the probability of misclassification, considered as a function f the class of all k by n matrices, is essentially minimized by B. A computer program, utilizing a gradient procedure, is developed to numerically maximize the B-average divergence and results are presented for the Cl flight line. For this example, corresponding to 9-distinct classes, most of the discriminatory information is found to lie in a 3-dimensional subspace, defined by an appropriately chosen 3 by 12 matrix B.
Variational equations are presented for maximizing the probability of correct classification as a function of a 1xn feature selection matrix B for the two-population problem. For the special case of equal covariance matrices the optimal B is unique up to scalar multiples and rank one sufficient. For equal population means, the best 1xn B is an eigenvector corresponding either to the largest or smallest eigenvalue of sigma sub 2 to the minus 1 power sigma sub 1 where sigma sub 1 and sigma sub 2 are the nxn covariance matrices of the two populations. The transformed probability of correct classification depends only on the eigenvalue. Finally, a procedure is proposed for constructing an optimal or nearly optimal kxn matrix of rank k without solving the k-dimensional variational equation.
The paper shows that it is possible to construct two k x n matrices, both of which maximize divergence in the transformed space of the linear feature selection problem in multiclass pattern recognition, and which are not row equivalent. Thus, even under extremely strong conditions, it is not possible to assume that all matrix solutions which maximize transformed divergence are row equivalent.
Energy discretization is a crucial component of deterministic neutron transport simulations. Metaheuristic (MH) optimizers are effective algorithms to determine group structures that maximize both solution accuracy and computational efficiency. This project establishes a framework for optimizing group structures for PARTISN simulations using the Python library MEALPY. Group structure optimization is formulated as a binary feature selection problem, and results are investigated with permutation and material importance techniques to determine physically relevant energy bounds. We conclude that MH optimizers find group structures that drastically improve flux calculations while preserving k-effective accuracy. Further, we find that individual energy bounds are not necessarily physically relevant, but rather specific energy ranges are.
A Technology Benefit Estimator (T/BEST) system has been developed to provide a formal method to assess advanced aerospace technologies and quantify the benefit contributions for prioritization. An open-ended, modular approach is used to allow for upgrade and insertion of advanced technology modules. T/BEST's software framework, beginner-to-expert operation, interface architecture, and key analysis modules are discussed. In this paper, selected features and applications of T/BEST are demonstrated. Sample cases pertaining to structural analysis of titanium and composite blades are presented. The performance of hot and cold composite fan blades is also discussed. The cost required to manufacture titanium and composite fan blades is estimated.
The use of a Monte Carlo model for generating sample directional reflectance data for two simplified target canopies at two different solar positions is reported. Successive iterations through the model permit the calculation of a mean vector and covariance matrix for canopy reflectance for varied sensor view angles. These data may then be used to calculate the divergence between the target distributions for various wavelength combinations and for these view angles. Results of a feature selection analysis indicate that different sets of wavelengths are optimum for target discrimination depending on sensor view angle and that the targets may be more easily discriminated for some scan angles than others. The time-varying behavior of these results is also pointed out.
An assessment is made of the information content of Thematic Mapper Simulator (TMS) data for the case of a forested region, in order to determine the sensitivity of such data to forest crown closure and tree size class. Principal components analysis and Monte Carlo simulation indicated that channels 4, 7, 5 and 3 were optimal for four-channel forest structure analysis. As the number of channels supplied to the Monte Carlo feature selection routine increased, classification accuracy increased. The greatest sensitivity to the forest structural parameters, which included succession within clearcuts as well as crown closure and size class, was obtained from the 7-channel TMS data.
Selected lunar topographic features by photometry with computer data reduction at 23 phase angles