Search NASA⌕ Search

SEARCH · Search NASA

Results for “linear discriminant analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Linear discriminant analysis with misallocation in training samples

Linear discriminant analysis for a two-class case is studied in the presence of misallocation in training samples. A general appraoch to modeling of mislocation is formulated, and the mean vectors and covariance matrices of the mixture distributions are derived. The asymptotic distribution of the discriminant boundary is obtained and the asymptotic first two moments of the two types of error rate given. Certain numerical results for the error rates are presented by considering the random and two non-random misallocation models. It is shown that when the allocation procedure for training samples is objectively formulated, the effect of misallocation on the error rates of the Bayes linear discriminant rule can almost be eliminated. If, however, this is not possible, the use of Fisher rule may be preferred over the Bayes rule.

Chhikara, R.↗

An automated land-use mapping comparison of the Bayesian maximum likelihood and linear discriminant analysis algorithms

The Bayesian maximum likelihood parametric classifier has been tested against the data-based formulation designated 'linear discrimination analysis', using the 'GLIKE' decision and "CLASSIFY' classification algorithms in the Landsat Mapping System. Identical supervised training sets, USGS land use/land cover classes, and various combinations of Landsat image and ancilliary geodata variables, were used to compare the algorithms' thematic mapping accuracy on a single-date summer subscene, with a cellularized USGS land use map of the same time frame furnishing the ground truth reference. CLASSIFY, which accepts a priori class probabilities, is found to be more accurate than GLIKE, which assumes equal class occurrences, for all three mapping variable sets and both levels of detail. These results may be generalized to direct accuracy, time, cost, and flexibility advantages of linear discriminant analysis over Bayesian methods.

Tom, C. H.↗

Advanced microwave soil moisture studies

Comparisons of low level L-band brightness temperature (TB) and thermal infrared (TIR) data as well as the following data sets: soil map and land cover data; direct soil moisture measurement; and a computer generated contour map were statistically evaluated using regression analysis and linear discriminant analysis. Regression analysis of footprint data shows that statistical groupings of ground variables (soil features and land cover) hold promise for qualitative assessment of soil moisture and for reducing variance within the sampling space. Dry conditions appear to be more conductive to producing meaningful statistics than wet conditions. Regression analysis using field averaged TB and TIR data did not approach the higher sq R values obtained using within-field variations. The linear discriminant analysis indicates some capacity to distinguish categories with the results being somewhat better on a field basis than a footprint basis.

Dalsted, K. J.↗

Linear Discriminant Analysis-Based Machine Learning and All-Atom Molecular Dynamics Simulations for Probing Electro-Osmotic Transport in Cationic-Polyelectrolyte-Brush-Grafted Nanochannels

Deciphering the correct mechanisms governing certain phenomena in polyelectrolyte (PE) brush grafted systems, revealed through atomistic simulations, is an extremely challenging problem. In a recent study, our all-atom molecular dynamics (MD) simulations revealed a non-linearly large electroosmotic (EOS) flow (in the presence of an applied electric field) in nanochannels grafted with PMETAC [Poly(2-(methacryloyloxy)ethyl trimethylammonium chloride] brushes. Given the lack of any formal procedure that would have directed us to identify the correct factors responsible for such an occurrence, we needed to spend several months and devote significant analyses to unravel the involved mechanisms. In this paper, we propose a Linear Discriminant Analysis (LDA) based Machine Learning (ML) approach to address this gap. At first, we obtain data on certain basic features from the all-atom MD data. These basic features represent the number of atoms of certain species around one atom of another (or same) species. Here, we obtain such data on basic features for a reference case (case of an EOS flow in PMETAC-brush-grafted nanochannels with a smaller electric field) and a perturbed case (case of an EOS flow in PMETAC-brush-grafted nanochannels with a larger electric field) in bins in which the nanochannel half height has been divided into. These datasets are high-dimensional dataset, to which the LDA is applied. This leads to the projection of the data (between the reference and the perturbed states) in a highly separated form on a 1D line. From such LDA calculations, we are able to identify the relative importance of the different basic features in ensuring this separation of the data (between the reference and the perturbed states) on the 1D line. This relative importance of the different basic features is quantified as “importance scores” for the different features, which in turn tell us what to study and where to study. Such knowledge enables us to rapidly identify the key factors responsible for the non-linearly large EOS transport in PMETAC-brush-grafted nanochannels.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Multivariate statistical analysis: Principles and applications to coorbital streams of meteorite falls

Multivariate statistical analysis techniques (linear discriminant analysis and logistic regression) can provide powerful discrimination tools which are generally unfamiliar to the planetary science community. Fall parameters were used to identify a group of 17 H chondrites (Cluster 1) that were part of a coorbital stream which intersected Earth's orbit in May, from 1855 - 1895, and can be distinguished from all other H chondrite falls. Using multivariate statistical techniques, it was demonstrated that a totally different criterion, labile trace element contents - hence thermal histories - or 13 Cluster 1 meteorites are distinguishable from those of 45 non-Cluster 1 H chondrites. Here, we focus upon the principles of multivariate statistical techniques and illustrate their application using non-meteoritic and meteoritic examples.

Wolf, S. F.↗

Refinement of a Method for Identifying Probable Archaeological Sites from Remotely Sensed Data

To facilitate locating archaeological sites before they are compromised or destroyed, we are developing approaches for generating maps of probable archaeological sites, through detecting subtle anomalies in vegetative cover, soil chemistry, and soil moisture by analyzing remotely sensed data from multiple sources. We previously reported some success in this effort with a statistical analysis of slope, radar, and Ikonos data (including tasseled cap and NDVI transforms) with Student's t-test. We report here on new developments in our work, performing an analysis of 8-band multispectral Worldview-2 data. The Worldview-2 analysis begins by computing medians and median absolute deviations for the pixels in various annuli around each site of interest on the 28 band difference ratios. We then use principle components analysis followed by linear discriminant analysis to train a classifier which assigns a posterior probability that a location is an archaeological site. We tested the procedure using leave-one-out cross validation with a second leave-one-out step to choose parameters on a 9,859x23,000 subset of the WorldView-2 data over the western portion of Ft. Irwin, CA, USA. We used 100 known non-sites and trained one classifier for lithic sites (n=33) and one classifier for habitation sites (n=16). We then analyzed convex combinations of scores from the Archaeological Predictive Model (APM) and our scores. We found that that the combined scores had a higher area under the ROC curve than either individual method, indicating that including WorldView-2 data in analysis improved the predictive power of the provided APM.

Tilton, James C.↗

Metric Learning for Hyperspectral Image Segmentation

We present a metric learning approach to improve the performance of unsupervised hyperspectral image segmentation. Unsupervised spatial segmentation can assist both user visualization and automatic recognition of surface features. Analysts can use spatially-continuous segments to decrease noise levels and/or localize feature boundaries. However, existing segmentation methods use tasks-agnostic measures of similarity. Here we learn task-specific similarity measures from training data, improving segment fidelity to classes of interest. Multiclass Linear Discriminate Analysis produces a linear transform that optimally separates a labeled set of training classes. The defines a distance metric that generalized to a new scenes, enabling graph-based segmentation that emphasizes key spectral features. We describe tests based on data from the Compact Reconnaissance Imaging Spectrometer (CRISM) in which learned metrics improve segment homogeneity with respect to mineralogical classes.

Compact Reconnaissance Imaging Spectrometer (CRISM↗

Geologic mapping using LANDSAT data

The feasibility of automated classification for lithologic mapping with LANDSAT digital data was evaluated using three classification algorithms. The two supervised algorithms analyzed, a linear discriminant analysis algorithm and a hybrid algorithm which incorporated the Parallelepiped algorithm and the Bayesian maximum likelihood function, were comparable in terms of accuracy; however, classification was only 50 per cent accurate. The linear discriminant analysis algorithm was three times as efficient as the hybrid approach. The unsupervised classification technique, which incorporated the CLUS algorithm, delineated the major lithologic boundaries and, in general, correctly classified the most prominent geologic units. The unsupervised algorithm was not as efficient nor as accurate as the supervised algorithms. Analysis of spectral data for the lithologic units in the 0.4 to 2.5 microns region indicated that a greater separability of the spectral signatures could be obtained using wavelength bands outside the region sensed by LANDSAT.

Siegal, B. S.↗

Multisensor classification of sedimentary rocks

A comparison is made between linear discriminant analysis and supervised classification results based on signatures from the Landsat TM, the Thermal Infrared Multispectral Scanner (TIMS), and airborne SAR, alone and combined into extended spectral signatures for seven sedimentary rock units exposed on the margin of the Wind River Basin, Wyoming. Results from a linear discriminant analysis showed that training-area classification accuracies based on the multisensor data were improved an average of 15 percent over TM alone, 24 percent over TIMS alone, and 46 percent over SAR alone, with similar improvement resulting when supervised multisensor classification maps were compared to supervised, individual sensor classification maps. When training area signatures were used to map spectrally similar materials in an adjacent area, the average classification accuracy improved 19 percent using the multisensor data over TM alone, 2 percent over TIMS alone, and 11 percent over SAR alone. It is concluded that certain sedimentary lithologies may be accurately mapped using a single sensor, but classification of a variety of rock types can be improved using multisensor data sets that are sensitive to different characteristics such as mineralogy and surface roughness.

Evans, Diane↗

Machine Learning–Augmented Laser-Induced Breakdown Spectroscopy for Spectral Discrimination of Iron Oxalates

Enhanced characterization and phase identification of post-PUREX Pu Oxalates (PuOXA) are pivotal for nonproliferation and pre-detonation nuclear forensics. Despite significant advances in the characterization of PuO 2 samples, little is known about the impact of both the chemical structure and oxidation states of PuOXA (i.e., Pu(III) and Pu(IV)) have on optical emission signatures. Here, we demonstrate the analytical capabilities of laser-induced breakdown spectroscopy (LIBS) applied to Fe(II) and Fe(III) oxalate samples as surrogates for PuOXA, highlighting the discriminating features in the LIBS emission spectra arising from differences in the oxidation states within mixed FeOXA samples. We report the enhancement of spectral feature selection using Principal Component Analysis (PCA), which enables the analytical superiority of machine learning algorithms such as Linear Discriminant Analysis (LDA), Quadratic Discriminant Analysis (QDA), Partial Least Squares Regression (PLSR), Support Vector Regression (SVR), and Random Forest Regression (RFR) over conventional univariate techniques for phase discrimination and chemometric analysis. Cluster analysis revealed how both matrix effects and laser ablation influence cluster separability by introducing spectral artifacts that misdirect the maximization of variance. PCA-selected emission lines were used in the regression models, demonstrating that both univariate and multivariate linear regression models (i.e., PLSR and SVR) can achieve acceptable performance, with machine learning models outperforming conventional calibration regressions. Furthermore, the application of non-linearly activated PCA-selected emission lines illustrates how simplifying the data while retaining captured variance enables the use of less complex and more computationally efficient models. Furthermore, this is particularly evident in the underperformance of RFR, which suffers from increased computational costs and overfitting owing to its high complexity.

Oxalates↗

A geobotanical investigation based on linear discriminant and profile analyses of airborne Thematic Mapper Simulator data

This paper discusses the application of linear discriminant and profile analyses to detailed investigation of an airborne Thematic Mapper Simulator (TMS) image collected over a geobotanical test site. The test site was located on the Keweenaw Peninsula of Michigan's Upper Peninsula, and remote sensing data collection coincided with the onset of leaf senescence in the regional deciduous flora. Linear discriminant analysis revealed that sites overlying soil geochemical anomalies were distinguishable from background sites by the reflectance and thermal emittance of the tree canopy imaged in the airborne TMS data. The correlation of individual bands with the linear discriminant function suggested that the TMS thermal Channel 7 (10.32-12.33 microns) contributed most, while TMS Bands 2 (0.53-0.60 microns), 3 (0.63-0.69 microns), and 5 (1.53-1.73 microns) contributed somewhat more modestly to the separation of anomalous and background sites imaged by the TMS. The observed changes in canopy reflectance and thermal emittance of the deciduous flora overlying geochemically anomalous areas are consistent with the biophysical changes which are known or presumed to occur as a result of injury induced in metal-stressed vegetation.

Schwaller, Mathew R.↗

Chemical studies of H chondrites. 6: Antarctic/non-Antarctic compositional differences revisited

We report data for the trace elements Au, Co, Sb, Ga, Rb, Ag, Se, Cs, Te, Zn, Cd, Bi, T1, and In (ordered by putative volatility during nebular condensation and accretion) determined by radiochemical neutron activation analysis of 14 additional H5 and H6 chondrite falls. Data for the 10 most volatile elements (Rb to In) treated by the multivariate techniques of linear discriminant analysis and logistic regression in these and 44 other falls are compared with those of 59 H4-6 chondrites from Antarctica. Various populations are tested by the multivariate techniques, using the previously developed method of randomization-simulation to assess significance levels. An earlier conclusion, based on fewer examples, that H4-6 chondrite falls are compositionally distinguishable from the Antarctic suite is verified by the additional data. This distinctiveness is highly significant because of the presence of samples from Victoria Land in the Antarctic population, which differ compositionally from falls beyond any reasonable doubt. However, it cannot be proven unequivocally that falls and Antarctic samples from Queen Maud Land are compositionally distinguishable. Trivial causes (e.g., analyst bias, weathering) cannot explain the Victoria Land (Antarctic)/non-Antarctic compositional difference for paradigmatic H4-6 chondrites. This seems to reflect a time-dependent variation of near-Earth meteoroid source regions differing in average thermal history.

Wolf, Stephen F.↗

Ordinary chondrites - Multivariate statistical analysis of trace element contents

The contents of mobile trace elements (Co, Au, Sb, Ga, Se, Rb, Cs, Te, Bi, Ag, In, Tl, Zn, and Cd) in Antarctic and non-Antarctic populations of H4-6 and L4-6 chondrites, were compared using standard multivariate discriminant functions borrowed from linear discriminant analysis and logistic regression. A nonstandard randomization-simulation method was developed, making it possible to carry out probability assignments on a distribution-free basis. Compositional differences were found both between the Antarctic and non-Antarctic H4-6 chondrite populations and between two L4-6 chondrite populations. It is shown that, for various types of meteorites (in particular, for the H4-6 chondrites), the Antarctic/non-Antarctic compositional difference is due to preterrestrial differences in the genesis of their parent materials.

Lipschutz, Michael E.↗

Chemical studies of H chondrites. 4: New data and comparison of Antarctic suites

We report data for the trace elements Au, Co, Sb, Ga, Rb, Ag, Se, Cs, Te, Zn, Cd, Bi, Ti, and In (ordered by putative volatility during nebular condensation and accretion) determined by neutron activation analysis in 13 H5 chondrites from Victoria Land and 20 H4-6 chondrites from Queen Maud Land, Antarctica. These and earlier results provide Antarctic sample suites of 34 chondrites from Victoria Land and 25 from Queen Maud Land. Treatment of data for the most volatile 10 elements (Rb to In) in these studies by multivariate statistical techniques more robust, as well as more conservative, than conventional linear discriminant analysis and logistic regression demonstrates that compositions differ at marginally significant levels. This difference cannot be explained by trivial (terrestrial) causes and becomes more significant, despite the smaller size of the database, when comparisons are limited to data from a single analyst and when all upper limits are eliminated from consideration. The Victoria Land and Queen Maud Land suites have different mean terrestrial ages (approximately 300 kyr and approximately 100 kyr, respectively) and age distributions, suggesting that a time-dependent variation of chondritic sources with different thermal histories is responsible. As a result, these two Antarctic suites are, on average, chemically distinguishable from each other. Since H chondrites serve as a paradigm for other meteorite classes, these results indicate that the near-Earth populations of planetary materials varied with time on the 10(exp 5)-year timescale.

Wolf, Stephen F.↗

Machine Learning Discrimination and Ultrasensitive Detection of Fentanyl Using Gold Nanoparticle-Decorated Carbon Nanotube-Based Field-Effect Transistor Sensors

The opioid overdose crisis is a global health challenge. Fentanyl, an exceedingly potent synthetic opioid, has emerged as a leading contributor to the surge in opioid-related overdose deaths. The surge in overdose fatalities, particularly due to illicitly manufactured fentanyl and its contamination of street drugs, emphasizes the urgency for drug-testing technologies that can quickly and accurately identify fentanyl from other drugs and quantify trace amounts of fentanyl. In this paper, gold nanoparticle (AuNP)-decorated single-walled carbon nanotube (SWCNT)-based field-effect transistors (FETs) are utilized for machine learning-assisted identification of fentanyl from codeine, hydrocodone, and morphine. The unique sensing performance of fentanyl led to use machine learning approaches for accurate identification of fentanyl. Employing linear discriminant analysis (LDA) with a leave-one-out cross-validation approach, a validation accuracy of 91.2% is achieved. Meanwhile, density functional theory (DFT) calculations reveal the factors that contributed to the enhanced sensitivity of the Au-SWCNT FET sensor toward fentanyl as well as the underlying sensing mechanism. Finally, fentanyl antibodies are introduced to the Au-SWCNT FET sensor as specific receptors, expanding the linear range of the sensor in the lower concentration range, and enabling ultrasensitive detection of fentanyl with a limit of detection at 10.8 fg mL –1 .

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Transcriptomics-based Machine Learning (ML) Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗