Search NASASearch

SEARCH · Search NASA

Results for “dimensionality reduction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

From microbial diversity to functional potential using dimensionality reduction

The high dimensionality of microbial diversity data from ‘omics observations can be reduced using Machine Learning, with many recent studies showcasing ML utility for exploratory ecological feature finding and process prediction. Here, we compare the Self Organizing Map (SOM) dimensionality reduction method to the well-documented sample-based Principal Coordinate Analysis (PCoA) and taxa-based Weighted Gene Correlation Network Analysis (WGCNA) using near daily 16S rRNA gene amplicon sequencing data from the 2019 to 2020 MOSAiC International Arctic Drift Expedition. We then map k-means clustering outputs from each method to available metagenomes, extracting functionally distinct seasonal microbial ecotypes in the surface Arctic Ocean. Our results indicate the SOM method better represented expected seasonal transitions and identified a greater number of metabolically distinct functional groups than the more traditional PCoA ordination. Ultimately, we identified four community ecotypes with distinct taxonomic and functional cut-offs driven by seasonality, water mass, and substrate turnover, highlighting the importance of succession in functional diversity for the central Arctic Ocean. These results reinforce ML dimensionality reduction as a meaningful translator in the mining of historical amplicon datasets to address modern mechanistic questions and potentially provide ’omics informed ecotype diversity to leverage in mechanistic biogeochemical models.

Arctic Ocean

Neural Active Manifolds: Nonlinear Dimensionality Reduction for Uncertainty Quantification

We present a new approach for nonlinear dimensionality reduction, specifically designed for computationally expensive mathematical models. We leverage autoencoders to discover a one-dimensional neural active manifold (NeurAM) capturing the model output variability, through the aid of a simultaneously learnt surrogate model with inputs on this manifold. Our method only relies on model evaluations and does not require the knowledge of gradients. The proposed dimensionality reduction framework can then be applied to assist outer loop many-query tasks in scientific computing, like sensitivity analysis and multifidelity uncertainty propagation. In particular, we prove, both theoretically under idealized conditions, and numerically in challenging test cases, how NeurAM can be used to obtain multifidelity sampling estimators with reduced variance by sampling the models on the discovered low-dimensional and shared manifold among models. Several numerical examples illustrate the main features of the proposed dimensionality reduction strategy and highlight its advantages with respect to existing approaches in the literature.

Autoencoders

Kernel Manifolds: Nonlinear‐Augmentation Dimensionality Reduction Using Reproducing Kernel Hilbert Spaces

This paper generalizes recent advances on quadratic manifold (QM) dimensionality reduction by developing kernel methods-based nonlinear-augmentation dimensionality reduction. QMs, and more generally feature map-based nonlinear corrections, augment linear dimensionality reduction with a nonlinear correction term in the reconstruction map to overcome approximation accuracy limitations of purely linear approaches. While feature map-based approaches typically learn a least squares optimal polynomial correction term, we generalize this approach by learning an optimal nonlinear correction from a user-defined reproducing kernel Hilbert space. Our approach allows one to impose arbitrary nonlinear structure on the correction term, including polynomial structure, and includes feature map and radial basis function-based corrections as special cases. Furthermore, our method has relatively low training cost and has monotonically decreasing error as the latent space dimension increases. In conclusion, we compare our approach to proper orthogonal decomposition and several recent QM approaches on data from several example problems.

kernel methods

Using the Dimensionality Reduction (DR) Approach to Treat Singular and Near-Singular Source and Test Integrals on Triangles for Moment Methods

Recently, the authors presented preliminary results of an initial study of a dimensional reduction scheme for treating (near-)singular integrals (D. R. Wilton et al., “Dimensionality Reduction Approach for Treating Singular and Near-Singular Multidimensional Integrals of Electromagnetics,” 2022 URSI USNC National Radio Science Meeting, Boulder, CO, Jan. 2022). The approach reported there refined and extended several ideas appearing in a recent paper (R. R. Chang, Z. Wang, and Q. Xie, "Fast Convergent Quadrature Method for Evaluating the RWG- and SWG-Related Convolutional Integrals," IEEE Trans. Antennas Propagat., 69, Dec. 2021). That paper reported a curated collection of important, previously published results all leading to very efficient and accurate line integral methods for handling the most common (near-)singular integrals associated with triangles and tetrahedrons with RWG and SWG bases, respectively. Our contributions provided a simpler exposition as well as a more unified, systematic, and robust overall framework for deriving and applying the approach. This presentation further elucidates our extensions to the approach and broadens our initial study to examine the convergence behaviors of the various potential forms over a wider range of triangular source element shape and observation point parameters for various smoothing transformations; we also examine convergence, not just at the subtriangle level, but also overcomplete triangles. In addition, we investigate the application of the recently reported “vertex function” concept to develop faster converging test integrals (Rivero et al., “Acceleration of the Surface Test Integral Using Vertex Functions,” 2021 IEEE Int’l Symp. Ant. Propagat. and USNC-URSI Rad. Sci. Meeting,” Singapore, 4-10 Dec. 2021). Combining these separate schemes has the potential for yielding a near-optimal approach for accurately evaluating singular and near-singular integrals for moment methods.

Singular Integrals

Landsat D Thematic Mapper image dimensionality reduction and geometric correction accuracy

To characterize and quantify the performance of the Landsat thematic mapper (TM), techniques for dimensionality reduction by linear transformation have been studied and evaluated and the accuracy of the correction of geometric errors in TM images analyzed. Theoretical evaluations and comparisons for existing methods for the design of linear transformation for dimensionality reduction are presented. These methods include the discrete Karhunen Loeve (KL) expansion, Multiple Discriminant Analysis (MDA), Thematic Mapper (TM)-Tasseled Cap Linear Transformation and Singular Value Decomposition (SVD). A unified approach to these design problems is presented in which each method involves optimizing an objective function with respect to the linear transformation matrix. From these studies, four modified methods are proposed. They are referred to as the Space Variant Linear Transformation, the KL Transform-MDA hybrid method, and the First and Second Version of the Weighted MDA method. The modifications involve the assignment of weights to classes to achieve improvements in the class conditional probability of error for classes with high weights. Experimental evaluations of the existing and proposed methods have been performed using the six reflective bands of the TM data. It is shown that in terms of probability of classification error and the percentage of the cumulative eigenvalues, the six reflective bands of the TM data require only a three dimensional feature space. It is shown experimentally as well that for the proposed methods, the classes with high weights have improvements in class conditional probability of error estimates as expected.

Ford, G. E.

The Impact of Dimensionality Reduction of Ion Counts Distributions on Preserving Moments, With Applications to Data Compression

The field of space physics has a long history of utilizing dimensionality reduction methods to distill data, including but not limited to spherical harmonics, the Fourier Transform, and the wavelet transform. Here, we present a technique for performing dimensionality reduction on ion counts distributions from the Multiscale Mission/Fast Plasma Investigation (MMS/FPI) instrument using a data-adaptive method powered by neural networks. This has applications to both feeding low-dimensional parameterizations of the counts distributions into other machine learning algorithms, and the problem of data compression to reduce transmission volume for space missions. The algorithm presented here is lossy, and in this work, we present the technique of validating the reconstruction performance with calculated plasma moments under the argument that preserving the moments also preserves fluid-level physics, and in turn a degree of scientific validity. The method presented here is an improvement over other lossy compressions in loss-tolerant scenarios like the Multiscale Mission/Fast Plasma Investigation Fast Survey or in non-research space weather applications.

D. da Silva

Lightning Charge Retrievals: Dimensional Reduction, LDAR Constraints, and a First Comparison with LIS Satellite Data

A "dimensional reduction" (DR) method is introduced for analyzing lightning field changes (DELTAEs) whereby the number of unknowns in a discrete two-charge model is reduced from the standard eight (x, y, z, Q, x', y', z', Q') to just four (x, y, z, Q). The four unknowns (x, y, z, Q) are found by performing a numerical minimization of a chi-square function. At each step of the minimization, an Overdetermined Fixed Matrix (OFM) method is used to immediately retrieve the best "residual source" (x', y', z', Q'), given the values of (x, y, z, Q). In this way, all 8 parameters (x, y, z, Q, x', y', z', Q') are found, yet a numerical search of only 4 parameters (x, y, z, Q) is required. The DR method has been used to analyze lightning-caused DeltaEs derived from multiple ground-based electric field measurements at the NASA Kennedy Space Center (KSC) and USAF Eastern Range (ER). The accuracy of the DR method has been assessed by comparing retrievals with data provided by the Lightning Detection And Ranging (LDAR) system at the KSC-ER, and from least squares error estimation theory, and the method is shown to be a useful "stand-alone" charge retrieval tool. Since more than one charge distribution describes a finite set of DELTAEs (i.e., solutions are non-unique), and since there can exist appreciable differences in the physical characteristics of these solutions, not all DR solutions are physically acceptable. Hence, an alternative and more accurate method of analysis is introduced that uses LDAR data to constrain the geometry of the charge solutions, thereby removing physically unacceptable retrievals. The charge solutions derived from this method are shown to compare well with independent satellite- and ground-based observations of lightning in several Florida storms.

Koshak, W. J.

Lightning Charge Retrievals: Dimensional Reduction, LDAR Constraints, and a First Comparison w/ LIS Satellite Data

A "dimensional reduction" (DR) method is introduced for analyzing lightning field changes whereby the number of unknowns in a discrete two-charge model is reduced from the standard eight to just four. The four unknowns are found by performing a numerical minimization of a chi-squared goodness-of-fit function. At each step of the minimization, an Overdetermined Fixed Matrix (OFM) method is used to immediately retrieve the best "residual source". In this way, all 8 parameters are found, yet a numerical search of only 4 parameters is required. The inversion method is applied to the understanding of lightning charge retrievals. The accuracy of the DR method has been assessed by comparing retrievals with data provided by the Lightning Detection And Ranging (LDAR) instrument. Because lightning effectively deposits charge within thundercloud charge centers and because LDAR traces the geometrical development of the lightning channel with high precision, the LDAR data provides an ideal constraint for finding the best model charge solutions. In particular, LDAR data can be used to help determine both the horizontal and vertical positions of the model charges, thereby eliminating dipole ambiguities. The results of the LDAR-constrained charge retrieval method have been compared to the locations of optical pulses/flash locations detected by the Lightning Imaging Sensor (LIS).

Koshak, William

Lightning Charge Retrievals: Dimensional Reduction, LDAR Constraints, and a First Comparison with LIS Satellite Data

A "dimensional reduction" ("DR") method is introduced for analyzing lightning field changes ((Delta)Es) whereby the number of unknowns in a discrete two-charge model is reduced from the standard eight (x, y, z, Q, x', y', z', Q') to just four (x, y, z, Q). The four unknowns (x, y, z, Q) are found by performing a numerical minimization of a chi-square function. At each step of the minimization, an overdetermined fixed matrix (OFM) method is used to immediately retrieve the best "residual source" (x', y', z', Q'), given the values of (x, y, z, Q). In this way, all eight parameters Ix. y, z, Q, x', y', z', Q') are found, yet a numerical search of only four parameters (x, y, z, Q) is required. The DR method has been used to analyze lightning-caused (Delta)Es derived from multiple ground-based electric field measurements at the NASA Kennedy Space Center (KSC) and U.S. Air Force Eastern Range (ER). The accuracy of the DR method has been assessed by comparing retrievals with data provided by the lightning detection and ranging (LDAR) system at the KSC-ER. and from least squares error estimation theory, and the method is shown to be a useful "stand alone" charge retrieval tool. Since more than one charge distribution describes a finite set of (Delta)Es (i.e., solutions are nonunique), and since there can be appreciable differences in the physical characteristics of these solutions, not all DR solutions are physically acceptable. Hence. an alternative and more accurate method of analysis is introduced that uses LDAR data to constrain the geometry of the charge solutions. thereby removing physically unacceptable retrievals. The charge solutions derived from this method are shown to compare well with independent satellite- and ground-based observations of lightning m several Florida storms.

Koshak, W. J.

Dimensional Reduction: A Method for Retrieving Lightning Charge

A method is introduced for retrieving the locations and magnitudes of charges deposited by a lightning flash using multiple ground-based electric field change measurements. The method, called Dimensional Reduction, reduces the number of unknowns in a discrete 2-charge model from the standard of eight (x, y, z, Q, x' , y' , z' , Q') to just four (x, y, z, Q). This reduction is accomplished by analyzing "residual measurements" that are formed by subtracting from each ground-based electric field change the contribution due to the source (x, y, z, Q). Using an improved analytic solution to the four-parameter point charge model (or "Q-model") the residual measurements are inverted to find the associated "residual source." For flash charge depositions that look approximately like two arbitrary charges, the residual source will be modeled accurately when (x, y, z, Q) is accurate. Hence, one need only minimize a chi-squared goodness-of-fit that is a function of the four variables (x, y, z, Q), rather than one that is a function of the eight variables (x, y, z, Q, x , y , z , Q ). The accuracy of the method is assessed by inverting computer-simulated electric field changes produced from known charge sources. The method is also applied to analyze real lightning electric field change data derived from the NASA Kennedy Space Center (KSC) and United States Air Force (USAF) Eastern Range (ER) ground-based field mill network It is shown that the charge retrievals compare favorably with associated ancillary ground- and satellite- based lightning measurements.

Koshak, William

Classification improvement by optimal dimensionality reduction when training sets are of small size

A computer simulation was performed to test the conjecture that, when the sizes of the training sets are small, classification in a subspace of the original data space may give rise to a smaller probability of error than the classification in the data space itself; this is because the gain in the accuracy of estimation of the likelihood functions used in classification in the lower dimensional space (subspace) offsets the loss of information associated with dimensionality reduction (feature extraction). A number of pseudo-random training and data vectors were generated from two four-dimensional Gaussian classes. A special algorithm was used to create an optimal one-dimensional feature space on which to project the data. When the sizes of the training sets are small, classification of the data in the optimal one-dimensional space is found to yield lower error rates than the one in the original four-dimensional space.

Starks, S. A.

Dimensional Reduction Guides Electronic Structure Evolution in the A n Cu 4–n SnS 4 Semiconductor Series

The search for new functional materials with tunable properties remains a central challenge in chemistry, particularly for applications in energy and electronics. In this work, we present a framework for predictive crystal design in alkali metal chalcogenides that enables controlled dimensional reduction of a parent covalent motif, yielding a broad range of electronic structures, which systematically evolve from one parent to the other. We present 11 new members of the A n Cu 4–n SnS 4 family (A = alkali metal; n = 0–4), which reduce the three-dimensional (3D) covalent network of Cu 4 SnS 4 into various 3D, 2D, 1D, and 0D [Cu 4–n SnS 4 ] n− motifs through the substitution of Cu with alkali metals of various radii. The end members of the family set the range in achievable band gaps at 0.99 eV for fully covalent Cu 4 SnS 4 (n = 0) and 3.38 eV for K 4 SnS 4 (n = 4) with 0D [SnS 4 ] n− tetrahedra. As the dimensionality of [Cu 4–n SnS 4 ] n− systematically reduces within A n Cu 4–n SnS 4 (n = 1–3), a stepwise increase in band gap energy occurs through a gradual decrease in the energy of the valence band maximum and an increase in the conduction band minimum, with an increase in the effective masses of charge carriers. Furthermore, irrespective of the alkali metal, the thermal stability decreases with decreasing [Cu 4–n SnS 4 ] n− dimensionality within the quaternary members. Most importantly, we demonstrate that predictable crystal structure and property evolution for a given composition space is possible by deriving a general formula based on substituting the covalent metals of a parent structure with alkali metals.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Optimal band selection for dimensionality reduction of hyperspectral imagery

Hyperspectral images have many bands requiring significant computational power for machine interpretation. During image pre-processing, regions of interest that warrant full examination need to be identified quickly. One technique for speeding up the processing is to use only a small subset of bands to determine the 'interesting' regions. The problem addressed here is how to determine the fewest bands required to achieve a specified performance goal for pixel classification. The band selection problem has been addressed previously Chen et al., Ghassemian et al., Henderson et al., and Kim et al.. Some popular techniques for reducing the dimensionality of a feature space, such as principal components analysis, reduce dimensionality by computing new features that are linear combinations of the original features. However, such approaches require measuring and processing all the available bands before the dimensionality is reduced. Our approach, adapted from previous multidimensional signal analysis research, is simpler and achieves dimensionality reduction by selecting bands. Feature selection algorithms are used to determine which combination of bands has the lowest probability of pixel misclassification. Two elements required by this approach are a choice of objective function and a choice of search strategy.

Stearns, Stephen D.

Systematic characterization of unknown compounds via dimensionality reduction of time series

Analysis of ambient aerosols provides valuable insight into particle sources and formation chemistry. However, due to the complexity of atmospheric data and the dynamic nature of aerosol composition, a substantial fraction of data often become discarded by conventional analysis methods. Furthermore, a large fraction of chemical species within those data are unidentifiable due to a lack of matching spectral information, resulting in suboptimal characterization of chemical composition. Previous work has demonstrated techniques for cataloging analytes in a chromatographic dataset by deconvolution of mass spectra, but integration of these analytes throughout a large dataset remains time consuming. Here, we present a method to automatically identify an ion for quantitation for single-ion chromatogram based peak fitting and integration, enabling comprehensive integration of analytes with minimal user interaction. The resulting time series are clustered with a machine-learning based dimensionality reduction technique to systematically investigate the underlying characteristics of the categorized analytes and gain new insights into the chemical composition and physicochemical properties of the unidentifiable analytes. We apply these methods to existing atmospheric datasets collected in Manacapuru, Brazil during the GoAmazon2014/5 campaign to identify new analytes and interpret their variability and transformations in the atmosphere. The analysis results generate 408 time series from cataloged analytes of interest, and the clustering of those time series with spherical k-means results in 8 distinct clusters. We find the analytes form clusters based on their distinct physicochemical properties, demonstrating the method’s ability to systematically identify and selectively filter contaminants and instrumental analytes and characterize the unidentifiable analytes.

54 ENVIRONMENTAL SCIENCES

LANDSAT-D Thematic Mapper image dimensionality reduction and geometric correction accuracy

Principal components transformations was applied to a Walnut Creek, Texas subscene to reduce the dimensionality of the multispectral sensor data. This transformation was also applied to a LANDSAT 3 MSS subscene of the same area acquired in a different season and year. Results of both procedures are tabulated and allow for comparisons between TM and MSS data. The TM correlation matrix shows that visible bands 1 to 3 exhibit a high degree of correlation in the range 0.92 to 0.96. Correlation for bands 5 to 7 is 0.93. Band 4 is not highly correlated with any other band, with corrections in the range 0.13 to 0.52. The thermal band (6) is not highly correlated with other bands in the range 0.13 to 0.46. The MSS correlation matrix shows that bands 4 and 5 are highly correlated (0.96) as are bands 6 and 7 with a correlation of 0.92.

Ford, G. E.

Dimensionality Reduction Through Classifier Ensembles

In data mining, one often needs to analyze datasets with a very large number of attributes. Performing machine learning directly on such data sets is often impractical because of extensive run times, excessive complexity of the fitted model (often leading to overfitting), and the well-known "curse of dimensionality." In practice, to avoid such problems, feature selection and/or extraction are often used to reduce data dimensionality prior to the learning step. However, existing feature selection/extraction algorithms either evaluate features by their effectiveness across the entire data set or simply disregard class information altogether (e.g., principal component analysis). Furthermore, feature extraction algorithms such as principal components analysis create new features that are often meaningless to human users. In this article, we present input decimation, a method that provides "feature subsets" that are selected for their ability to discriminate among the classes. These features are subsequently used in ensembles of classifiers, yielding results superior to single classifiers, ensembles that use the full set of features, and ensembles based on principal component analysis on both real and synthetic datasets.

Oza, Nikunj C.