Search NASASearch

SEARCH · Search NASA

Results for “high dimensional data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Feature Selection in High-Dimensional Space with Applications to Gene Expression Data

Recent years have seen rapid growth in high-dimensional datasets. Most existing machine learning (ML) algorithms fail in high-dimensional settings where many features could be redundant. A critical process of feature selection is thus applied in such a setting that helps in identifying the most relevant features while removing redundant ones. With the increase in high dimensionality, one is also faced with problems of efficiency and interpretation in performing such selection methods. Therefore, this paper proposes a “novel” feature selection framework that uses an ensemble of interpretable ML algorithms to perform feature selection and the ranking of final features. Finally, this framework is applied to a gene expression dataset obtained through collaboration with the National Aeronautics and Space Administration (NASA)’s Biological and Physical Sciences (BPS) team and helps identify important and relevant genes contributing to specific target attributes through classification tasks.

Nishan Pantha

High-Performance Computing and Four-Dimensional Data Assimilation: The Impact on Future and Current Problems

This is the final technical report for the project entitled: "High-Performance Computing and Four-Dimensional Data Assimilation: The Impact on Future and Current Problems", funded at NPAC by the DAO at NASA/GSFC. First, the motivation for the project is given in the introductory section, followed by the executive summary of major accomplishments and the list of project-related publications. Detailed analysis and description of research results is given in subsequent chapters and in the Appendix.

Makivic, Miloje S.

System and Safety Analysis with SysAI A Statistical Learning Framework

This is a tutorial on how to use the SYSAI (System Analysis using Statistical AI), a flexible statistical learning framework for the V&V and analysis of complex and high-dimensional Aerospace systems with DNN and AI components. SYSAI provides functionality for a variety of analyses and V&V tasks, including statistical data analysis, high dimensional safety-envelope and time-series analysis, property checking, as well as intelligent test-case generation. The tutorial will demonstrate SYSAI with our industrial partner’s Autonomous Centerline Tracking system, which uses a DNN to enable autonomous aircraft taxiing as an example. Video & Tutorial

Statistical V&V for Complex safety-critical system

Advances in Hyperspectral Image Classification Methods for Vegetation and Agricultural Cropland Studies

Hyperspectral data are becoming more widely available via sensors on airborne and unmanned aerial vehicle (UAV) platforms, as well as proximal platforms. While space-based hyperspectral data continue to be limited in availability, multiple spaceborne Earth-observing missions on traditional platforms are scheduled for launch, and companies are experimenting with small satellites for constellations to observe the Earth, as well as for planetary missions. Land cover mapping via classification is one of the most important applications of hyperspectral remote sensing and will increase in significance as time series of imagery are more readily available. However, while the narrow bands of hyperspectral data provide new opportunities for chemistry-based modeling and mapping, challenges remain. Hyperspectral data are high dimensional, and many bands are highly correlated or irrelevant for a given classification problem. For supervised classification methods, the quantity of training data is typically limited relative to the dimension of the input space. The resulting Hughes phenomenon, often referred to as the curse of dimensionality, increases potential for unstable parameter estimates, overfitting, and poor generalization of classifiers. This is particularly problematic for parametric approaches such as Gaussian maximum likelihood–based classifiers that have been the backbone of pixel-based multispectral classification methods. This issue has motivated investigation of alternatives, including regularization of the class covariance matrices, ensembles of weak classifiers, development of feature selection and extraction methods, adoption of nonparametric classifiers, and exploration of methods to exploit unlabeled samples via semi-supervised and active learning. Data sets are also quite large, motivating computationally efficient algorithms and implementations. This chapter provides an overview of the recent advances in classification methods for mapping vegetation using hyperspectral data. Three data sets that are used in the hyperspectral classification literature (e.g., Botswana Hyperion satellite data and AVIRIS airborne data over both Kennedy Space Center and Indian Pines) are described in Section 3.2 and used to illustrate methods described in the chapter. An additional high-resolution hyperspectral data set acquired by a SpecTIR sensor on an airborne platform over the Indian Pines area is included to exemplify the use of new deep learning approaches, and a multiplatform example of airborne hyperspectral data is provided to demonstrate transfer learning in hyperspectral image classification. Classical approaches for supervised and unsupervised feature selection and extraction are reviewed in Section 3.3. In particular, nonlinearities exhibited in hyperspectral imagery have motivated development of nonlinear feature extraction methods in manifold learning, which are outlined in Section 3.3.1.4. Spatial context is also important in classification of both natural vegetation with complex textural patterns and large agricultural fields with significant local variability within fields. Approaches to exploit spatial features at both the pixel level (e.g., co-occurrence–based texture and extended morphological attribute profiles [EMAPs]) and integration of segmentation approaches (e.g., HSeg) are discussed in this context in Section 3.3.2. Recently, classification methods that leverage nonparametric methods originating in the machine learning community have grown in popularity. An overview of both widely used and newly emerging approaches, including support vector machines (SVMs), Gaussian mixture models, and deep learning based on convolutional neural networks is provided in Section 3.4. Strategies to exploit unlabeled samples, including active learning and metric learning, which combine feature extraction and augmentation of the pool of training samples in an active learning framework, are outlined in Section 3.5. Integration of image segmentation with classification to accommodate spatial coherence typically observed in vegetation is also explored, including as an integrated active learning system. Exploitation of multisensor strategies for augmenting the pool of training samples is investigated via a transfer learning framework in Section 3.5.1.2. Finally, we look to the future, considering opportunities soon to be provided by new paradigms, as hyperspectral sensing is becoming common at multiple scales from ground-based and airborne autonomous vehicles to manned aircraft and space-based platforms.

Pasolli, Edoardo

Using Image Tour to Explore Multiangle, Multispectral Satellite Image

This viewgraph presentation reviews the use of Image Tour to explore the multiangle, multispectral satellite imagery. Remote sensing data are spatial arrays of p-dimensional vectors where each component corresponds to one of p variables. Applying the same R(exp p) to R(exp d) projection to all pixels creates new images, which may be easier to analyze than the original because d < p. Image grand tour (IGT) steps through the space of projections, and d=3 outputs a sequence of RGB images, one for each step. In this talk, we apply IGT to multiangle, multispectral data from NASA's MISR instrument. MISR views each pixel in four spectral bands at nine view angles. Multiple views detect photon scattering in different directions and are indicative of physical properties of the scene. IGT allows us to explore MISR's data structure while maintaining spatial context; a key requirement for physical interpretation. We report results highlighting the uniqueness of multiangle data and how IGT can exploit it.

visualization

Consensus theoretic classification methods

Consensus theory is adopted as a means of classifying geographic data from multiple sources. The foundations and usefulness of different consensus theoretic methods are discussed in conjunction with pattern recognition. Weight selections for different data sources are considered and modeling of non-Gaussian data is investigated. The application of consensus theory in pattern recognition is tested on two data sets: 1) multisource remote sensing and geographic data and 2) very-high-dimensional remote sensing data. The results obtained using consensus theoretic methods are found to compare favorably with those obtained using well-known pattern recognition methods. The consensus theoretic methods can be applied in cases where the Gaussian maximum likelihood method cannot. Also, the consensus theoretic methods are computationally less demanding than the Gaussian maximum likelihood method and provide a means for weighting data sources differently.

Benediktsson, Jon A.

Toward Soil Spatial Information Systems (SSIS) for global modeling and ecosystem management

The general objective is to conduct research to contribute toward the realization of a world soils and terrain (SOTER) database, which can stand alone or be incorporated into a more complete and comprehensive natural resources digital information system. The following specific objectives are focussed on: (1) to conduct research related to (a) translation and correlation of different soil classification systems to the SOTER database legend and (b) the inferfacing of disparate data sets in support of the SOTER Project; (2) to examine the potential use of AVHRR (Advanced Very High Resolution Radiometer) data for delineating meaningful soils and terrain boundaries for small scale soil survey (range of scale: 1:250,000 to 1:1,000,000) and terrestrial ecosystem assessment and monitoring; and (3) to determine the potential use of high dimensional spectral data (220 reflectance bands with 10 m spatial resolution) for delineating meaningful soils boundaries and conditions for the purpose of detailed soil survey and land management.

Baumgardner, Marion F.

A spectral feature design system for the HIRIS/MODIS era

A spectral feature design system for high-dimensional multispectral data is described. This system utilizes hard-limited or infinitely clipped optimal transforms and canonical analysis to extract the spectral features for data volume reduction and classification purposes. The design procedure is intended to be application-specific in order to make it maximally effective for each use. Although it could be used in a variety of circumstances, the procedure was designed with satellite data collection in mind, such as will be needed with the High-Resolution Imaging Spectrometer (HIRIS), a land-oriented Earth observational sensor intended for launch in the mid-1990s. The computation required for the steps prior to the actual satellite data collection are straightforward and could be done with readily available subroutines. The procedure is also defined in such a way as to require only very simple onboard calculations. The tests reported provide substantial data volume reduction in the satellite-to-Earth downlink and subsequent computation phases, while maintaining satisfactory classification accuracy.

Chen, Chih-Chien Thomas

Developing Fast and Accurate Radiative Transfer Models to Meet the Needs of Modern Satellite Remote Sensing Applications

Modern hyperspectral satellite remote sensors provide highly accurate measurements the Earth’s Top-of-Atmosphere (TOA) radiance, reflectance, or polarized spectra with hundreds to thousands of spectral channels and with millions of observations per day. The large data volume and high spectral dimensionality of the data pose challenges for retrieval algorithms. To process the satellite Level-1 data (e.g. calibrated TOA spectra) into Level-2 products (e.g. atmospheric and surface properties) using physical-based retrieval algorithms, accurate and fast Radiative Transfer Models (RTMs) are needed. RTMs are usually the limiting factor in determining the speed of a level-2 algorithm. For example, more than one million Line-by-Line (LBL) radiative transfer (RT) calculations are needed in order to properly capture the spectral contributions of important atmospheric molecules for an IR hyperspectral sensor with a spectral coverage from 3.5 m to 15 m or a solar hyperspectral sensor with spectral coverage from 0.25 m to 2.5 m. In this presentation, we will discuss advantages and disadvantages of different ways (e.g. correlated k and effective transmittance) to accelerate the speed of a fast RTM. We finally describe a Principal Component-based Radiative Transfer Model (PCRTM), which can calculate TOA radiance or reflectance spectra from 50 cm-1 to 40,000 cm-1 (200 m to 0.25 m). It has demonstrated very good accuracy relative to reference LBL RTMs and saves orders of magnitude in computational time. The PCRTM has been used in many satellite remote sensing applications. Examples include forward modeling in Level-2 and Level-3 retrieval algorithms, high fidelity satellite instrument simulators and instrument performance trade studies, spectral and radiometric accuracy characterizations of satellite Level-1 data, tools for inter-satellite calibrations, tools for satellite RTM lookup table generations, and tools for generating physically based training datasets for Artificial Intelligence (AI) algorithms.

climate data record

High dimensional reflectance analysis of soil organic matter

Recent breakthroughs in remote-sensing technology have led to the development of high spectral resolution imaging sensors for observation of earth surface features. This research was conducted to evaluate the effects of organic matter content and composition on narrowband soil reflectance across the visible and reflective infrared spectral ranges. Organic matter from four Indiana agricultural soils, ranging in organic C content from 0.99 to 1.72 percent, was extracted, fractionated, and purified. Six components of each soil were isolated and prepared for spectral analysis. Reflectance was measured in 210 narrow bands in the 400- to 2500-nm wavelength range. Statistical analysis of reflectance values indicated the potential of high dimensional reflectance data in specific visible, near-infrared, and middle-infrared bands to provide information about soil organic C content, but not organic matter composition. These bands also responded significantly to Fe- and Mn-oxide content.

Henderson, T. L.

A comparison of spectral mixture analysis an NDVI for ascertaining ecological variables

In this study, we compare the performance of spectral mixture analysis to the Normalized Difference Vegetation Index (NDVI) in detecting change in a grassland across topographically-induced nutrient gradients and different management schemes. The Konza Prairie Research Natural Area, Kansas, is a relatively homogeneous tallgrass prairie in which change in vegetation productivity occurs with respect to topographic positions in each watershed. The area is the site of long-term studies of the influence of fire and grazing on tallgrass production and was the site of the First ISLSCP (International Satellite Land Surface Climatology Project) Field Experiment (FIFE) from 1987 to 1989. Vegetation indices such as NDVI are commonly used with imagery collected in few (less than 10) spectral bands. However, the use of only two bands (e.g. NDVI) does not adequately account for the complex of signals making up most surface reflectance. Influences from background spectral variation and spatial heterogeneity may confound the direct relationship with biological or biophysical variables. High dimensional multispectral data allows for the application position of techniques such as derivative analysis and spectral curve fitting, thereby increasing the probability of successfully modeling the reflectance from mixed surfaces. The higher number of bands permits unmixing of a greater number of surface components, separating the vegetation signal for further analyses relevant to biological variables.

Wessman, Carol A.

Transcriptomics-based Machine Learning (ML) Analysis Predicts Space-Exposed Murine Livers

NASA has employed high-throughput molecular assays to identify sub-cellular changes impacting human physiology during spaceflight. Machine learning (ML) methods hold the promise to improve our ability to identify important signals within highly dimensional molecular data. However, the inherent limitation of study subject numbers within a spaceflight mission minimizes the utility of ML approaches. To overcome the sample power limitations, data from multiple spaceflight missions must be aggregated while appropriately addressing intra- and inter-study variabilities. Here we describe an approach to log transform, scale and normalize data from six heterogeneous, mouse liver derived transcriptomics datasets (ntotal=137) which enabled ML-methods to perform well (AUC ≥ 0.87) in classifying spaceflown vs ground control animals rather than mission-of-origin. Concordance was found between liver-specific biological processes identified from harmonized ML-based analysis and study-by-study classical omics analysis. This work demonstrates the feasibility of applying ML methods on integrated, heterogeneous datasets of small sample size.

Machine Learning

High dimensional feature reduction via projection pursuit

The recent development of more sophisticated remote sensing systems enables the measurement of radiation in many more spectral intervals than previously possible. An example of that technology is the AVIRIS system, which collects image data in 220 bands. As a result of this, new algorithms must be developed in order to analyze the more complex data effectively. Data in a high dimensional space presents a substantial challenge, since intuitive concepts valid in a 2-3 dimensional space to not necessarily apply in higher dimensional spaces. For example, high dimensional space is mostly empty. This results from the concentration of data in the corners of hypercubes. Other examples may be cited. Such observations suggest the need to project data to a subspace of a much lower dimension on a problem specific basis in such a manner that information is not lost. Projection Pursuit is a technique that will accomplish such a goal. Since it processes data in lower dimensions, it should avoid many of the difficulties of high dimensional spaces. In this paper, we begin the investigation of some of the properties of Projection Pursuit for this purpose.

Jimenez, Luis

High Performance Input/Output Systems for High Performance Computing and Four-Dimensional Data Assimilation

The approach of this task was to apply leading parallel computing research to a number of existing techniques for assimilation, and extract parameters indicating where and how input/output limits computational performance. The following was used for detailed knowledge of the application problems: 1. Developing a parallel input/output system specifically for this application 2. Extracting the important input/output characteristics of data assimilation problems; and 3. Building these characteristics s parameters into our runtime library (Fortran D/High Performance Fortran) for parallel input/output support.

Fox, Geoffrey C.

Onboard Hyperspectral Image Classification via Transfer Learning for Communication-Limited Spacecraft

Employing deep-learning and artificial-intelligence (AI) techniques onboard spacecraft can dramatically improve priority data selection to ensure more effective use of the available downlink. However, deployment of effective deep-learning models requires significant training on the ground, which may not be feasible, due to limited data available in an unexplored environment. Therefore, this research explores building robust classification models for onboard data processing where training data is highly limited using transfer-learning techniques. In this paper, we focus on the use case of hyperspectral imaging for remote sensing, a domain where the high dimensionality of the data from the sensor can rapidly saturate the downlink bandwidth. With this bottleneck, there is an impending need to autonomously and robustly classify data onboard to optimize downlink of high-impact measurements, thus maximizing the scientific utility per bit transmitted to the ground. This paper examines the use of deep neural networks onboard for hyperspectral image classification in a communication-limited scenario to analyze how the models perform with limited training data. The use of transfer learning can ameliorate the issue of poor generalization by transferring features learned from training on a large source dataset for one classification task to the target classification task with limited training data. For two deep-learning models from literature, we compare the accuracy of the models trained using transfer learning to models trained from scratch using a random weight initialization with varying amounts of training data. We demonstrate the feasibility and performance of running inference of the deep-learning models on representative flight-like hardware.

Advanced Avionics, Machine Learning, Data Processi

A Spatiotemporal Indexing Approach for Efficient Processing of Big Array-Based Climate Data with MapReduce

Climate observations and model simulations are producing vast amounts of array-based spatiotemporal data. Efficient processing of these data is essential for assessing global challenges such as climate change, natural disasters, and diseases. This is challenging not only because of the large data volume, but also because of the intrinsic high-dimensional nature of geoscience data. To tackle this challenge, we propose a spatiotemporal indexing approach to efficiently manage and process big climate data with MapReduce in a highly scalable environment. Using this approach, big climate data are directly stored in a Hadoop Distributed File System in its original, native file format. A spatiotemporal index is built to bridge the logical array-based data model and the physical data layout, which enables fast data retrieval when performing spatiotemporal queries. Based on the index, a data-partitioning algorithm is applied to enable MapReduce to achieve high data locality, as well as balancing the workload. The proposed indexing approach is evaluated using the National Aeronautics and Space Administration (NASA) Modern-Era Retrospective Analysis for Research and Applications (MERRA) climate reanalysis dataset. The experimental results show that the index can significantly accelerate querying and processing (10 speedup compared to the baseline test using the same computing cluster), while keeping the index-to-data ratio small (0.0328). The applicability of the indexing approach is demonstrated by a climate anomaly detection deployed on a NASA Hadoop cluster. This approach is also able to support efficient processing of general array-based spatiotemporal data in various geoscience domains without special configuration on a Hadoop cluster.

big data

Surface pressure measurements on the blade of an operating Mod-2 wind turbine with and without vortex generators

Pressure measurements covering a range of wind velocities were made at one span location on the surface of an operating Mod-2, 2500 kW, wind turbine blade. The data, which were taken with and without vortex generators installed on the leading edge, show the existence of higher pressure coefficients than would be expected from two-dimensional wind tunnel data. These high pressure ratios may be the result of three-dimensional flow over the blade, which delays flow separation. Data are presented showing the repetitiveness of abrupt changes in the pressure distribution that occur as the blade rotates. Calculated values of suction and flap coefficients are also presented.

Nyland, Ted W.