SEARCH · Search NASA
Results for “principal component analysis”
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Solving Forward and Inverse Problems for Hyperspectral Satellite Remote Sensors using Principal Component Analysis
Explore the source record for details and available documents.
Revisiting AVHRR tropospheric aerosol trends using principal component analysis
Explore the source record for details and available documents.
Principal Component Analysis applied on MASCS/MESSENGER data for the spectral investigation of Mercury's surface
Explore the source record for details and available documents.
An evaluation of air quality in major urban areas of India
Rapid economic growth and burgeoning population have contributed to enhanced levels of PM 2.5 concentrations in urban regions of India. Evaluation of ambient air quality facilitates the assessment of effectiveness of emission control measures and early identification of new sources. This study provides a comprehensive statistical analysis of PM 2.5 concentrations in key urban areas across India, including Delhi, Kolkata, Mumbai, Chennai, Hyderabad, and several regional centers. Data from 2017 to 2023 was analyzed using trend analysis, cluster analysis, principal component analysis, and geostatistical interpolation to understand spatiotemporal variations and sources. The analysis reveals significant differences in spatial distribution of PM 2.5 concentrations with high annual averages in urban regions in Indo-Gangetic plain (82–123 μg m −3 ) and relatively lower concentrations (29–46 μg m −3 ) in southern urban areas of Kerala, Tamil Nadu and Andhra Pradesh. Delhi state had the highest 24-averaged PM 2.5 concentrations (112 μg m −3 ) followed by urban regions in Uttar Pradesh, Bihar and West Bengal (94 μg m −3 ). Trend analysis from 2017 to 2023 revealed an overall 2.5% decline in site-wide PM2.5 concentrations, with the exception of Ludhiana, which exhibited a consistent annual increase of 10%. Principal component analysis (PCA) attributes 30% of the variance to wintertime emissions, 13% to biomass burning, and 18% to the regional haze in the northern Indo-Gangetic Plain. Different analyses clearly demonstrates the contribution of biomass burning to pollution in Delhi and surrounding cities. Transboundary pollution to Kolkata is likely from the highly polluted region in Indo-Gangetic Plain. Coastal cities of Mumbai and Chennai has relatively lower pollution attributed to the influence of sea breeze dilution, with mostly local contribution and some potential transport from upwind industry clusters. Hyderabad also has local contribution due to high density of vehicular traffic and local small industries. This study shows that mitigation efforts targeting clusters of regions should be undertaken to curb the high PM2.5 pollution. Policy measures should be implemented both at local and the intra-state level to address shared sources and transport of pollution.
Uncertainty Quantification for Smooth Functional Data with Application to Material Properties
This document outlines a method for processing functional output (i.e., curves) for the ultimate purpose of sampling curves under specified input conditions for use in modeling and simulation uncertainty quantification (UQ) studies. A set of benchmark curves sufficiently representative of the relevant scenario(s) being simulated are provided to the process and formatted as described in Section 1. Principal Component Analysis (PCA) is utilized to discover the components of uncertainty in the benchmark curves and is outlined in Section 2. Section 3 describes the application of uncertainty quantification to the PCA results for the purpose of sampling curves to be used in UQ analysis. Section 4 applies these techniques to an example benchmark dataset. Concluding remarks are provided in the final section.
Probabilisitc Geobiological Classification Using Elemental Abundance Distributions and Lossless Image Compression in Recent and Modern Organisms
Last year we presented techniques for the detection of fossils during robotic missions to Mars using both structural and chemical signatures[Storrie-Lombardi and Hoover, 2004]. Analyses included lossless compression of photographic images to estimate the relative complexity of a putative fossil compared to the rock matrix [Corsetti and Storrie-Lombardi, 2003] and elemental abundance distributions to provide mineralogical classification of the rock matrix [Storrie-Lombardi and Fisk, 2004]. We presented a classification strategy employing two exploratory classification algorithms (Principal Component Analysis and Hierarchical Cluster Analysis) and non-linear stochastic neural network to produce a Bayesian estimate of classification accuracy. We now present an extension of our previous experiments exploring putative fossil forms morphologically resembling cyanobacteria discovered in the Orgueil meteorite. Elemental abundances (C6, N7, O8, Na11, Mg12, Ai13, Si14, P15, S16, Cl17, K19, Ca20, Fe26) obtained for both extant cyanobacteria and fossil trilobites produce signatures readily distinguishing them from meteorite targets. When compared to elemental abundance signatures for extant cyanobacteria Orgueil structures exhibit decreased abundances for C6, N7, Na11, All3, P15, Cl17, K19, Ca20 and increases in Mg12, S16, Fe26. Diatoms and silicified portions of cyanobacterial sheaths exhibiting high levels of silicon and correspondingly low levels of carbon cluster more closely with terrestrial fossils than with extant cyanobacteria. Compression indices verify that variations in random and redundant textural patterns between perceived forms and the background matrix contribute significantly to morphological visual identification. The results provide a quantitative probabilistic methodology for discriminating putatitive fossils from the surrounding rock matrix and &om extant organisms using both structural and chemical information. The techniques described appear applicable to the geobiological analysis of meteoritic samples or in situ exploration of the Mars regolith. Keywords: cyanobacteria, microfossils, Mars, elemental abundances, complexity analysis, multifactor analysis, principal component analysis, hierarchical cluster analysis, artificial neural networks, paleo-biosignatures
Combining ToF‐SIMS and Multivariate Analysis to Resolve Active Sites on Ni‐Based HER Catalysts
Unambiguous identification of active sites in heterogeneous catalysis remains a major challenge, particularly for materials with ultrathin, chemically mixed surface layers. Here, we demonstrate a generalizable approach that combines time-of-flight secondary ion mass spectrometry (ToF-SIMS) with multivariate statistical analysis (principal component analysis [PCA] and multivariate curve resolution [MCR]) to resolve catalytically relevant motifs at the nanoscale. Using Ni electrodes as a model system, PCA distinguished hydroxide-enriched domains from oxide- and metal-rich regions, while MCR decomposed depth profiles and 3D images into hydroxide, oxide, and metallic layers with nanometer resolution. A unique secondary-ion fragment, NiO 3 H 3 − (m/z 108.94), emerged as a marker of hydroxide-rich environments and correlated with hydrogen evolution reaction (HER) activity across a series of Ni electrodes. Complementary density functional theory (DFT) calculations revealed that Ni(OH) 2 clusters adjacent to metallic Ni offer the most favorable water dissociation energetics, establishing the structural origin of the marker. While illustrated here for Ni-based HER, this workflow provides a broadly applicable framework to isolate and rank near-surface patterns that govern catalytic activity, thereby extending ToF-SIMS from a qualitative probe to a predictive tool for active site identification.
NASA Instrument Cost/Schedule Model
NASA's Office of Independent Program and Cost Evaluation (IPCE) has established a number of initiatives to improve its cost and schedule estimating capabilities. 12One of these initiatives has resulted in the JPL developed NASA Instrument Cost Model. NICM is a cost and schedule estimator that contains: A system level cost estimation tool; a subsystem level cost estimation tool; a database of cost and technical parameters of over 140 previously flown remote sensing and in-situ instruments; a schedule estimator; a set of rules to estimate cost and schedule by life cycle phases (B/C/D); and a novel tool for developing joint probability distributions for cost and schedule risk (Joint Confidence Level (JCL)). This paper describes the development and use of NICM, including the data normalization processes, data mining methods (cluster analysis, principal components analysis, regression analysis and bootstrap cross validation), the estimating equations themselves and a demonstration of the NICM tool suite.
MSL Telecom Automated Anomaly Detection
The Mars Science Laboratory (MSL) Telecom Operations Team at the Jet Propulsion Laboratory (JPL) has implemented a machine learning system in order to automate the anomaly detection process as a part of daily operations. Machine learning enables reliable detection of anomalies in Telecom-related telemetry and automated reporting of Telecom subsystem status, resulting in an 90% reduction in team workload and improved anomaly detection reliability. At present, machine learning methods are used to detect: 1. Anomalous long-term trends in telemetry data 2. Anomalous time-domain evolution of telemetry values Both types of anomalies pose their own unique challenges that are addressed in different ways. In the first case, long term trending of daily minima, maximum, and mean telemetry values in temperatures, currents, voltages, and radio frequency (RF) power levels is used in addition to hard threshold safety checks to look for changes in long-term equipment health and performance. Long-term trending methods allow for ordinary seasonal variations in these quantities caused by temperature changes over the course of the Martian year while allowing operators to determine whether current performance remains in line with historical values from previous years. Changes in long-term trends can provide important insights into the health and status of the rover's on-board systems as well as valuable early warning if subtle degradation begins to take hold. But while trending of daily statistics is valuable, it does not detect anomalies in the short-term time evolution of data over the course of minutes or hours during a day, and this task is handled with short-term shape analysis. Principal components analysis (PCA) has been found to provide robust detection of short-term anomalies, and several examples of the use of PCA to detect actual anomalous events will be provided here. In using PCA, we use both the percentage of explained variance and also a log likelihood test on the PCA expansion coefficients to flag telemetry data for human review. Previous work in the field of spacecraft anomaly detection includes [1] for MSL and [2] for some other JPL missions.
Preliminary Comparisons of the Information Content and Utility of TM Versus MSS Data
Comparisons were made between subscenes from the first TM scene acquired of the Washington, D.C. area and a MSS scene acquired approximately one year earlier. Three types of analyses were conducted to compare TM and MSS data: a water body analysis, a principal components analysis and a spectral clustering analysis. The water body analysis compared the capability of the TM to the MSS for detecting small uniform targets. Of the 59 ponds located on aerial photographs 34 (58%) were detected by the TM with six commission errors (15%) and 13 (22%) were detected by the MSS with three commission errors (19%). The smallest water body detected by the TM was 16 meters; the smallest detected by the MSS was 40 meters. For the principal components analysis, means and covariance matrices were calculated for each subscene, and principal components images generated and characterized. In the spectral clustering comparison each scene was independently clustered and the clusters were assigned to informational classes. The preliminary comparison indicated that TM data provides enhancements over MSS in terms of (1) small target detection and (2) data dimensionality (even with 4-band data). The extra dimension, partially resultant from TM band 1, appears useful for built-up/non-built-up area separation.
Algorithms for Spectral Decomposition with Applications to Optical Plume Anomaly Detection
The analysis of spectral signals for features that represent physical phenomenon is ubiquitous in the science and engineering communities. There are two main approaches that can be taken to extract relevant features from these high-dimensional data streams. The first set of approaches relies on extracting features using a physics-based paradigm where the underlying physical mechanism that generates the spectra is used to infer the most important features in the data stream. We focus on a complementary methodology that uses a data-driven technique that is informed by the underlying physics but also has the ability to adapt to unmodeled system attributes and dynamics. We discuss the following four algorithms: Spectral Decomposition Algorithm (SDA), Non-Negative Matrix Factorization (NMF), Independent Component Analysis (ICA) and Principal Components Analysis (PCA) and compare their performance on a spectral emulator which we use to generate artificial data with known statistical properties. This spectral emulator mimics the real-world phenomena arising from the plume of the space shuttle main engine and can be used to validate the results that arise from various spectral decomposition algorithms and is very useful for situations where real-world systems have very low probabilities of fault or failure. Our results indicate that methods like SDA and NMF provide a straightforward way of incorporating prior physical knowledge while NMF with a tuning mechanism can give superior performance on some tests. We demonstrate these algorithms to detect potential system-health issues on data from a spectral emulator with tunable health parameters.
Dimensionality Reduction Through Classifier Ensembles
In data mining, one often needs to analyze datasets with a very large number of attributes. Performing machine learning directly on such data sets is often impractical because of extensive run times, excessive complexity of the fitted model (often leading to overfitting), and the well-known "curse of dimensionality." In practice, to avoid such problems, feature selection and/or extraction are often used to reduce data dimensionality prior to the learning step. However, existing feature selection/extraction algorithms either evaluate features by their effectiveness across the entire data set or simply disregard class information altogether (e.g., principal component analysis). Furthermore, feature extraction algorithms such as principal components analysis create new features that are often meaningless to human users. In this article, we present input decimation, a method that provides "feature subsets" that are selected for their ability to discriminate among the classes. These features are subsequently used in ensembles of classifiers, yielding results superior to single classifiers, ensembles that use the full set of features, and ensembles based on principal component analysis on both real and synthetic datasets.
Towards Solving the Mixing Problem in the Decomposition of Geophysical Time Series by Independent Component Analysis
The use of the Principal Component Analysis technique for the analysis of geophysical time series has been questioned in particular for its tendency to extract components that mix several physical phenomena even when the signal is just their linear sum. We demonstrate with a data simulation experiment that the Independent Component Analysis, a recently developed technique, is able to solve this problem. This new technique requires the statistical independence of components, a stronger constraint, that uses higher-order statistics, instead of the classical decorrelation a weaker constraint, that uses only second-order statistics. Furthermore, ICA does not require additional a priori information such as the localization constraint used in Rotational Techniques.
Summertime Influence of Asian Pollution in the Free Troposphere over North America
We analyze aircraft observations obtained during INTEX-A (1 July 14 - August 2004) to examine the summertime influence of Asian pollution in the free troposphere over North America. By applying correlation analysis and Principal Component Analysis (PCA) to the observations between 6-12 km, we find dominant influences from recent convection and lightning (13 percent of observations), Asia (7 percent), the lower stratosphere (7 percent), and boreal forest fires (2 percent), with the remaining 71 percent assigned to background. Asian airmasses are marked by high levels of CO, O3, HCN, PAN, acetylene, benzene, methanol, and SO4(2-). The partitioning of reactive nitrogen species in the Asian plumes is dominated by peroxyacetyl nitrate (PAN) (approximately 600 pptv), with varying NO(x)/HNO3 ratios in individual plumes consistent with different plumes ages ranging from 3 to 9 days. Export of Asian pollution in warm conveyor belts of mid-latitude cyclones, deep convection, and lifting in typhoons all contributed to the five major Asian pollution plumes. Compared to past measurement campaigns of Asian outflow during spring, INTEX-A observations display unique characteristics: lower levels of anthropogenic pollutants (CO, propane, ethane, benzene) due to their shorter summer lifetimes; higher levels of biogenic tracers (methanol and acetone) because of a more active biosphere; as well as higher levels of PAN, NO(x), HNO3, and O3 (more active photochemistry possibly enhanced by injection of lightning NO(x)). The high delta O3/delta CO ratio (0.76 mol mol(exp -1)) of Asian plumes during INTEX-A is due to a combination of strong photochemical production and mixing with stratospheric air along isentropic surfaces. The GEOS-Chem global chemical transport model captures the timing and location of the Asian plumes remarkably well. However, it significantly underestimates the magnitude of the enhancements.
Principal components technique analysis for vegetation and land use discrimination
Automatic pre-processing technique called Principal Components (PRINCO) in analyzing LANDSAT digitized data, for land use and vegetation cover, on the Brazilian cerrados was evaluated. The chosen pilot area, 223/67 of MSS/LANDSAT 3, was classified on a GE Image-100 System, through a maximum-likehood algorithm (MAXVER). The same procedure was applied to the PRINCO treated image. PRINCO consists of a linear transformation performed on the original bands, in order to eliminate the information redundancy of the LANDSAT channels. After PRINCO only two channels were used thus reducing computer effort. The original channels and the PRINCO channels grey levels for the five identified classes (grassland, "cerrado", burned areas, anthropic areas, and gallery forest) were obtained through the MAXVER algorithm. This algorithm also presented the average performance for both cases. In order to evaluate the results, the Jeffreys-Matusita distance (JM-distance) between classes was computed. The classification matrix, obtained through MAXVER, after a PRINCO pre-processing, showed approximately the same average performance in the classes separability.
Geobotanical discrimination of ultramafic parent materials An evaluation of remote sensing techniques
Color and color infrared aerial photography and imagery acquired from a Daedalus DEI-1260 multispectral airborne scanner were employed in an investigation to discriminate ultramafic rock types in a test site in southwest Oregon. An analysis of the relationships between vegetation characteristics and parent materials was performed using a vegetation classification and map developed for the project, lithologic information derived from published geologic maps of the region, and terrain information gathered in the field. Several analytical methods, including visual image analysis, band ratioing, principal components analysis, and contrast enhancement and subsequent color composite generation were used in the investigation. There was a close correspondence between vegetation types and major rock types. These were readily discriminated by the remote sensing techniques. It was found that ultramafic rock types were separable from non-ultramafic rock types and serpentine was distinguishable from non-serpentinized peridotite. Further investigations involving spectroradiometric and digital classification techniques are being performed to further identify rock types and to discriminate chromium and nickel-bearing rock types.