Search NASA⌕ Search

SEARCH · Search NASA

Results for “principal component analysis (PCA)”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Using Principal Component Analysis (PCA) to Speed up Radiative Transfer (RT) Computations

Multiple scattering RT calculations time-consuming. Need a speed improvement of about 1000 (for OCO)! Solution: Make use of redundancies in spectra. Correlated-k (Lacis and Wang, Lacis and Oinas, Goody et al, Fu and Liou) Problem: Assume that spectral variation of atmospheric optical properties spatially correlated at all points along optical path. High accuracy (HI) and 2-stream (2S) calculations have high correlation. Single scattering (SS) computations highly scenario-dependent, but not time consuming. Perform SS and 2S calculations at every wavelength. Perform small number of HI computations. Need to compute correction factor B at every wavelength.

data sets↗

Discerning the Impact of Powder Feedstock Variability on Structure, Property, and Performance of Selective Laser Melted Alloy 718: A Principal Component Analysis (PCA) of Feedstock Variability

Extensive mechanical, chemical and microstructural analyses were conducted on additively manufactured Alloy 718 to characterize powders from multiple vendors to determine the effects of variations observed in the powders had on the consolidated material. With over 190 variables examined, it was necessary to reduce the number of variables and identify the variables and classes of variables that had the greatest effect. Principle Component Analysis (PCA) was used to reduce the number of variable to effectively 12 while identifying several classes of variables as most important.

Ellis, David↗

Data structure characterization of miltispectral data using principal component and principal factor analysis

Both principal component analysis (PCA) and principal factor analysis (PFA) were used to analyze an experimental multispectral data structure in terms of common and unique variance. Only the common variance of the multispectral data was associated with the principal factor, while higher-order principal components were associated with both common and unique variance. The unique variance was found to represent small spectral variations within each cover type as well as noise vectors, and was most abundant in the lower-order principal components. The lower-order principal components can be useful in research designed to discriminate minor physical variations within features, and to highlight localized change when using multitemporal-multispectral data. Conversely, PFA of the multispectral data provided an insight into a great potential for discriminating basic land-cover types by excluding the unique variance which was related to the noise and minor spectral variations.

Lee, Jae K.↗

SO(3)-invariant PCA with application to molecular data

Principal component analysis (PCA) is a fundamental technique for dimensionality reduction and denoising; however, its application to three-dimensional data with arbitrary orientations -- common in structural biology -- presents significant challenges. A naive approach requires augmenting the dataset with many rotated copies of each sample, incurring prohibitive computational costs. In this paper, we extend PCA to 3D volumetric datasets with unknown orientations by developing an efficient and principled framework for SO(3)-invariant PCA that implicitly accounts for all rotations without explicit data augmentation. By exploiting underlying algebraic structure, we demonstrate that the computation involves only the square root of the total number of covariance entries, resulting in a substantial reduction in complexity. We validate the method on real-world molecular datasets, demonstrating its effectiveness and opening up new possibilities for large-scale, high-dimensional reconstruction problems.

Fraiman, Michael [Tel Aviv Univ., Tel Aviv (Israel↗

Continuum Fitting HST QSO Spectra

The Principal Component Analysis (PCA) method which we are using to fit and describe QSO spectra relies upon the fact that QSO continuum are generally very smooth and simple except for emission and absorption lines. To see this we need high signal-to-noise (S/N) spectra of QSOs at low redshift which have relatively few absorption lines in the Lyman-a forest. We need a large number of such spectra to use as the basis set for the PCA analysis which will find the set of principal component spectra which describe the QSO family as a whole. We have found that too few HST spectra have the required S/N and hence we need to supplement them with ground based spectra of QSOs at higher redshift. We have many such spectra and we have been working to make them suitable for this analysis. We have concentrated on this topic since 12/15/01.

Tytler, David↗

What drives the variance of galaxy spectra?

We present a study aimed at understanding the physical phenomena underlying the formation and evolution of galaxies following a data-driven analysis of spectroscopic data based on the variance in a carefully selected sample. We apply principal component analysis (PCA) independently to three subsets of continuum-subtracted optical spectra, segregated into their nebular emission activity as quiescent, star-forming, and active galactic nuclei (AGNs). We emphasize that the variance of the input data in this work only relates to the absorption lines in the photospheres of the stellar populations. The sample is taken from the Sloan Digital Sky Survey (SDSS) in the stellar velocity dispersion range 100–150 km s −1 , to minimize the ‘blurring’ effect of the stellar motion. We restrict the analysis to the first three principal components (PCs) and find that PCA segregates the three types with the highest variance mapping SSP-equivalent age, along with an inextricable degeneracy with metallicity, even when all three PCs are included. Spectral fitting shows that stellar age dominates PC1, whereas PC2 and PC3 have a mixed dependence of age and metallicity. The trends support – independently of any model fitting – the hypothesis of an evolutionary sequence from star formation to AGN to quiescence. As a further test of the consistency of the analysis, we apply the same methodology in different spectral windows, finding similar trends, but the variance is maximal in the blue wavelength range, roughly around the 4000 Å break.

79 ASTRONOMY AND ASTROPHYSICS↗

Mapping Rare Earths and Toxics in E-Waste via Hyperspectral Imaging and Machine Learning

Electronic waste (e-waste) presents a mounting challenge to environmental sustainability due to its complex composition, which includes high-value rare earth elements, hazardous organic compounds, and non-recyclable plastics. Accurate and scalable material classification is essential for enabling efficient resource recovery and safe recycling practices. This study introduces a confidence-aware classification pipeline that combines mid-infrared hyperspectral imaging (HSI), spectral angle mapping (SAM), and iterative machine learning to perform pixel-level material identification across e-waste devices. A curated spectral library encompassing artificial materials (e.g., plastic iron oxide, galvanized metals), minerals (e.g., allanite, hematite), and organic compounds (e.g., benzanthracene, toluene) was used to generate pseudo-labels, each assigned a confidence score based on SAM-derived spectral similarity. High-confidence samples from seven consumer electronics—digital cameras, keyboards, laptop fans, modems, motherboards, TV remotes, and speakers—were iteratively expanded and classified using models such as Support Vector Machine (SVM), Random Forest, Gradient Boosting Classifier, Partial Least Squares Discriminant Analysis (PLSDA) and Logistic Regression. The best-performing classifiers achieved macro F1 scores approaching 1.0. Results revealed widespread plastic content (dominated by plastic iron oxide), the presence of rare earth-bearing minerals like cerium-containing allanite, and pervasive detection of hazardous organics such as benzanthracene. Principal Component Analysis (PCA) visualizations and confusion matrices confirmed high separability and robust classification performance. This methodology enables precise, non-destructive, and scalable classification of heterogeneous e-waste streams. It supports automated, hazard-aware sorting in recycling workflows, facilitating selective recovery of critical materials and compliance with circular economy goals. The confidence-aware framework provides a foundation for real-time deployment in industrial settings, offering significant implications for smart e-recycling infrastructure and policy-driven material stewardship.

Circular economy↗

Structuring Nutrient Yields throughout Mississippi/Atchafalaya River Basin Using Machine Learning Approaches

To minimize the eutrophication pressure along the Gulf of Mexico or reduce the size of the hypoxic zone in the Gulf of Mexico, it is important to understand the underlying temporal and spatial variations and correlations in excess nutrient loads, which are strongly associated with the formation of hypoxia. This study’s objective was to reveal and visualize structures in high-dimensional datasets of nutrient yield distributions throughout the Mississippi/Atchafalaya River Basin (MARB). For this purpose, the annual mean nutrient concentrations were collected from thirty-three US Geological Survey (USGS) water stations scattered in the upper and lower MARB from 1996 to 2020. Eight surface water quality indicators were selected to make comparisons among water stations along the MARB over the past two decades. Principal component analysis (PCA) was used to comprehensively evaluate the nutrient yields across thirty-three USGS monitoring stations and identify the major contributing nutrient loads. The results showed that all samples could be analyzed using two main components, which accounted for 81.6% of the total variance. The PCA results showed that yields of orthophosphate (OP), silica (SI), nitrate–nitrites (NO 3 -NO 2 ), and total suspended sediment (TSS) are major contributors to nutrient yields. It also showed that land-planted crops, density of population, domestic and industrial discharges, and precipitation are fundamental causes of excess nutrient loads in MARB. These factors are of great significance for the excess nutrient load management and pollution control of the Mississippi River. It was found that the average nutrient yields were stable within the sub-MARB area, but the large nitrogen yields in the upper MARB and the large phosphorus yields in the lower MARB were of great concern. t-distributed stochastic neighbor embedding (t-SNE) revealed interesting nonlinear and local structures in nutrient yield distributions. Clustering analysis (CA) showed the detailed development of similarities in the nutrient yield distribution. Moreover, PCA, t-SNE, and CA showed consistent clustering results. This study demonstrated that the integration of dimension reduction techniques, PCA, and t-SNE with CA techniques in machine learning are effective tools for the visualization of the structures of the correlations in high-dimensional datasets of nutrient yields and provide a comprehensive understanding of the correlations in the distributions of nutrient loads across the MARB.

54 ENVIRONMENTAL SCIENCES↗

Spatiotemporal Filtering Using Principal Component Analysis and Karhunen-Loeve Expansion Approaches for Regional GPS Network Analysis

Spatial filtering is an effective way to improve the precision of coordinate time series for regional GPS networks by reducing so-called common mode errors, thereby providing better resolution for detecting weak or transient deformation signals. The commonly used approach to regional filtering assumes that the common mode error is spatially uniform, which is a good approximation for networks of hundreds of kilometers extent, but breaks down as the spatial extent increases. A more rigorous approach should remove the assumption of spatially uniform distribution and let the data themselves reveal the spatial distribution of the common mode error. The principal component analysis (PCA) and the Karhunen-Loeve expansion (KLE) both decompose network time series into a set of temporally varying modes and their spatial responses. Therefore they provide a mathematical framework to perform spatiotemporal filtering.We apply the combination of PCA and KLE to daily station coordinate time series of the Southern California Integrated GPS Network (SCIGN) for the period 2000 to 2004. We demonstrate that spatially and temporally correlated common mode errors are the dominant error source in daily GPS solutions. The spatial characteristics of the common mode errors are close to uniform for all east, north, and vertical components, which implies a very long wavelength source for the common mode errors, compared to the spatial extent of the GPS network in southern California. Furthermore, the common mode errors exhibit temporally nonrandom patterns.

displacement↗

Portland Urban Development: Quantifying and Visualizing Urban Heat with Compounding Vulnerabilities to Support Community Depaving Initiatives

Urban heat is a pressing concern in Portland, Oregon as climate change induced heat waves increase. Cities experience higher temperatures due to the urban heat island effect (UHI), and environmental injustice and disenfranchisement in minority communities expose low-income and Black, Indigenous, and People of Color (BIPOC) residents to more extreme and debilitating heat events. Our team identified Portland’s communities on the frontlines of urban heat impacts by overlapping environmental and social vulnerabilities using NASA Earth observations. We partnered with Depave, a Portland-based nonprofit that works alongside communities to replace pavement with greenspace in historically disenfranchised areas. Using Landsat 8 Thermal Infrared Sensor (TIRS) imagery, we mapped Land Surface Temperature (LST) and developed a heat-specific Social Vulnerability Index (SVI) through a Principal Component Analysis (PCA) to identify Portland’s communities with the highest potential heat vulnerability. Then, we calculated the temperature change of depaving in six case studies to quantify Depave's efforts in heat mitigation and environmental justice. Our analysis demonstrated that, throughout Portland, there are frontline communities experiencing high potential social vulnerability to extreme temperatures due to environmental injustices and over-pavement. Finally, Depave’s impact on urban heat is observable and quantifiable using remote-sensing data and tools, with an average of 1ºF LST decrease across the six case studies. We illustrated the significance of local urban heat mitigation efforts and propose next steps for conducting inclusive and intentional research that highlights the lived experiences and resilience of frontline communities.

Environmental justice↗

Reducing the matrix effect in mass spectral imaging of biofilms using flow-cell culture

The interactions between soil microorganisms and soil minerals play a crucial role in the formation and evolution of minerals and the stability of soil aggregates. Due to the heterogeneity and diversity of the soil environment, the under-standing of the functions of bacterial biofilms in soil minerals at the microscale is limited. A soil mineral-bacterial biofilm system was used as a model in this study, and it was analyzed by time-of-flight secondary ion mass spectrometry (ToF-SIMS) to acquire molecular level information. Static culture in multi-wells and dynamic flow-cell culture in microfluidics of biofilms were investigated. Our results show that more characteristic molecules of biofilms can be observed in SIMS spectra of the flow-cell culture. In contrast, biofilm signature peaks are buried under the mineral components in SIMS spectra in the static culture case. Spectral overlay was used in peak selection prior to performing Principal component analysis (PCA). Comparisons of the PCA results between the static and flow-cell culture show more pronounced molecular features and higher loadings of organic peaks of the dynamic cultured specimens. For example, fatty acids secreted from bacterial biofilm extracellular polymeric substance are likely to be responsible for biofilm dispersal due to mineral treatment up to 48 h. Such findings suggest that the use of microfluidic cells to dynamically culture biofilms be a more suitable method for reducing the matrix effect arisen from the growth medium and minerals as a perturbation fac-tor for improved spectral and multivariate analysis of complex mass spectral data in ToF-SIMS. These results show that the interaction mechanism between biofilms and soil minerals at the molecular level can be better studied using the flow-cell culture and advanced mass spectral imaging techniques like ToF-SIMS.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Predicting non-linear stress–strain response of mesostructured cellular materials using supervised autoencoder

Recent breakthroughs in advanced manufacturing capabilities have made it possible to design and print sophisticated topologies of cellular structures using diverse engineering materials such as metals, polymers, and ceramics. In these architectured materials, it is often desirable to tailor the mechanical properties by altering the unit cell topology. This necessitates an in-depth understanding of how the topology of the unit cell structure affects the macroscopic behavior of the material in both the linear and the non-linear regimes encountered under large compression. Here, we have developed a machine learning (ML) approach capable of accelerating the prediction of the stress–strain response of a polymer-based cellular structure under uniaxial confined compression. As part of generating the training data for ML, 60,000 mesostructures were generated using a relatively novel approach based on cellular automata, and their corresponding stress–strain responses were obtained from the finite element simulations. Principal component analysis (PCA) was used to reduce the dimensionality of the stress–strain curves. With only 20 principal components, PCA captured 99.89% of the variance in the stress–strain curves while reducing the dimensionality by 5X. ML using supervised autoencoder was able to successfully speed up the prediction of the non-linear stress–strain response of a unit cell by up to 4600X. The proposed method can serve as an efficient data generation tool and a rapid means for predicting the structure–property relationship through accelerated forward modeling of cellular materials under compaction, in cases where the macroscopic stress–strain response is governed by the unit-cell topology.

36 MATERIALS SCIENCE↗

Forward and Inverse Models for Satellite Remote Sensors using Principal Component Analysis

Satellite remote sensors such as AIRS on Aqua, CrIS on S-NPP, NOAA20 and JPSS-2, IASI on Metop A, B, and C make millions of observations each day with thousands of spectral channels for each observation; this poses challenges for efficiently inversion of the inherently large dataset as needed to retrieve atmospheric and surface properties. This presentation will illustrate the use of Principal Component Analysis (PCA) to speed up radiative transfer forward model calculations and to stabilize the inversion algorithms. A Principal Component-based radiative transfer model (PCRTM) developed at NASA Langley Research Center can simulate top of atmosphere (TOA) radiance or reflectance spectra from 50 cm-1 to 50000 cm-1 (200 m to 0.20 m quickly and accurately. PCRTM demonstrated very high accuracy relative to reference line-by-line radiative transfer models and it saves orders of magnitude computational time. Examples of the PCRTM model developed for hyperspectral sensors such as AIRS, CrIS, IASI, NAST-I, SHIS, CPF, TEMPO, SBG, OMI, and SCIAMACHY will be presented. In addition to using the PCRTM as forward model, the NASA Langley developed inversion algorithm also uses PCA to compress the state vector into a compressed dimension to speed up and stabilize the inversion process. Examples of retrieved atmospheric temperature, water vapor, CO2, CO, CH4, N2O, and O3 profiles, cloud properties (optical depth, size, phase, and height), and surface properties (surface emissivity spectra and skin temperatures) will be presented. This algorithm is being transitioned to the NASA Sounder SIPS and NASA's Goddard Earth Sciences Data and Information Services Center (GES DISC).

forward model↗

Automated Analysis, Classification, and Display of Waveforms

A computer program partly automates the analysis, classification, and display of waveforms represented by digital samples. In the original application for which the program was developed, the raw waveform data to be analyzed by the program are acquired from space-shuttle auxiliary power units (APUs) at a sampling rate of 100 Hz. The program could also be modified for application to other waveforms -- for example, electrocardiograms. The program begins by performing principal-component analysis (PCA) of 50 normal-mode APU waveforms. Each waveform is segmented. A covariance matrix is formed by use of the segmented waveforms. Three eigenvectors corresponding to three principal components are calculated. To generate features, each waveform is then projected onto the eigenvectors. These features are displayed on a three-dimensional diagram, facilitating the visualization of the trend of APU operations.

Kwan, Chiman↗

Assessing Wetland Hydroperiod and Soil Moisture With Remote Sensing: A Demonstration for the NASA Plum Brook Station Year 2

Primary Goal: Assist with the evaluation and measuring of wetlands hydroperiod at the PlumBrook Station using multi-source remote sensing data as part of a larger effort on projecting climate change-related impacts on the station's wetland ecosystems. MTRI expanded on the multi-source remote sensing capabilities to help estimate and measure hydroperiod and the relative soil moisture of wetlands at NASA's Plum Brook Station. Multi-source remote sensing capabilities are useful in estimating and measuring hydroperiod and relative soil moisture of wetlands. This is important as a changing regional climate has several potential risks for wetland ecosystem function. The year two analysis built on the first year of the project by acquiring and analyzing remote sensing data for additional dates and types of imagery, combined with focused field work. Five deliverables were planned and completed: 1) Show the relative length of hydroperiod using available remote sensing datasets 2) Date linked table of wetlands extent over time for all feasible non-forested wetlands 3) Utilize LIDAR data to measure topographic height above sea level of all wetlands, wetland to catchment area radio, slope of wetlands, and other useful variables 4) A demonstration of how analyzed results from multiple remote sensing data sources can help with wetlands vulnerability assessment 5) A MTRI style report summarizing year 2 results. This report serves as a descriptive summary of our completion of these our deliverables. Additionally, two formal meetings were held with Larry Liou and Amanda Sprinzl to provide project updates and receive direction on outputs. These were held on 2/26/15 and 9/17/15 at the Plum Brook Station. Principal Component Analysis (PCA) is a multivariate statistical technique used to identify dominant spatial and temporal backscatter signatures. PCA reduces the information contained in the temporal dataset to the first few new Principal Component (PC) images. Some advantages of PCA include the ability to filter out temporal autocorrelation and reduce speckle to the higher order PC images. A PCA was performed using ERDAS Imagine on a time series of PALSAR dates. Hydroperiod maps were created by separating the PALSAR dates into two date ranges, 2006-2008 and 2010, and performing an unsupervised classification on the PCAs.

Brooks, Colin↗

Principal Component Analysis of Arctic Solar Irradiance Spectra

During the FIRE (First ISCPP Regional Experiment) Arctic Cloud Experiment and coincident SHEBA (Surface Heat Budget of the Arctic Ocean) campaign, detailed moderate resolution solar spectral measurements were made to study the radiative energy budget of the coupled Arctic Ocean - Atmosphere system. The NASA Ames Solar Spectral Flux Radiometers (SSFRs) were deployed on the NASA ER-2 and at the SHEBA ice camp. Using the SSFRs we acquired continuous solar spectral irradiance (380-2200 nm) throughout the atmospheric column. Principal Component Analysis (PCA) was used to characterize the several tens of thousands of retrieved SSFR spectra and to determine the number of independent pieces of information that exist in the visible to near-infrared solar irradiance spectra. It was found in both the upwelling and downwelling cases that almost 100% of the spectral information (irradiance retrieved from 1820 wavelength channels) was contained in the first six extracted principal components. The majority of the variability in the Arctic downwelling solar irradiance spectra was explained by a few fundamental components including infrared absorption, scattering, water vapor and ozone. PCA analysis of the SSFR upwelling Arctic irradiance spectra successfully separated surface ice and snow reflection from overlying cloud into distinct components.

Rabbette, Maura↗

DEEPEN Global Standardized Categorical Exploration Datasets for Magmatic Plays

DEEPEN stands for DE-risking Exploration of geothermal Plays in magmatic ENvironments. As part of the development of the DEEPEN 3D play fairway analysis (PFA) methodology for magmatic plays (conventional hydrothermal, superhot EGS, and supercritical), weights needed to be developed for use in the weighted sum of the different favorability index models produced from geoscientific exploration datasets. This was done using two different approaches: one based on expert opinions, and one based on statistical learning. This GDR submission includes the datasets used to produce the statistical learning-based weights. While expert opinions allow us to include more nuanced information in the weights, expert opinions are subject to human bias. Data-centric or statistical approaches help to overcome these potential human biases by focusing on and drawing conclusions from the data alone. The drawback is that, to apply these types of approaches, a dataset is needed. Therefore, we attempted to build comprehensive standardized datasets mapping anomalies in each exploration dataset to each component of each play. This data was gathered through a literature review focused on magmatic hydrothermal plays along with well-characterized areas where superhot or supercritical conditions are thought to exist. Datasets were assembled for all three play types, but the hydrothermal dataset is the least complete due to its relatively low priority. For each known or assumed resource, the dataset states what anomaly in each exploration dataset is associated with each component of the system. The data is only a semi-quantitative, where values are either high, medium, or low, relative to background levels. In addition, the dataset has significant gaps, as not every possible exploration dataset has been collected and analyzed at every known or suspected geothermal resource area, in the context of all possible play types. The following training sites were used to assemble this dataset: - Conventional magmatic hydrothermal: Akutan (from AK PFA), Oregon Cascades PFA, Glass Buttes OR, Mauna Kea (from HI PFA), Lanai (from HI PFA), Mt St Helens Shear Zone (from WA PFA), Wind River Valley (From WA PFA), Mount Baker (from WA PFA). - Superhot EGS: Newberry (EGS demonstration project), Coso (EGS demonstration project), Geysers (EGS demonstration project), Eastern Snake River Plain (EGS demonstration project), Utah FORGE, Larderello, Kakkonda, Taupo Volcanic Zone, Acoculco, Krafla. - Supercritical: Coso, Geysers, Salton Sea, Larderello, Los Humeros, Taupo Volcanic Zone, Krafla, Reyjanes, Hengill. **Disclaimer: Treat the supercritical fluid anomalies with skepticism. They are based on assumptions due to the general lack of confirmed supercritical fluid encounters and samples at the sites included in this dataset, at the time of assembling the dataset. The main assumption was that the supercritical fluid in a given geothermal system has shared properties with the hydrothermal fluid, which may not be the case in reality. Once the datasets were assembled, principal component analysis (PCA) was applied to each. PCA is an unsupervised statistical learning technique, meaning that labels are not required on the data, that summarized the directions of variance in the data. This approach was chosen because our labels are not certain, i.e., we do not know with 100% confidence that superhot resources exist at all the assumed positive areas. We also do not have data for any known non-geothermal areas, meaning that it would be challenging to apply a supervised learning technique. In order to generate weights from the PCA, an analysis of the PCA loading values was conducted. PCA loading values represent how much a feature is contributing to each principal component, and therefore the overall variance in the data.

15 GEOTHERMAL ENERGY↗