Search NASA⌕ Search

SEARCH · Search NASA

Results for “functional principal component analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Elastic functional changepoint detection of climate impacts from localized sources

Detecting changepoints in functional data has become an important problem as interest in monitoring of climate phenomenon has increased, where the data is functional in nature. Here, the observed data often contains both amplitude (y-axis) and phase (x-axis) variability. If not accounted for properly, true changepoints may be undetected, and the estimated underlying mean change functions will be incorrect. In this article, an elastic functional changepoint method is developed which properly accounts for these types of variability. The method can detect amplitude and phase changepoints which current methods in the literature do not, as they focus solely on the amplitude changepoint. This method can easily be implemented using the functions directly or can be computed via functional principal component analysis to ease the computational burden. We apply the method and its nonelastic competitors to both simulated data and observed data to show its efficiency in handling data with phase variation with both amplitude and phase changepoints. We use the method to evaluate potential changes in stratospheric temperature due to the eruption of Mt. Pinatubo in the Philippines in June 1991. Using an epidemic changepoint model, we find evidence of a increase in stratospheric temperature during a period that contains the immediate aftermath of Mt. Pinatubo, with most detected changepoints occurring in the tropics as expected.

54 ENVIRONMENTAL SCIENCES↗

Covariate Dependent Sparse Functional Data Analysis

This study proposes a method to incorporate covariate information into sparse functional data analysis. The method aims at cases where each subject has a limited number of longitudinal measurements and is associated with static covariates. This research is motivated by several use cases in practice. One representative example is void swelling, a nuclear-specific material degradation mechanism. Void swelling is affected by many covariates, including alloy composition and irradiation type. How to accurately model the complicated joint effects of such covariates on the swelling process is the key to mitigating the effect of swelling and ensuring safe operation. Unlike most of the existing methods, the proposed method can handle high-dimensional covariates with the informative covariate identification procedure and sparse and irregularly spaced measurements, that is, does not require complete or dense observations. The main innovation of the proposed method is that we model the variation coming from covariates and the variation left conditioned on covariates, such that the functional principal component analysis and Gaussian process can be conducted in a unified manner. Further, we also propose a systematic approach to identify important covariates in the hypothesis testing context. The methodology is demonstrated on applications in nuclear engineering and healthcare and simulation studies.

42 ENGINEERING↗

Explainable Machine Learning for Functional Data

Black-box machine learning models are recognized as useful tools for prediction applications, but the algorithmic complexity of some models causes interpretation challenges. Explainability methods have been proposed to provide insight into these models, but there is little research focused on supervised modeling with functional data inputs. We argue that, especially in applications of high consequence, it is important to explicitly model the functional dependence in a black-box analysis to not obscure or misrepresent patterns in explanations. As such, we propose the V ariable importance E xplainable E lastic S hape A nalysis (VEESA) pipeline for training supervised machine learning models with functional inputs. The pipeline is an analysis process that includes the data preprocessing, modeling, and post-hoc explanations. The preprocessing is done using elastic functional principal components analysis, which accounts for vertical and horizontal variability in functional data and, ultimately, allows for explanations in the original data space that identify the important functional variability without bias due to correlated variables. Here, we demonstrate the pipeline on two high-consequence applications: explosives classification for national security and inkjet printer identification in forensic science. The applications exhibit the VEESA pipeline’s ability to provide an understanding of the characteristics of the functional data useful for prediction. Code for implementing the pipeline is available in the veesa R package (and supplemental python code).

Elastic Shape Analysis↗

VEESA R package

SAND2024-04584O R package for applying the VEESA pipeline method is a technique used for explainable machine learning with functional data. The VEESA pipeline makes use of the elastic-shape analysis framework for functional data. It also implements functional principal component analysis and permutation feature importance. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525.

Tucker, James↗

High-Throughput Field Plant Phenotyping: A Self-Supervised Sequential CNN Method to Segment Overlapping Plants

High-throughput plant phenotyping—the use of imaging and remote sensing to record plant growth dynamics—is becoming more widely used. The first step in this process is typically plant segmentation, which requires a well-labeled training dataset to enable accurate segmentation of overlapping plants. However, preparing such training data is both time and labor intensive. To solve this problem, we propose a plant image processing pipeline using a self-supervised sequential convolutional neural network method for in-field phenotyping systems. This first step uses plant pixels from greenhouse images to segment nonoverlapping in-field plants in an early growth stage and then applies the segmentation results from those early-stage images as training data for the separation of plants at later growth stages. The proposed pipeline is efficient and self-supervising in the sense that no human-labeled data are needed. We then combine this approach with functional principal components analysis to reveal the relationship between the growth dynamics of plants and genotypes. We show that the proposed pipeline can accurately separate the pixels of foreground plants and estimate their heights when foreground and background plants overlap and can thus be used to efficiently assess the impact of treatments and genotypes on plant growth in a field environment by computer vision techniques. This approach should be useful for answering important scientific questions in the area of high-throughput phenotyping.

59 BASIC BIOLOGICAL SCIENCES↗

Augmenting machine learning of Grad–Shafranov equilibrium reconstruction with Green's functions

This work presents a method for predicting plasma equilibria in tokamak fusion experiments and reactors. The approach involves representing the plasma current as a linear combination of basis functions using principal component analysis of plasma toroidal current densities (J t ) from the EFIT-AI equilibrium database. Then utilizing EFIT's Green's function tables, basis functions are created for the poloidal flux (ψ) and diagnostics generated from the toroidal current (J t ). Similar to the idea of a physics-informed neural network (NN), this physically enforces consistency between ψ, J t , and the synthetic diagnostics. First, the predictive capability of a least squares technique to minimize the error on the synthetic diagnostics is employed. The results show that the method achieves high accuracy in predicting ψ and moderate accuracy in predicting J t with median R 2 = 0.9993 and R 2 = 0.978, respectively. A comprehensive NN using a network architecture search is also employed to predict the coefficients of the basis functions. The NN demonstrates significantly better performance compared to the least squares method with median R 2 = 0.9997 and 0.9916 for J t and ψ, respectively. The robustness of the method is evaluated by handling missing or incorrect data through the least squares filling of missing data, which shows that the NN prediction remains strong even with a reduced number of diagnostics. Additionally, the method is tested on plasmas outside of the training range showing reasonable results.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Machine learning inversion from scattering for mechanically driven polymers

A machine learning inversion method is developed for analyzing scattering functions of mechanically driven polymers and extracting the corresponding feature parameters, which include energy parameters and conformation variables. The polymer is modeled as a chain of fixed-length bonds constrained by bending energy, and it is subject to external forces such as stretching and shear. We generate a data set consisting of random combinations of energy parameters, including bending modulus, stretching and shear force, along with Monte Carlo-calculated scattering functions and conformation variables such as end-to-end distance, radius of gyration and off-diagonal component of the gyration tensor. The effects of the energy parameters on the polymer are captured by the scattering function, and principal component analysis ensures the feasibility of the machine learning inversion. Finally, we train a Gaussian process regressor using part of the data set as a training set and validate the trained regressor for inversion using the rest of the data. The regressor successfully extracts the feature parameters.

Gaussian process regressors↗

Description of sunspot cycles by orthogonal functions

Based on the principal component analysis technique and evidence for a 22-yr double-sunspot cycle periodicity. The time series of sunspot numbers is represented as a sum of mutually orthogonal eigenvectors in the time domain. It is shown that the first two eigenvectors account for about 90 percent of the cumulative 'signal power,' and that this is sufficient for reconstruction of the raw data curve. It is also noted that the second eigenvector behaves as the time derivative of the first, and that a phase-plane plot of these eigenvectors (i.e. a plot of a variable vs. its rate of change) suggests that the sun's sunspot cycle is driven by an oscillator; the implication is that, embedded within the sun, a chronometer is at work (e.g. Dicke, 1979).

Teuber, D. L.↗

Uncertainty Quantification for Smooth Functional Data with Application to Material Properties

This document outlines a method for processing functional output (i.e., curves) for the ultimate purpose of sampling curves under specified input conditions for use in modeling and simulation uncertainty quantification (UQ) studies. A set of benchmark curves sufficiently representative of the relevant scenario(s) being simulated are provided to the process and formatted as described in Section 1. Principal Component Analysis (PCA) is utilized to discover the components of uncertainty in the benchmark curves and is outlined in Section 2. Section 3 describes the application of uncertainty quantification to the PCA results for the purpose of sampling curves to be used in UQ analysis. Section 4 applies these techniques to an example benchmark dataset. Concluding remarks are provided in the final section.

36 MATERIALS SCIENCE↗

Relative effectiveness of kinetic analysis vs single point readings for classifying environmental samples based on community-level physiological profiles (CLPP)

The relative effectiveness of average-well-color-development-normalized single-point absorbance readings (AWCD) vs the kinetic parameters mu(m), lambda, A, and integral (AREA) of the modified Gompertz equation fit to the color development curve resulting from reduction of a redox sensitive dye from microbial respiration of 95 separate sole carbon sources in microplate wells was compared for a dilution series of rhizosphere samples from hydroponically grown wheat and potato ranging in inoculum densities of 1 x 10(4)-4 x 10(6) cells ml-1. Patterns generated with each parameter were analyzed using principal component analysis (PCA) and discriminant function analysis (DFA) to test relative resolving power. Samples of equivalent cell density (undiluted samples) were correctly classified by rhizosphere type for all parameters based on DFA analysis of the first five PC scores. Analysis of undiluted and 1:4 diluted samples resulted in misclassification of at least two of the wheat samples for all parameters except the AWCD normalized (0.50 abs. units) data, and analysis of undiluted, 1:4, and 1:16 diluted samples resulted in misclassification for all parameter types. Ordination of samples along the first principal component (PC) was correlated to inoculum density in analyses performed on all of the kinetic parameters, but no such influence was seen for AWCD-derived results. The carbon sources responsible for classification differed among the variable types with the exception of AREA and A, which were strongly correlated. These results indicate that the use of kinetic parameters for pattern analysis in CLPP may provide some additional information, but only if the influence of inoculum density is carefully considered. c2001 Elsevier Science Ltd. All rights reserved.

NASA Center KSC↗

Regional climate change predictions from the Goddard Institute for Space Studies high resolution GCM

A new diagnostic tool is developed for examining relationships between the synoptic scale circulation and regional temperature distributions in GCMs. The 4 x 5 deg GISS GCM is shown to produce accurate simulations of the variance in the synoptic scale sea level pressure distribution over the U.S. An analysis of the observational data set from the National Meteorological Center (NMC) also shows a strong relationship between the synoptic circulation and grid point temperatures. This relationship is demonstrated by deriving transfer functions between a time-series of circulation parameters and temperatures at individual grid points. The circulation parameters are derived using rotated principal components analysis, and the temperature transfer functions are based on multivariate polynomial regression models. The application of these transfer functions to the GCM circulation indicates that there is considerable spatial bias present in the GCM temperature distributions. The transfer functions are also used to indicate the possible changes in U.S. regional temperatures that could result from differences in synoptic scale circulation between a 1XCO2 and a 2xCO2 climate, using a doubled CO2 version of the same GISS GCM.

Crane, Robert G.↗

SpecDis: Value Added Distance Catalog for 4 Million Stars from DESI Year-1 Data

We present the SpecDis value-added stellar distance catalog accompanying DESI Data Release 1. SpecDis trains a feed-forward neural network (NN) with Gaia parallaxes and gets the distance estimates. To build up an unbiased training sample, we do not apply selections on parallax error or signal-to-noise (S/N) of the stellar spectra, and instead, we incorporate parallax error into the loss function. Moreover, we employ principal component analysis to reduce the noise and dimensionality of stellar spectra. Validated by independent external samples of member stars with precise distances from globular clusters, dwarf galaxies, stellar streams, combined with blue horizontal branch stars, we demonstrate that our distance measurements show no significant bias up to 100 kpc, and are much more precise than Gaia parallax beyond 7 kpc. The median distance uncertainties are 23%, 19%, 11%, and 7% for S/N < 20, 20 ≤ S/N < 60, 60 ≤ S/N < 100, and S/N ≥ 100. Selecting stars with ${\mathrm{log}}\,g\lt 3.8$ and distance uncertainties smaller than 25%, we have more than 74,000 giant candidates within 50 kpc of the Galactic center and 1500 candidates beyond this distance. Additionally, we develop a Gaussian mixture model to identify unresolvable equal-mass binaries by modeling the discrepancy between the NN-predicted and the geometric absolute magnitudes from Gaia parallaxes and identify 120,000 equal-mass binary candidates. Our final catalog provides distances and distance uncertainties for >4 million stars, offering a valuable resource for Galactic astronomy.

astronomy data analysis↗

A unified development of several techniques for the representation of random vectors and data sets

Linear vector space theory is used to develop a general representation of a set of data vectors or random vectors by linear combinations of orthonormal vectors such that the mean squared error of the representation is minimized. The orthonormal vectors are shown to be the eigenvectors of an operator. The general representation is applied to several specific problems involving the use of the Karhunen-Loeve expansion, principal component analysis, and empirical orthogonal functions; and the common properties of these representations are developed.

Bundick, W. T.↗

PCAfold 2.0—Novel tools and algorithms for low-dimensional manifold assessment and optimization

We describe an update to our open-source Python package, PCAfold, designed to help researchers generate, analyze and improve low-dimensional data manifolds. In the current version, PCAfold 2.0, we introduce novel tools and algorithms for assessing and optimizing low-dimensional manifolds. This includes a method that generates a “map” of local feature sizes that can help pinpoint researchers to problematic regions on a manifold. We introduce a novel cost function that characterizes the quality of a manifold topology with a single number. We develop two algorithms for feature selection based on principal component analysis (PCA) that use the cost function as an objective function to minimize. We introduce a quantity of interest (QoI)-aware dimensionality reduction strategy where data projections are computed using an artificial neural network and are directly optimized towards representing various projection-independent and projection-dependent QoIs. We also introduce an implementation of partition of unity networks (POUnets) for efficient reconstruction of QoIs from low-dimensional manifolds based on combining neural network classification with localized polynomial regression. Our software can be broadly applicable in all domains of science and engineering that aim to reduce data dimensionality, as well as in the fundamental research on representation learning.

97 MATHEMATICS AND COMPUTING↗

Application of Spectral Analysis Techniques in the Intercomparison of Aerosol Data: 1. an EOF Approach to the Spatial-Temporal Variability of Aerosol Optical Depth Using Multiple Remote Sensing Data Sets

Many remote sensing techniques and passive sensors have been developed to measure global aerosol properties. While instantaneous comparisons between pixel-level data often reveal quantitative differences, here we use Empirical Orthogonal Function (EOF) analysis, also known as Principal Component Analysis, to demonstrate that satellite-derived aerosol optical depth (AOD) data sets exhibit essentially the same spatial and temporal variability and are thus suitable for large-scale studies. Analysis results show that the first four EOF modes of AOD account for the bulk of the variance and agree well across the four data sets used in this study (i.e., Aqua MODIS, Terra MODIS, MISR, and SeaWiFS). Only SeaWiFS data over land have slightly different EOF patterns. Globally, the first two EOF modes show annual cycles and are mainly related to Sahara dust in the northern hemisphere and biomass burning in the southern hemisphere, respectively. After removing the mean seasonal cycle from the data, major aerosol sources, including biomass burning in South America and dust in West Africa, are revealed in the dominant modes due to the different interannual variability of aerosol emissions. The enhancement of biomass burning associated with El Niño over Indonesia and central South America is also captured with the EOF technique.

variability↗

Applications of array processors in the analysis of remote sensing images

The architectures, programming characteristics, and ranges of application of past, present, and planned array processors for the digital processing of remote-sensing images are compared. Such functions as radiometric and geometric corrections, principal-components analysis, cluster coding, histogram generation, grey-level mapping, convolution, classification, and mensuration and modeling operations are considered, and both pipeline-type and single-instruction/multiple-data-stream (SIMD) arrays are evaluated. Numerical results are presented in a table, and it is found that the pipeline-type arrays normally used with minicomputers increase their speed significantly at low cost, while even further gains are provided by the more expensive SIMD arrays. Most image-processing operations become I/O-limited when SIMD arrays are used with current I/O devices.

Ramapriyan, H. K.↗

Sea Ice Motion from Wavelet Analysis of Satellite Data

Wavelet analysis of NASA scatterometer (NSCAT) backscatter and Defense Meteorological Satellite Program (DMSP) Special Sensor Microwave/Imager (SSM/I) radiance data can be used to obtain daily sea ice drift information for the Arctic region. This technique provides improved spatial coverage over the existing array of Arctic Ocean buoys and better temporal resolution over techniques utilizing data from satellite synthetic aperture radars. Comparisons with ice motion derived from ocean buoys give good quantitative agreement. Both comparison results from NSCAT and SSM/I are compatible, and the results from NSCAT can definitely complement that from SSM/I when there are cloud or surface effects. Then three sea-ice drift daily results from NSCAT, SSM/I, and buoy data can be merged as a composite map by some data fusion techniques. The ice flow streamlines are highly correlated with surface air pressure contours. Examples of derived ice-drift maps in December 1996 illustrate large-scale circulation reversals over a period of four days. A method for deriving divergence and shear at the large-scale has been developed and comparison between buoys and satellite results shows a good agreement. These calibrated/validated results indicate that NSCAT, SSM/I merged daily ice motion are suitably accurate to identify and closely locate sea ice processes, and to improve our current knowledge of sea ice drift and related processes through the data assimilation of ocean-ice numerical model. For demonstration purpose, the ice velocities derived from satellite data are compared with the ice velocities derived from a coupled ice-ocean interaction model. The comparison reveals that the general circulation patterns of the two are quite similar but the ice velocity differences between the two are quite significant. In order to quantify the wind effects on ice motion, empirical orthogonal functions (EOF) are used in the principal component analysis for both ice motion and pressure field. Some preliminary results of sea-ice motion from QuikScat will also be presented.

Liu, Antony K.↗

Intraseasonal oscillations in the global atmosphere. I - Northern Hemisphere and tropics

Oscillatory modes in the Northern Hemisphere and in the tropics were examined systematically. The 700 mb heights were used to analyze extratropical oscillations, and the outgoing longwave radiation to study tropical oscillations in convection. All datasets were band-pass filtered to focus on the intraseasonal (IS) band of 10-120 days. Leading spatial patterns of variability were obtained by applying empirical orthogonal function analysis to these IS data. The leading principal components were subjected to singular spectrum analysis. In the Northern Hemisphere, there are two important modes of oscillation with periods near 48 and 23 days, respectively. The 48-day mode is the most important of the two. It has both traveling and standing components, and is dominated by a zonal wavenumber two. The 23-day mode has the spatial structure and propagation properties described by Branstator and by Kushnir (1987). In the tropics, the 40-50 day oscillation documented by Madden and Julian (1972), Weickmann (1983), Lau, and their colleagues, dominates the Indian and Pacific oceans from 60 deg E to the date line. From 170 deg W to 90 deg W, however, a 24-28 day oscillation is equally strong. The extratropical modes are often independent of, and sometimes lead, the tropical modes.

Ghil, Michael↗