Search NASA⌕ Search

SEARCH · Search NASA

Results for “functional principal component analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Elastic functional changepoint detection of climate impacts from localized sources

Detecting changepoints in functional data has become an important problem as interest in monitoring of climate phenomenon has increased, where the data is functional in nature. Here, the observed data often contains both amplitude (y-axis) and phase (x-axis) variability. If not accounted for properly, true changepoints may be undetected, and the estimated underlying mean change functions will be incorrect. In this article, an elastic functional changepoint method is developed which properly accounts for these types of variability. The method can detect amplitude and phase changepoints which current methods in the literature do not, as they focus solely on the amplitude changepoint. This method can easily be implemented using the functions directly or can be computed via functional principal component analysis to ease the computational burden. We apply the method and its nonelastic competitors to both simulated data and observed data to show its efficiency in handling data with phase variation with both amplitude and phase changepoints. We use the method to evaluate potential changes in stratospheric temperature due to the eruption of Mt. Pinatubo in the Philippines in June 1991. Using an epidemic changepoint model, we find evidence of a increase in stratospheric temperature during a period that contains the immediate aftermath of Mt. Pinatubo, with most detected changepoints occurring in the tropics as expected.

54 ENVIRONMENTAL SCIENCES↗

Covariate Dependent Sparse Functional Data Analysis

This study proposes a method to incorporate covariate information into sparse functional data analysis. The method aims at cases where each subject has a limited number of longitudinal measurements and is associated with static covariates. This research is motivated by several use cases in practice. One representative example is void swelling, a nuclear-specific material degradation mechanism. Void swelling is affected by many covariates, including alloy composition and irradiation type. How to accurately model the complicated joint effects of such covariates on the swelling process is the key to mitigating the effect of swelling and ensuring safe operation. Unlike most of the existing methods, the proposed method can handle high-dimensional covariates with the informative covariate identification procedure and sparse and irregularly spaced measurements, that is, does not require complete or dense observations. The main innovation of the proposed method is that we model the variation coming from covariates and the variation left conditioned on covariates, such that the functional principal component analysis and Gaussian process can be conducted in a unified manner. Further, we also propose a systematic approach to identify important covariates in the hypothesis testing context. The methodology is demonstrated on applications in nuclear engineering and healthcare and simulation studies.

42 ENGINEERING↗

Explainable Machine Learning for Functional Data

Black-box machine learning models are recognized as useful tools for prediction applications, but the algorithmic complexity of some models causes interpretation challenges. Explainability methods have been proposed to provide insight into these models, but there is little research focused on supervised modeling with functional data inputs. We argue that, especially in applications of high consequence, it is important to explicitly model the functional dependence in a black-box analysis to not obscure or misrepresent patterns in explanations. As such, we propose the V ariable importance E xplainable E lastic S hape A nalysis (VEESA) pipeline for training supervised machine learning models with functional inputs. The pipeline is an analysis process that includes the data preprocessing, modeling, and post-hoc explanations. The preprocessing is done using elastic functional principal components analysis, which accounts for vertical and horizontal variability in functional data and, ultimately, allows for explanations in the original data space that identify the important functional variability without bias due to correlated variables. Here, we demonstrate the pipeline on two high-consequence applications: explosives classification for national security and inkjet printer identification in forensic science. The applications exhibit the VEESA pipeline’s ability to provide an understanding of the characteristics of the functional data useful for prediction. Code for implementing the pipeline is available in the veesa R package (and supplemental python code).

Elastic Shape Analysis↗

VEESA R package

SAND2024-04584O R package for applying the VEESA pipeline method is a technique used for explainable machine learning with functional data. The VEESA pipeline makes use of the elastic-shape analysis framework for functional data. It also implements functional principal component analysis and permutation feature importance. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525.

Tucker, James↗

High-Throughput Field Plant Phenotyping: A Self-Supervised Sequential CNN Method to Segment Overlapping Plants

High-throughput plant phenotyping—the use of imaging and remote sensing to record plant growth dynamics—is becoming more widely used. The first step in this process is typically plant segmentation, which requires a well-labeled training dataset to enable accurate segmentation of overlapping plants. However, preparing such training data is both time and labor intensive. To solve this problem, we propose a plant image processing pipeline using a self-supervised sequential convolutional neural network method for in-field phenotyping systems. This first step uses plant pixels from greenhouse images to segment nonoverlapping in-field plants in an early growth stage and then applies the segmentation results from those early-stage images as training data for the separation of plants at later growth stages. The proposed pipeline is efficient and self-supervising in the sense that no human-labeled data are needed. We then combine this approach with functional principal components analysis to reveal the relationship between the growth dynamics of plants and genotypes. We show that the proposed pipeline can accurately separate the pixels of foreground plants and estimate their heights when foreground and background plants overlap and can thus be used to efficiently assess the impact of treatments and genotypes on plant growth in a field environment by computer vision techniques. This approach should be useful for answering important scientific questions in the area of high-throughput phenotyping.

59 BASIC BIOLOGICAL SCIENCES↗

Augmenting machine learning of Grad–Shafranov equilibrium reconstruction with Green's functions

This work presents a method for predicting plasma equilibria in tokamak fusion experiments and reactors. The approach involves representing the plasma current as a linear combination of basis functions using principal component analysis of plasma toroidal current densities (J t ) from the EFIT-AI equilibrium database. Then utilizing EFIT's Green's function tables, basis functions are created for the poloidal flux (ψ) and diagnostics generated from the toroidal current (J t ). Similar to the idea of a physics-informed neural network (NN), this physically enforces consistency between ψ, J t , and the synthetic diagnostics. First, the predictive capability of a least squares technique to minimize the error on the synthetic diagnostics is employed. The results show that the method achieves high accuracy in predicting ψ and moderate accuracy in predicting J t with median R 2 = 0.9993 and R 2 = 0.978, respectively. A comprehensive NN using a network architecture search is also employed to predict the coefficients of the basis functions. The NN demonstrates significantly better performance compared to the least squares method with median R 2 = 0.9997 and 0.9916 for J t and ψ, respectively. The robustness of the method is evaluated by handling missing or incorrect data through the least squares filling of missing data, which shows that the NN prediction remains strong even with a reduced number of diagnostics. Additionally, the method is tested on plasmas outside of the training range showing reasonable results.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Machine learning inversion from scattering for mechanically driven polymers

A machine learning inversion method is developed for analyzing scattering functions of mechanically driven polymers and extracting the corresponding feature parameters, which include energy parameters and conformation variables. The polymer is modeled as a chain of fixed-length bonds constrained by bending energy, and it is subject to external forces such as stretching and shear. We generate a data set consisting of random combinations of energy parameters, including bending modulus, stretching and shear force, along with Monte Carlo-calculated scattering functions and conformation variables such as end-to-end distance, radius of gyration and off-diagonal component of the gyration tensor. The effects of the energy parameters on the polymer are captured by the scattering function, and principal component analysis ensures the feasibility of the machine learning inversion. Finally, we train a Gaussian process regressor using part of the data set as a training set and validate the trained regressor for inversion using the rest of the data. The regressor successfully extracts the feature parameters.

Gaussian process regressors↗

Uncertainty Quantification for Smooth Functional Data with Application to Material Properties

This document outlines a method for processing functional output (i.e., curves) for the ultimate purpose of sampling curves under specified input conditions for use in modeling and simulation uncertainty quantification (UQ) studies. A set of benchmark curves sufficiently representative of the relevant scenario(s) being simulated are provided to the process and formatted as described in Section 1. Principal Component Analysis (PCA) is utilized to discover the components of uncertainty in the benchmark curves and is outlined in Section 2. Section 3 describes the application of uncertainty quantification to the PCA results for the purpose of sampling curves to be used in UQ analysis. Section 4 applies these techniques to an example benchmark dataset. Concluding remarks are provided in the final section.

36 MATERIALS SCIENCE↗

SpecDis: Value Added Distance Catalog for 4 Million Stars from DESI Year-1 Data

We present the SpecDis value-added stellar distance catalog accompanying DESI Data Release 1. SpecDis trains a feed-forward neural network (NN) with Gaia parallaxes and gets the distance estimates. To build up an unbiased training sample, we do not apply selections on parallax error or signal-to-noise (S/N) of the stellar spectra, and instead, we incorporate parallax error into the loss function. Moreover, we employ principal component analysis to reduce the noise and dimensionality of stellar spectra. Validated by independent external samples of member stars with precise distances from globular clusters, dwarf galaxies, stellar streams, combined with blue horizontal branch stars, we demonstrate that our distance measurements show no significant bias up to 100 kpc, and are much more precise than Gaia parallax beyond 7 kpc. The median distance uncertainties are 23%, 19%, 11%, and 7% for S/N < 20, 20 ≤ S/N < 60, 60 ≤ S/N < 100, and S/N ≥ 100. Selecting stars with ${\mathrm{log}}\,g\lt 3.8$ and distance uncertainties smaller than 25%, we have more than 74,000 giant candidates within 50 kpc of the Galactic center and 1500 candidates beyond this distance. Additionally, we develop a Gaussian mixture model to identify unresolvable equal-mass binaries by modeling the discrepancy between the NN-predicted and the geometric absolute magnitudes from Gaia parallaxes and identify 120,000 equal-mass binary candidates. Our final catalog provides distances and distance uncertainties for >4 million stars, offering a valuable resource for Galactic astronomy.

astronomy data analysis↗

PCAfold 2.0—Novel tools and algorithms for low-dimensional manifold assessment and optimization

We describe an update to our open-source Python package, PCAfold, designed to help researchers generate, analyze and improve low-dimensional data manifolds. In the current version, PCAfold 2.0, we introduce novel tools and algorithms for assessing and optimizing low-dimensional manifolds. This includes a method that generates a “map” of local feature sizes that can help pinpoint researchers to problematic regions on a manifold. We introduce a novel cost function that characterizes the quality of a manifold topology with a single number. We develop two algorithms for feature selection based on principal component analysis (PCA) that use the cost function as an objective function to minimize. We introduce a quantity of interest (QoI)-aware dimensionality reduction strategy where data projections are computed using an artificial neural network and are directly optimized towards representing various projection-independent and projection-dependent QoIs. We also introduce an implementation of partition of unity networks (POUnets) for efficient reconstruction of QoIs from low-dimensional manifolds based on combining neural network classification with localized polynomial regression. Our software can be broadly applicable in all domains of science and engineering that aim to reduce data dimensionality, as well as in the fundamental research on representation learning.

97 MATHEMATICS AND COMPUTING↗

The influence of physical and algorithmic factors on simulated far-field waveforms and source–time functions of underground explosions using unsupervised machine learning

SUMMARY Characterizing explosion sources and differentiating between earthquake and underground explosions using distributed seismic networks becomes non-trivial when explosions are detonated in cavities or heterogeneous ground material. Moreover, there is little understanding of how changes in subsurface physical properties affect the far-field waveforms we record and use to infer information about the source. Simulations of underground explosions and the resultant ground motions can be a powerful tool to systematically explore how different subsurface properties affect far-field waveform features, but there are added variables that arise from how we choose to model the explosions that can confound interpretation. To assess how both subsurface properties and algorithmic choices affect the seismic wavefield and the estimated source functions, we ran a series of 2-D axisymmetric non-linear numerical explosion experiments and wave propagation simulations that explore a wide array of parameters. We then inverted the synthetic far-field waveform data using a linear inversion scheme to estimate source–time functions (STFs) for each simulation case. We applied principal component analysis (PCA), an unsupervised machine learning method, to both the far-field waveforms and STFs to identify the most important factors that control variance in the waveform data and differences between cases. For the far-field waveforms, the largest variance occurs in the shallower radial receiver channels in the 0–50 Hz frequency band. For the STFs, both peak amplitude and rise times across different frequencies contribute to the variance. We find that the ground equation of state (i.e. lithology and rheology) and the explosion emplacement conditions (i.e. tamped versus cavity) have the greatest effect on the variance of the far-field waveforms and STFs, with the ground yield strength and fracture pressure being secondary factors. Differences in the PCA results between the far-field waveforms and STFs could possibly be due to near-field non-linearities of the source that are not accounted for in the estimation of STFs and could be associated with yield strength, fracture pressure, cavity radius and cavity shape parameters. Other algorithmic parameters are found to be less important and cause less variance in both the far-field waveforms and STFs, meaning algorithmic choices in how we model explosions are less important, which is encouraging for the further use of explosion simulations to study how physical Earth properties affect seismic waveform features and estimated STFs.

58 GEOSCIENCES↗

Design Basis Document / Owner’s Technical Specification for Nitrate Salt Systems in CSP Projects: Volume 1 Specifications for Parabolic Trough Projects, Volume 2 - Specifications for Central Receiver Projects, Volume 3 - Narrative

Design Basis Documents / Owner’s Technical Specifications are developed for parabolic trough and central receiver power plants using nitrate salt as the heat transport fluid and the thermal storage medium. The goals were to 1) distill the successful experience with nitrate salt systems from as many commercial projects as possible, 2) provide technical bases for equipment design/selection that an owner can impose on an EPC contractor, 3) compile this information in one location to provide guidance on salt systems that is as broadly applicable to as many projects as possible, and 4) work toward an industry consensus on a Design Basis Document that reflects the lessons learned from earlier commercial projects. The document allows an owner to provide design guidance, and to impose a minimum set of requirements, above and beyond those in the normal Codes and Standards, on an EPC contractor in an effort to avoid a repeat of past mistakes. If successful, this would allow salt systems to achieve the same levels of reliability and availability as commercial parabolic trough plants using diphenyl oxide / biphenyl (i.e., Dowtherm A, Therminol VP-1) as the heat transfer fluid. The report consists of 3 volumes: • Volume 1 - Specifications for Parabolic Trough Projects • Volume 2 - Specifications for Central Receiver Projects • Volume 3 - Narrative Volumes 1 and 2 include discussions of the following topics: • Plant functional requirements for the salt systems • Plant operating states, and transitions between states, for the salt systems • Risk analysis of the principal salt components • Plant requirements for the salt systems to meet the functional and the operating requirements • Type of specification for the principal salt components: functional; or prescriptive • Current state of the art for salt systems. Volume 3 discusses a number of cases in which the salt equipment at commercial parabolic trough and central receiver projects has not met the projected levels of reliability and availability. Possible reasons for the sources of the problems are discussed, as is a range of possible alternate designs that could avoid the known problems.

14 SOLAR ENERGY↗

Principal Component Analysis of azimuthal flow in intermediate-energy heavy-ion reactions

Principal Component Analysis (PCA) via Singular Value Decomposition (SVD) of large datasets is an adaptive exploratory method to uncover natural patterns underlying the data. Several recent applications of the PCA-SVD to event-by-event single-particle azimuthal angle distribution matrices in ultra-relativistic heavy-ion collisions at RHIC-LHC energies indicate that the sine and cosine functions chosen a priori in the traditional Fourier analysis are naturally the most optimal basis for azimuthal flow studies according to the data itself. We perform PCA-SVD analyses of mid-central Au+Au collisions at $E$ $beam$ / $A$ =1.23 GeV simulated using an isospin-dependent Boltzmann-Uehling-Uhlenbeck (IBUU) transport model to address the following two questions: (1) if the principal components of the covariance matrix of nucleon azimuthal angle distributions in heavy-ion reactions around 1 GeV/nucleon are naturally sine and/or cosine functions and (2) what if any advantages the PCA-SVD may have over the traditional flow analysis using the Fourier expansion for studying the EOS of dense nuclear matter. In conclusion, we find that (1) in none of our analyses the principal components come out naturally as sine and/or cosine functions, (2) while both the eigenvectors and eigenvalues of the covariance matrix are appreciably EOS dependent, the PCA-SVD has no apparent advantage over the traditional Fourier analysis for studying the EOS of dense nuclear matter using the azimuthal collective flow in heavy-ion collisions.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Image Distinguishability Analysis Testing Through Principal Components and Its Application to Hot Spot Scale Invariance

Hot spots are spatial regions of intense energy localization that govern initiation of secondary high explosives. Studies that characterize or compare simulated hot spots are frequently either qualitatively descriptive or resort to quantitative distribution functions that neglect stochastic variations and spatial correlations—effects that are also neglected in common comparison tests like the Kolmogorov–Smirnov test. To this end, we develop an image distinguishability analysis (IDA) test based on principal component (PC) analysis that makes pixel-by-pixel comparisons between small, for example, O(<10), image data sets. The IDA test makes comparisons through a generalized distance metric in the PC space and a test statistic that is derived to calculate mathematical equation-values. Here, we derive a statistical distribution and criticality criterion to determine whether images are distinguishable from established baselines. We apply the IDA test on images generated from molecular dynamics simulations of hot spots from pore collapse in TATB to assess scale invariance in the complex patterns of hot spots that form in a representative high explosive crystal. The IDA test shows that TATB hot spot spatial temperature fields and their derived temperature histograms exhibit scale-invariant features over specific intervals of shock orientation, strength, and initial pore diameter. However, the IDA test also shows that qualitatively different conclusions regarding invariance can be reached depending on whether the hot spot is treated as a spatially correlated field as opposed to a distribution function that lacks spatial information.

organic↗

Using the optimal combined index weight ratio to improve the probability of anomaly detection in big area additive manufacturing

Big Area Additive Manufacturing (BAAM) of composites requires significant time, energy, and material, so it is critical to reduce production inefficiencies to make functional parts without multiple iterations. Statistical process control coupled with Principal Component Analysis (PCA) is a powerful technique that provides a quick, computationally inexpensive, and intuitive way for operators to detect defects that form in a manufacturing process without massive datasets. Recently, a combined index that is a weighted sum of the Hotelling's T 2 and squared residual error statistics has been proposed that can be monitored in one chart, improving interpretation accuracy and simplicity. However, the literature does not offer a formal method to optimise the weights. Here, we introduce two new approaches to the traditional weight selection approach using simulated and BAAM image data. Approach 1 uses a theoretically motivated optimum inspired by probabilistic principal component analysis. Approach 2 systematically varies the ratio of the weights to find the optimum. We show that approach 1 delivers optimal anomaly detection performance in select cases while approach 2 fares better in practice. Surprisingly, we also show that choosing a more complex PCA model has a minimal negative impact on anomaly detection performance compared to a more simplistic model.

3-dimensional printing↗

Beyond PCA: Additional Dimension Reduction Techniques to Consider in the Development of Climate Fingerprints

Abstract Dimension reduction techniques are an essential part of the climate analyst’s toolkit. Due to the enormous scale of climate data, dimension reduction methods are used to identify major patterns of variability within climate dynamics, to create compelling and informative visualizations, and to quantify major named modes such as El Niño–Southern Oscillation. Principal components analysis (PCA), also known as the method of empirical orthogonal functions (EOFs), is the most commonly used form of dimension reduction, characterized by a remarkable confluence of attractive mathematical, statistical, and computational properties. Despite its ubiquity, PCA suffers from several difficulties relevant to climate science: high computational burden with large datasets, decreased statistical accuracy in high dimensions, and difficulties comparing across multiple datasets. In this paper, we introduce several variants of PCA that are likely to be of use in climate sciences and address these problems. Specifically, we introduce non-negative , sparse , and tensor PCA and demonstrate how each approach provides superior pattern recognition in climate data. We also discuss approaches to comparing PCA-family results within and across datasets in a domain-relevant manner. We demonstrate these approaches through an analysis of several runs of the E3SM climate model from 1991 to 1995, focusing on the simulated response to the Mt. Pinatubo eruption; our findings are consistent with a recently identified stratospheric warming fingerprint associated with this type of stratospheric aerosol injection.

Weylandt, Michael↗

Assisted migration in a warmer and drier climate: less climate buffering capacity, less facilitation and more fires at temperate latitudes?

Assisted tree migration has been proposed as a conceptual solution to mitigate lags in biotic responses to anthropogenic climate change. The rationale behind this concept is that tree species currently growing under warmer and drier climates will be more resistant and resilient to the new climatic conditions than tree species naturally growing in currently wetter and colder climates. However, we hypothesize that, by being more stress‐tolerant to warmer and drier conditions, translocated species should exhibit different functional attributes, which could induce important ecological and societal costs and overcome the desired benefits of maintaining wood production and other ecosystem services. We used principal component analysis (PCA) to analyze variation in seven traits of 106 tree and tall shrub species from contrasting latitudinal distributions in western North America and Europe to predict the potential functional changes of forest ecosystems due to the translocation of tree species from low to high latitudes. We show that species from both continents differed primarily by their position on the leaf economy spectrum (LES) and their size traits. Even though, in Europe, differences in LES were significantly correlated to species southern latitudinal positions, in both continents differences in size traits were significantly correlated to latitude. These results suggest that assisted migration by translocating more conservative species of shorter stature in currently cooler climates should decrease the buffering capacity of forest canopies, decrease facilitation for understory species, and increase wildfire risks, whose effects have the potential to accelerate climate warming through negative atmospheric feedback processes. As an alternative solution to assisted migration that may accelerate rather than mitigate climate change, we recommend that foresters gradually diversify the vertical structure and layering of the existing forest canopy to maintain a sustainable water cycle and energy balance between the soil, the tree and the atmosphere without increasing the wildfire risk.

Environmental Sciences & Ecology↗

Global teleconnections influencing large-scale drought in the United States using SVDI

Understanding recent large-scale drought patterns and the mechanisms producing extreme drought events is vital for future drought forecasts and understanding future drought risks. Increasingly, vapor pressure deficit (VPD) has been used as an important measure of evaporative demand and proxy for drought detection. In this study, VPD is used to calculate the new Standardized VPD Drought Index (SVDI) with NASA North American Land Data Assimilation System (NLDAS) data. Previous studies have shown that SVDI accurately identifies the timing and magnitude short-term droughts in the United States (U.S). In the present study, SVDI is now used to identify large-scale drought patterns between 1980 and 2021 and drought variability driven by selected global teleconnections originating in the Pacific and Atlantic Oceans. Spatial drought characteristics were extracted from SVDI using empirical orthogonal function (EOF) analysis. Then a k-means clustering algorithm was applied to both EOF principal components and primary teleconnections, including the El Nino-Southern Oscillation (ENSO) and Pacific Decadal Oscillation (PDO) to identify drought events driven by the Pacific Ocean. Results show that the SVDI is useful in evaluating large-scale drought variability in the U.S. related to global teleconnections, and that mechanisms influencing summer drought patterns in the Western and Southwestern U.S. are driven by a tropical-extratropical interactions originating in the equatorial Pacific Ocean related to ENSO dynamics with interdecadal variability modulated by PDO. The large-scale droughts in the Central and Southern U.S., like those in 2011 and 2012, on the other hand, are driven by the North Pacific Ocean warm pool during a strong negative PDO, which subsequently influenced variability in the Bermuda-Azores High in the Atlantic Ocean. In summer 2011, the Bermuda-Azores High weakened, reducing the onshore winds and moisture transport along the eastern Gulf of Mexico and contributing to ongoing drought in the region. The Northern Pacific and Atlantic Ocean sea surface temperatures (SSTs) have increased between 1980 and 2021. In conclusion, as SSTs continue to rise in the Northern Pacific Ocean, one consequence of the coupled North Pacific warm pool and atmospheric dynamics, is to increase summer drought variability over a large region in the southern and midwestern U.S. under global warming.

54 ENVIRONMENTAL SCIENCES↗