Search NASASearch

SEARCH · Search NASA

Results for “principal component analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Characterizing skyrmion flow phases with principal component analysis

Principal component analysis (PCA) is a powerful method that can identify patterns in large, complex data sets by constructing low-dimensional order parameters from higher-dimensional feature vectors. There are increasing efforts to use space-and-time-dependent PCA to detect transitions in nonequilibrium systems that are difficult to characterize with equilibrium methods. Here, we demonstrate that feature vectors incorporating the position and velocity information of driven skyrmions moving through random disorder permit PCA to resolve different types of disordered skyrmion motion as a function of driving force and the ratio of the Magnus force to the dissipation. Since the Magnus force creates gyroscopic motion and a finite Hall angle, skyrmions can exhibit a greater range of flow phases than what is observed in overdamped driven systems with quenched disorder. We show that in addition to identifying previously known skyrmion flow phases, PCA detects several additional phases, including different types of channel flow, moving fluids, and partially ordered states. Guided by the PCA analysis, we further characterize the disordered flow phases to elucidate the different microscopic dynamics and show that the changes in the PCA-derived order parameters can be connected to features in bulk transport measures, including the transverse and longitudinal velocity-force curves, differential conductivity, topological defect density, and changes in the skyrmion Hall angle as a function of drive. We discuss how asymmetric feature vectors can be used to improve the resolution of the PCA analysis, and how this technique can be extended to find disordered phases in other nonequilibrium systems with time-dependent dynamics.

36 MATERIALS SCIENCE

Modeling MTS pyrolysis and SiC deposition kinetics using principal component analysis and neural networks

Accurate chemical kinetics modeling is crucial for improving the efficiency of chemical processing and synthesis of ceramic matrix composites. Detailed kinetic models are computationally expensive due to the large number of transported chemical species, while the simplified physics-based models, such as single-step global mechanisms, are efficient but often overlook key chemical intermediates and pathways. Recent deep learning approaches promise accurate and cost-effective models. Yet, they require additional closures for the transported nonlinear latent variables, complicating integration with existing solvers. In this work, we develop a hybrid linear—nonlinear reduced model for silicon carbide deposition from methyltrichlorosilane precursor by combining principal component analysis (PCA) and autoencoder (AE) neural network (NN) approaches. PCA is used to identify a smaller set of linear transport variables, enabling direct reuse of conventional transport solvers. NNs then reconstruct the full chemical state from these reduced variables. We demonstrate the method on a chemical vapor deposition reactor—comprising a gas-phase pyrolysis plug flow reactor and a heterogeneous surface reactor—over a wide range of temperatures, pressures, and residence times. Our PCA–AE model achieves high accuracy with only five transported scalars, achieving an eightfold cost reduction compared to detailed mechanisms, in both a priori (using data from the test set only) and a posteriori (coupled with a differential equation solver). In conclusion, notable errors arise primarily near training domain boundaries and for long residence times, indicating the need for domain shift indicators and better long-horizon predictions in future reduced chemistry model development.

autoencoder neural networks

Using principal component analysis to distinguish different dynamic phases in superconducting vortex matter

Vortices in type-II superconductors driven over random disorder are known to exhibit a remarkable variety of distinct nonequilibrium dynamical phases that arise owing to the competition between vortex-vortex interactions, the quenched disorder, and the drive. These include pinned states, elastic flows, plastic or disordered flows, and dynamically reordered moving crystal or moving smectic states. The plastic flow phases can be particularly difficult to characterize since the flows are strongly disordered. Here, we perform principal component analysis (PCA) on the positions and velocities of vortex matter moving over random disorder for different disorder strengths and drives. We find that PCA can distinguish the known dynamic phases as well as or better than previous measures based on transport signatures or topological defect densities. In addition, PCA recognizes distinct plastic flow regimes, a slowly changing channel flow and a moving amorphous fluid flow, that do not produce distinct signatures in the standard measurements. In conclusion, our results suggest that this position and velocity-based PCA approach could be used to characterize dynamic phases in a broader class of systems that exhibit depinning and nonequilibrium phase transitions.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND

Identifying Nuclear Data Correlated Through Predicting Bias in Integral Experiments via Applying Principal Component Analysis to Random Forest

ABSTRACT Nuclear data (ND) are the input data for neutron‐transport simulations to answer questions related to nuclear technologies. Subsets of ND, here > 20,000 data points, are validated with respect to thousands of criticality experiments that represent various applications on a small scale. The aim of validation with these experiments is to find errors in ND or methods. The key challenge here is that several hundreds of ND are used to simulate one integral value. Hence, one cannot clearly identify what ND are leading to bias in criticality measurements. In fact, a mistake in one nuclear‐data observable can be compensated with an error in another, and the predicted criticality value would still be predicted in agreement with experimental data. Random forest (RF) was previously employed to predict bias in criticality measurements using sensitivities of simulated criticality experiments to ND. The SHapley Additive exPlanations (SHAP) metric was then applied to attribute the importance of each ND experiment and observable to bias prediction. This, however, did not highlight what ND were jointly related to predicting bias. This is important as it could inform us about where compensating errors in ND could hide. We tackle this shortcoming here by first decomposing the ND sensitivities to integral‐experiment simulations into principal components. Then we use principal component projections to predict bias via the RF and SHAP. The SHAP values and principal components are employed to reconstruct detailed SHAP values for each ND observable. We demonstrate that these extended SHAP bias predictions are more robust, less noisy, and more efficient. In addition, we show that this approach accounts for covariance in ND sensitivities and automates the identification of where compensating errors could hide in ND.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

An evaluation of air quality in major urban areas of India

Rapid economic growth and burgeoning population have contributed to enhanced levels of PM 2.5 concentrations in urban regions of India. Evaluation of ambient air quality facilitates the assessment of effectiveness of emission control measures and early identification of new sources. This study provides a comprehensive statistical analysis of PM 2.5 concentrations in key urban areas across India, including Delhi, Kolkata, Mumbai, Chennai, Hyderabad, and several regional centers. Data from 2017 to 2023 was analyzed using trend analysis, cluster analysis, principal component analysis, and geostatistical interpolation to understand spatiotemporal variations and sources. The analysis reveals significant differences in spatial distribution of PM 2.5 concentrations with high annual averages in urban regions in Indo-Gangetic plain (82–123 μg m −3 ) and relatively lower concentrations (29–46 μg m −3 ) in southern urban areas of Kerala, Tamil Nadu and Andhra Pradesh. Delhi state had the highest 24-averaged PM 2.5 concentrations (112 μg m −3 ) followed by urban regions in Uttar Pradesh, Bihar and West Bengal (94 μg m −3 ). Trend analysis from 2017 to 2023 revealed an overall 2.5% decline in site-wide PM2.5 concentrations, with the exception of Ludhiana, which exhibited a consistent annual increase of 10%. Principal component analysis (PCA) attributes 30% of the variance to wintertime emissions, 13% to biomass burning, and 18% to the regional haze in the northern Indo-Gangetic Plain. Different analyses clearly demonstrates the contribution of biomass burning to pollution in Delhi and surrounding cities. Transboundary pollution to Kolkata is likely from the highly polluted region in Indo-Gangetic Plain. Coastal cities of Mumbai and Chennai has relatively lower pollution attributed to the influence of sea breeze dilution, with mostly local contribution and some potential transport from upwind industry clusters. Hyderabad also has local contribution due to high density of vehicular traffic and local small industries. This study shows that mitigation efforts targeting clusters of regions should be undertaken to curb the high PM2.5 pollution. Policy measures should be implemented both at local and the intra-state level to address shared sources and transport of pollution.

Hysplitbacktrajectories

Uncertainty Quantification for Smooth Functional Data with Application to Material Properties

This document outlines a method for processing functional output (i.e., curves) for the ultimate purpose of sampling curves under specified input conditions for use in modeling and simulation uncertainty quantification (UQ) studies. A set of benchmark curves sufficiently representative of the relevant scenario(s) being simulated are provided to the process and formatted as described in Section 1. Principal Component Analysis (PCA) is utilized to discover the components of uncertainty in the benchmark curves and is outlined in Section 2. Section 3 describes the application of uncertainty quantification to the PCA results for the purpose of sampling curves to be used in UQ analysis. Section 4 applies these techniques to an example benchmark dataset. Concluding remarks are provided in the final section.

36 MATERIALS SCIENCE

Combining ToF‐SIMS and Multivariate Analysis to Resolve Active Sites on Ni‐Based HER Catalysts

Unambiguous identification of active sites in heterogeneous catalysis remains a major challenge, particularly for materials with ultrathin, chemically mixed surface layers. Here, we demonstrate a generalizable approach that combines time-of-flight secondary ion mass spectrometry (ToF-SIMS) with multivariate statistical analysis (principal component analysis [PCA] and multivariate curve resolution [MCR]) to resolve catalytically relevant motifs at the nanoscale. Using Ni electrodes as a model system, PCA distinguished hydroxide-enriched domains from oxide- and metal-rich regions, while MCR decomposed depth profiles and 3D images into hydroxide, oxide, and metallic layers with nanometer resolution. A unique secondary-ion fragment, NiO 3 H 3 − (m/z 108.94), emerged as a marker of hydroxide-rich environments and correlated with hydrogen evolution reaction (HER) activity across a series of Ni electrodes. Complementary density functional theory (DFT) calculations revealed that Ni(OH) 2 clusters adjacent to metallic Ni offer the most favorable water dissociation energetics, establishing the structural origin of the marker. While illustrated here for Ni-based HER, this workflow provides a broadly applicable framework to isolate and rank near-surface patterns that govern catalytic activity, thereby extending ToF-SIMS from a qualitative probe to a predictive tool for active site identification.

HER active sites

Using the optimal combined index weight ratio to improve the probability of anomaly detection in big area additive manufacturing

Big Area Additive Manufacturing (BAAM) of composites requires significant time, energy, and material, so it is critical to reduce production inefficiencies to make functional parts without multiple iterations. Statistical process control coupled with Principal Component Analysis (PCA) is a powerful technique that provides a quick, computationally inexpensive, and intuitive way for operators to detect defects that form in a manufacturing process without massive datasets. Recently, a combined index that is a weighted sum of the Hotelling's T 2 and squared residual error statistics has been proposed that can be monitored in one chart, improving interpretation accuracy and simplicity. However, the literature does not offer a formal method to optimise the weights. Here, we introduce two new approaches to the traditional weight selection approach using simulated and BAAM image data. Approach 1 uses a theoretically motivated optimum inspired by probabilistic principal component analysis. Approach 2 systematically varies the ratio of the weights to find the optimum. We show that approach 1 delivers optimal anomaly detection performance in select cases while approach 2 fares better in practice. Surprisingly, we also show that choosing a more complex PCA model has a minimal negative impact on anomaly detection performance compared to a more simplistic model.

3-dimensional printing

DESI Emission-line Galaxies: Unveiling the Diversity of [O II ] Profiles and Its Links to Star Formation and Morphology

We study the [O II ] profiles of emission-line galaxies (ELGs) from the Early Data Release of the Dark Energy Spectroscopic Instrument (DESI). To this end, we decompose and classify the shape of [O II ] profiles with the first two eigenspectra derived from principal component analysis. Our results show that DESI ELGs have diverse line profiles, which can be categorized into three main types: (1) narrow lines with a median width of ∼50 km s −1 , (2) broad lines with a median width of ∼80 km s −1 , and (3) two redshift systems with a median velocity separation of ∼150 km s −1 , i.e., double-peak galaxies. To investigate the connections between the line profiles and galaxy properties, we utilize the information from the COSMOS data set and compare the properties of ELGs, including star formation rate (SFR) and galaxy morphology, with the average properties of reference star-forming galaxies with similar stellar mass, sizes, and redshifts. Our findings show that, on average, DESI ELGs have a higher SFR and more asymmetrical/disturbed morphology than the reference galaxies. Moreover, we uncover a relationship between the line profiles, the excess SFR, and the excess asymmetry parameter, showing that DESI ELGs with broader [O II ] line profiles have more disturbed morphology and higher SFR than the reference star-forming galaxies. Finally, we discuss possible physical mechanisms giving rise to the observed relationship and the implications of our findings on the galaxy clustering measurements, including the halo occupation distribution modeling of DESI ELGs and the observed excess velocity dispersion of the satellite ELGs.

79 ASTRONOMY AND ASTROPHYSICS

Mapping Rare Earths and Toxics in E-Waste via Hyperspectral Imaging and Machine Learning

Electronic waste (e-waste) presents a mounting challenge to environmental sustainability due to its complex composition, which includes high-value rare earth elements, hazardous organic compounds, and non-recyclable plastics. Accurate and scalable material classification is essential for enabling efficient resource recovery and safe recycling practices. This study introduces a confidence-aware classification pipeline that combines mid-infrared hyperspectral imaging (HSI), spectral angle mapping (SAM), and iterative machine learning to perform pixel-level material identification across e-waste devices. A curated spectral library encompassing artificial materials (e.g., plastic iron oxide, galvanized metals), minerals (e.g., allanite, hematite), and organic compounds (e.g., benzanthracene, toluene) was used to generate pseudo-labels, each assigned a confidence score based on SAM-derived spectral similarity. High-confidence samples from seven consumer electronics—digital cameras, keyboards, laptop fans, modems, motherboards, TV remotes, and speakers—were iteratively expanded and classified using models such as Support Vector Machine (SVM), Random Forest, Gradient Boosting Classifier, Partial Least Squares Discriminant Analysis (PLSDA) and Logistic Regression. The best-performing classifiers achieved macro F1 scores approaching 1.0. Results revealed widespread plastic content (dominated by plastic iron oxide), the presence of rare earth-bearing minerals like cerium-containing allanite, and pervasive detection of hazardous organics such as benzanthracene. Principal Component Analysis (PCA) visualizations and confusion matrices confirmed high separability and robust classification performance. This methodology enables precise, non-destructive, and scalable classification of heterogeneous e-waste streams. It supports automated, hazard-aware sorting in recycling workflows, facilitating selective recovery of critical materials and compliance with circular economy goals. The confidence-aware framework provides a foundation for real-time deployment in industrial settings, offering significant implications for smart e-recycling infrastructure and policy-driven material stewardship.

Circular economy

Elemental and isotopic signatures of Asteroid Ryugu support three early Solar System reservoirs

Understanding the number and locations of different reservoirs present in the early Solar System is crucial to understanding the Solar System’s origin and evolution. Previous work has suggested that three unique isotopic reservoirs existed in the early Solar System but subsequent works have challenged that idea. Here we present elemental abundances along with Ca, Ti, Cr, Fe, Ni, and Zn isotopic data from primitive material returned by the Japan Aerospace Exploration Agency’s (JAXA) Hayabusa2 mission to asteroid (162173) Ryugu to make inferences on the Solar System’s early architecture. Data from Ryugu particle A0208 are consistent with a close genetic heritage between Ryugu and CI chondrites. Here, we employ principal component analysis (PCA) on these Ryugu and published meteorite data to demonstrate that Ryugu and CI chondrites are distinct from other known astromaterials, strongly supporting the existence of a third major isotopic reservoir in the early Solar System.

Isotopes

Accelerating the Structure Exploration of Diverse Bi–Pt Nanoclusters via Physics‐Informed Machine Learning Potential and Particle Swarm Optimization

Bimetallic Bi–Pt nanoclusters exhibit diverse structural motifs, including core-shell, Janus, and mixed alloy configurations, due to the unique bonding characteristics between Bi and Pt atoms. Using density functional theory refinements from ChIMES physically machine-learned potential and CALYPSO particle swarm optimization global searches, 34 Bi20-Pt20 nanoclusters are systematically classified. The results reveal that Bi atoms predominantly occupy surface sites, driven by charge transfer effects. Cohesive energy trends alone prove insufficient for structure differentiation, necessitating a data-driven approach employing principal component analysis and K-means clustering. Furthermore, vibrational, electronic, and infrared spectral analyses provide additional insights into structure-property relationships. The findings offer an original framework for the automated classification and analysis of bimetallic nanoclusters, enhancing the understanding of their stability and functional properties.

bimetallic nanoparticles

Advancements in Constitutive Model Calibration: Leveraging the Power of Full‐Field DIC Measurements and In Situ Load Path Selection for Reliable Parameter Inference

Accurate material characterization and model calibration are essential for computationally supported high-consequence engineering decisions. Historically, characterization and calibration methods (1) use simplified test specimen geometries and global data, (2) cannot guarantee that sufficient characterization data are collected for a specific model of interest, (3) use deterministic methods that provide best-fit parameter values with no uncertainty quantification, and (4) are sequential, inflexible, and time-consuming. This work brings together several recent advancements into an improved workflow called interlaced characterization and calibration (ICC) that advances the state-of-the-art in constitutive model calibration. The ICC paradigm (1) employs tools to efficiently use full-field data to calibrate high-fidelity material models, (2) aligns the data needed with the data collected by adopting an optimal experimental design protocol, (3) quantifies parameter uncertainty through Bayesian inference and (4) incorporates these advancements into a quasi real-time feedback loop. The ICC framework is demonstrated here on the calibration of a material model using simulated full-field data for an aluminium cruciform specimen being deformed biaxially. The cruciform is actively driven through the myopically preferred load path using Bayesian optimal experimental design, which selects load steps that yield the maximum expected information gain (EIG). Principal component analysis (PCA) is performed on the model predictions of full-field displacements, and fast surrogate models are built to approximate the input-output relationships of the expensive finite element model. Furthermore, the tools developed and demonstrated here show that high-fidelity constitutive models can be efficiently and reliably calibrated with quantified uncertainty, thus supporting credible decision-making and potentially increasing the agility of solid mechanics modelling by enabling utilization of computational simulations at earlier stages of the design cycle.

Bayesian optimal experimental design

Uncovering Sequence and Structural Characteristics of Fungal Expansin‐Related Proteins With Potential to Drive Substrate Targeting

Expansins loosen plant cell wall networks through disrupting non-covalent bonds between cellulose microfibrils and matrix polysaccharides. Whereas expansins were first discovered in plants, expansin-related proteins have since been identified in bacteria and fungi. The biological function of microbial expansins remains unclear; however, several studies have shown distinct binding preferences toward different structural polysaccharides. Earlier studies of bacterial expansin-related proteins uncovered sequence and structural features that correlate to substrate binding. Herein, 20 fungal expansin-related sequences were recombinantly produced in Komagataella phaffii, and the purified proteins were compared in terms of substrate binding to cellulosic and chitinous substrates. The impact of pH on the zeta potential of prioritized substrates was also measured, and Principal Component Analysis was performed to uncover correlations between protein characteristics (e.g., pI, hydrophobicity, surface charge distribution) and measured substrate binding preferences. Whereas acidic proteins with a predicted pI less than 5.0 preferentially bound to chitin, basic proteins with pI greater than 8.0 preferentially bound to xylan and xylan-containing fiber. Similar to many cellulases, binding to cellulose was correlated to relatively high aromatic amino acid content in the protein sequence and presence of a carbohydrate binding module (CBM), which in the case of expansins is a C-terminal CBM63. Whereas overall sequence characteristics could be correlated to substrate binding preference, the identity of amino acids occupying conserved positions that impact protein activity was better correlated with loosenin versus expansin classifications.

chitin

Multivariate Analysis as a Tool for Validating Tester Matching

A method of applying Principal Component Analysis, Soft Independent Modeling of Class Analysis, and statistical analysis is described that can be applied to many types of testers to ascertain how well matched the performance of the testers in the analysis are to one another or how well matched a tester is to itself at a later time. This method is most useful for situations for which the same units have not been run across the testers being analyzed for matched performance.

Multari, Rosalie A [Sandia National Laboratories (

Explainable Machine Learning for Functional Data

Black-box machine learning models are recognized as useful tools for prediction applications, but the algorithmic complexity of some models causes interpretation challenges. Explainability methods have been proposed to provide insight into these models, but there is little research focused on supervised modeling with functional data inputs. We argue that, especially in applications of high consequence, it is important to explicitly model the functional dependence in a black-box analysis to not obscure or misrepresent patterns in explanations. As such, we propose the V ariable importance E xplainable E lastic S hape A nalysis (VEESA) pipeline for training supervised machine learning models with functional inputs. The pipeline is an analysis process that includes the data preprocessing, modeling, and post-hoc explanations. The preprocessing is done using elastic functional principal components analysis, which accounts for vertical and horizontal variability in functional data and, ultimately, allows for explanations in the original data space that identify the important functional variability without bias due to correlated variables. Here, we demonstrate the pipeline on two high-consequence applications: explosives classification for national security and inkjet printer identification in forensic science. The applications exhibit the VEESA pipeline’s ability to provide an understanding of the characteristics of the functional data useful for prediction. Code for implementing the pipeline is available in the veesa R package (and supplemental python code).

Elastic Shape Analysis

Image Distinguishability Analysis Testing Through Principal Components and Its Application to Hot Spot Scale Invariance

Hot spots are spatial regions of intense energy localization that govern initiation of secondary high explosives. Studies that characterize or compare simulated hot spots are frequently either qualitatively descriptive or resort to quantitative distribution functions that neglect stochastic variations and spatial correlations—effects that are also neglected in common comparison tests like the Kolmogorov–Smirnov test. To this end, we develop an image distinguishability analysis (IDA) test based on principal component (PC) analysis that makes pixel-by-pixel comparisons between small, for example, O(<10), image data sets. The IDA test makes comparisons through a generalized distance metric in the PC space and a test statistic that is derived to calculate mathematical equation-values. Here, we derive a statistical distribution and criticality criterion to determine whether images are distinguishable from established baselines. We apply the IDA test on images generated from molecular dynamics simulations of hot spots from pore collapse in TATB to assess scale invariance in the complex patterns of hot spots that form in a representative high explosive crystal. The IDA test shows that TATB hot spot spatial temperature fields and their derived temperature histograms exhibit scale-invariant features over specific intervals of shock orientation, strength, and initial pore diameter. However, the IDA test also shows that qualitatively different conclusions regarding invariance can be reached depending on whether the hot spot is treated as a spatially correlated field as opposed to a distribution function that lacks spatial information.

organic

Machine Learning for Mapping Multipactor Susceptibility in RF Systems: Capabilities and Generalization Constraints

Multipactor is a surface-driven electron avalanche phenomenon that degrades the performance and reliability of radio-frequency (RF) systems in particle accelerator and vacuum electronics applications. Multipactor behavior in a given device structure is conventionally assessed through susceptibility charts, which provide a parameter-space characterization of the instability. In this work, we assess the capabilities of machine-learning (ML) models to learn and predict such susceptibility charts and analyze the constraints governing their generalization across materials. Using a simulation-derived dataset spanning six distinct secondary-electron-yield material profiles in a canonical two-surface planar geometry, we train supervised regression models and artificial neural networks to predict the time-averaged electron growth rate, δavg, across the relevant parameter space. Model performance is evaluated using metrics that explicitly probe the structure of susceptibility charts, including Intersection over Union, Structural Similarity Index, and correlation analysis. Tree-based ensemble models outperform neural-network models in reconstructing susceptibility regions and in generalizing across material domains. Principal-component analysis reveals disjoint material feature distributions, indicating that the piecewise mode structure of multipactor susceptibility is difficult to represent with a single global model and that generalization is constrained by data coverage rather than by model complexity. An exhaustive reduced-coverage study further shows that sparse material-space coverage can yield mean performance in the same general range but producing large variability in the susceptibility-region overlap. These results clarify the capabilities of ML-based surrogate models for parameter-space characterization of multipactor discharge. They also provide guidance for their appropriate use in RF system design.

43 PARTICLE ACCELERATORS