Search NASASearch

SEARCH · Search NASA

Results for “cluster analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Cluster Analysis of IRIS Spectroscopic Line Profiles and SDO/AIA EUV Emission in Observations and RMHD Simulations of the Solar Atmosphere

Spatially-resolved observations from the IRIS and SDO/AIA satellites, especially when coupled with realistic 3D RMHD simulations, are a powerful tool for analysis of processes in the solar chromosphere, transition region, and corona. However, the complexity of the data makes understanding the observations and modeling results difficult. In this work, we apply unsupervised clustering algorithms for analysis of observational and synthetic chromospheric Mg II h&k 2796Å&2803Å and transition region C II 1334Å&1335Å line profiles observed by IRIS, and extreme ultraviolet (EUV) emission observed by SDO/AIA, for various types of problems. The synthetic line profiles are computed for simulations of the quiescent solar atmosphere (using the StellarBox and RH1.5 codes). The K-Means clustering algorithm is applied, and the selection of an optimal number of clusters is supported by the average silhouette width technique. We discuss applications of the line profile clustering method to 1) visualization of computational and observational spectroscopic imaging data; 2) understanding of evolutionary trends and behavior patterns of quiet Sun emission and during solar flares; and 3) recognition of heating events and shock waves.

Sadykov, Viacheslav

Cluster analysis based on dimensional information with applications to feature selection and classification

A new clustering algorithm is presented that is based on dimensional information. The algorithm includes an inherent feature selection criterion, which is discussed. Further, a heuristic method for choosing the proper number of intervals for a frequency distribution histogram, a feature necessary for the algorithm, is presented. The algorithm, although usable as a stand-alone clustering technique, is then utilized as a global approximator. Local clustering techniques and configuration of a global-local scheme are discussed, and finally the complete global-local and feature selector configuration is shown in application to a real-time adaptive classification scheme for the analysis of remote sensed multispectral scanner data.

Eigen, D. J.

Application of High-Dimensional Fuzzy K-Means Cluster Analysis to CALIOP/CALIPSO Version 4.1 Cloud-Aerosol Discrimination

This study applies fuzzy k-means (FKM) cluster analyses to a subset of the parameters reported in the CALIPSO lidar level 2 data products in order to classify the layers detected as either clouds or aerosols. The results obtained are used to assess the reliability of the cloud–aerosol discrimination (CAD) scores reported in the version 4.1 release of the CALIPSO data products. FKM is an unsupervised learning algorithm, whereas the CALIPSO operational CAD algorithm (COCA) takes a highly supervised approach. Despite these substantial computational and architectural differences, our statistical analyses show that the FKM classifications agree with the COCA classifications for more than 94 % of the cases in the troposphere. This high degree of similarity is achieved because the lidar-measured signatures of the majority of the clouds and the aerosols are naturally distinct, and hence objective methods can independently and effectively separate the two classes in most cases. Classification differences most often occur in complex scenes (e.g., evaporating water cloud filaments embedded in dense aerosol) or when observing diffuse features that occur only intermittently (e.g., volcanic ash in the tropical tropopause layer). The two methods examined in this study establish overall classification correctness boundaries due to their differing algorithm uncertainties. In addition to comparing the outputs from the two algorithms, analysis of sampling, data training, performance measurements, fuzzy linear discriminants, defuzzification, error propagation, and key parameters in feature type discrimination with the FKM method are further discussed in order to better understand the utility and limits of the application of clustering algorithms to space lidar measurements. In general, we find that both FKM and COCA classification uncertainties are only minimally affected by noise in the CALIPSO measurements, though both algorithms can be challenged by especially complex scenes containing mixtures of discrete layer types. Our analysis results show that attenuated backscatter and color ratio are the driving factors that separate water clouds from aerosols; backscatter intensity, depolarization, and mid-layer altitude are most useful in discriminating between aerosols and ice clouds; and the joint distribution of backscatter intensity and depolarization ratio is critically important for distinguishing ice clouds from water clouds.

Zeng, Shan

Cluster Analysis of Spectroscopic Line Profiles and EUV Emission in RMHD Simulations and Observations of the Solar Atmosphere

Spatially-resolved observations from the IRIS, SDO/AIA, and other space mission and ground-based telescopes, coupled with realistic 3D RMHD simulations, are a powerful tool for analysis of processes in the solar atmosphere. To better understand the dynamical and thermodynamic properties in the simulation data and their connection to observations, it is essential to determine similarities in the behaviors of the synthesized and observed emission. However, the complexity of observational data and physical processes makes comparison of observations and modeling results difficult. In this work, we show the initial results of application of K-Means clustering (unsupervised machine learning) algorithm to two different problems: 1) recognition of the typical spectroscopic line profiles observed by IRIS during solar flares and their typical dynamic behavior; 2) recognition of shocks and heating events in synthetic AIA emission data obtained from StellarBox quiet-Sun simulations. The average silhouette width technique for the KMeans algorithm is utilized in different ways to obtain optimal numbers of clusters. We discuss application of the emission clustering to visualizations of the computational volume, understanding its evolutionary trends and behavior patterns, and inversion (reconstruction) of physical properties of the solar atmosphere from synthesizes emission data.

Sadykov, Viacheslav

Failure Mode Identification Through Clustering Analysis

Research has shown that nearly 80% of the costs and problems are created in product development and that cost and quality are essentially designed into products in the conceptual stage. Currently, failure identification procedures (such as FMEA (Failure Modes and Effects Analysis), FMECA (Failure Modes, Effects and Criticality Analysis) and FTA (Fault Tree Analysis)) and design of experiments are being used for quality control and for the detection of potential failure modes during the detail design stage or post-product launch. Though all of these methods have their own advantages, they do not give information as to what are the predominant failures that a designer should focus on while designing a product. This work uses a functional approach to identify failure modes, which hypothesizes that similarities exist between different failure modes based on the functionality of the product/component. In this paper, a statistical clustering procedure is proposed to retrieve information on the set of predominant failures that a function experiences. The various stages of the methodology are illustrated using a hypothetical design example.

Arunajadai, Srikesh G.

Exploring HOD-dependent systematics for the DESI 2024 Full-Shape galaxy clustering analysis

We analyze the robustness of the DESI 2024 cosmological inference from the full shape of the galaxy power spectrum to uncertainties in the Halo Occupation Distribution (HOD) model of the galaxy-halo connection and the choice of priors on nuisance parameters. We assess variations in the recovered cosmological parameters across a range of mocks populated with different HOD models and find that shifts are often greater than 20% of the expected statistical uncertainties from the DESI data. We encapsulate the effect of such shifts in terms of a systematic covariance term, C HOD , and an additional diagonal contribution quantifying the impact of our choice of nuisance parameter priors on the ability of the effective field theory (EFT) model to correctly recover the cosmological parameters of the simulations. These two covariance contributions are designed to be added to the usual covariance term, C stat , describing the statistical uncertainty in the power spectrum measurement, in order to fairly represent these sources of systematic uncertainty. This novel approach should be more general and robust to the choice of model or additional external datasets used in cosmological fits than the alternative approach of adding systematic uncertainties to the recovered marginalised parameter posteriors. We compare the approaches within the context of a fixed ΛCDM model and demonstrate that our method gives conservative estimates of the systematic uncertainty that nevertheless have little impact on the final posteriors obtained from DESI data.

79 ASTRONOMY AND ASTROPHYSICS

Time-dependent clustering analysis of the second BATSE gamma-ray burst catalog

A time-dependent two-point correlation-function analysis of the Burst and Transient Source Experiment (BATSE) 2B catalog finds no evidence of burst repetition. As part of this analysis, we discuss the effects of sky exposure on the observability of burst repetition and present the equation describing the signature of burst repetition in the data. For a model of all burst repetition from a source occurring in less than five days we derive upper limits on the number of bursts in the catalog from repeaters and model-dependent upper limits on the fraction of burst sources that produce multiple outbursts.

Brainerd, J. J.

The Ocean Carbon States Database: A Proof-of-Concept Application of Cluster Analysis in the Ocean Carbon Cycle

In this paper, we present a database of the basic regimes of the carbon cycle in the ocean, the 'ocean carbon states', as obtained using a data mining/pattern recognition technique in observation-based as well as model data. The goal of this study is to establish a new data analysis methodology, test it and assess its utility in providing more insights into the regional and temporal variability of the marine carbon cycle. This is important as advanced data mining techniques are becoming widely used in climate and Earth sciences and in particular in studies of the global carbon cycle, where the interaction of physical and biogeochemical drivers confounds our ability to accurately describe, understand, and predict CO2 concentrations and their changes in the major planetary carbon reservoirs. In this proof-of-concept study, we focus on using well-understood data that are based on observations, as well as model results from the NASA Goddard Institute for Space Studies (GISS) climate model. Our analysis shows that ocean carbon states are associated with the subtropical-subpolar gyre during the colder months of the year and the tropics during the warmer season in the North Atlantic basin. Conversely, in the Southern Ocean, the ocean carbon states can be associated with the subtropical and Antarctic convergence zones in the warmer season and the coastal Antarctic divergence zone in the colder season. With respect to model evaluation, we find that the GISS model reproduces the cold and warm season regimes more skillfully in the North Atlantic than in the Southern Ocean and matches the observed seasonality better than the spatial distribution of the regimes. Finally, the ocean carbon states provide useful information in the model error attribution. Model air-sea CO2 flux biases in the North Atlantic stem from wind speed and salinity biases in the subpolar region and nutrient and wind speed biases in the subtropics and tropics. Nutrient biases are shown to be most important in the Southern Ocean flux bias.

carbon cycle

Cluster Analysis of Thermal Icequakes Using the Seismometer to Investigate Ice and Ocean Structure (SIIOS): Implications for Ocean World Seismology

Ocean Worlds are of high interest to the planetary community due to the potential habitability of their subsurface oceans. Over the next few decades several missions will be sent to ocean worlds including the Europa Clipper, Dragonfly, and possibly a Europa lander. The Dragonfly and Europa lander missions will carry seismic payloads tasked with detecting and locating seismic sources. The Seismometer to Investigate Ice and Ocean Structure (SIIOS) is a NASA PSTAR funded project that investigates ocean world seismology using terrestrial analogs. The goals of the SIIOS experiment include quantitatively comparing flight-candidate seismometers to traditional instruments, comparing single-station approaches to a small-aperture array, and characterizing the local seismic environment of our field sites. Here we present an analysis of detected local events at our field sites at Gulkana Glacier in Alaska and in Northwest Greenland approximately 80 km North of Qaanaaq, Greenland. Both field sites passively recorded data for about two weeks. We deployed our experiment on Gulkana Glacier in September 2017 and in Greenland in June 2018. At Gulkana there was a nearby USGS weather station which recorded wind data. Temperature data was collected using the MERRA satellite. In Greenland we deployed our own weather station to collect temperature and wind data. Gulkana represents a noisier and more active environment. Temperatures fluctuated around 0°C, allowing for surface runoff to occur during the day. The glacier had several moulins, and during deployment we heard several rockfalls from nearby mountains. In addition to the local environment, Gulkana is located close to an active plate boundary (relative to Greenland). This meant that there were more regional events recorded over two weeks, than in Greenland. Greenland’s local environment was also quieter, and less active. Temperatures remained below freezing. The Greenland ice was much thicker than Gulkana (~850 m versus ~100 m) and our stations were above a subglacial lake. Both conditions can reduce event detections from basal motion. Lastly, we encased our Greenland array in an aluminum vault and buried it beneath the surface unlike our array in Gulkana where the instruments were at the surface and covered with plastic bins. The vault further insulated the array from thermal and atmospheric events.

Marusiak, A. G.

Machine Learning–Augmented Laser-Induced Breakdown Spectroscopy for Spectral Discrimination of Iron Oxalates

Enhanced characterization and phase identification of post-PUREX Pu Oxalates (PuOXA) are pivotal for nonproliferation and pre-detonation nuclear forensics. Despite significant advances in the characterization of PuO 2 samples, little is known about the impact of both the chemical structure and oxidation states of PuOXA (i.e., Pu(III) and Pu(IV)) have on optical emission signatures. Here, we demonstrate the analytical capabilities of laser-induced breakdown spectroscopy (LIBS) applied to Fe(II) and Fe(III) oxalate samples as surrogates for PuOXA, highlighting the discriminating features in the LIBS emission spectra arising from differences in the oxidation states within mixed FeOXA samples. We report the enhancement of spectral feature selection using Principal Component Analysis (PCA), which enables the analytical superiority of machine learning algorithms such as Linear Discriminant Analysis (LDA), Quadratic Discriminant Analysis (QDA), Partial Least Squares Regression (PLSR), Support Vector Regression (SVR), and Random Forest Regression (RFR) over conventional univariate techniques for phase discrimination and chemometric analysis. Cluster analysis revealed how both matrix effects and laser ablation influence cluster separability by introducing spectral artifacts that misdirect the maximization of variance. PCA-selected emission lines were used in the regression models, demonstrating that both univariate and multivariate linear regression models (i.e., PLSR and SVR) can achieve acceptable performance, with machine learning models outperforming conventional calibration regressions. Furthermore, the application of non-linearly activated PCA-selected emission lines illustrates how simplifying the data while retaining captured variance enables the use of less complex and more computationally efficient models. Furthermore, this is particularly evident in the underperformance of RFR, which suffers from increased computational costs and overfitting owing to its high complexity.

Oxalates

An evaluation of air quality in major urban areas of India

Rapid economic growth and burgeoning population have contributed to enhanced levels of PM 2.5 concentrations in urban regions of India. Evaluation of ambient air quality facilitates the assessment of effectiveness of emission control measures and early identification of new sources. This study provides a comprehensive statistical analysis of PM 2.5 concentrations in key urban areas across India, including Delhi, Kolkata, Mumbai, Chennai, Hyderabad, and several regional centers. Data from 2017 to 2023 was analyzed using trend analysis, cluster analysis, principal component analysis, and geostatistical interpolation to understand spatiotemporal variations and sources. The analysis reveals significant differences in spatial distribution of PM 2.5 concentrations with high annual averages in urban regions in Indo-Gangetic plain (82–123 μg m −3 ) and relatively lower concentrations (29–46 μg m −3 ) in southern urban areas of Kerala, Tamil Nadu and Andhra Pradesh. Delhi state had the highest 24-averaged PM 2.5 concentrations (112 μg m −3 ) followed by urban regions in Uttar Pradesh, Bihar and West Bengal (94 μg m −3 ). Trend analysis from 2017 to 2023 revealed an overall 2.5% decline in site-wide PM2.5 concentrations, with the exception of Ludhiana, which exhibited a consistent annual increase of 10%. Principal component analysis (PCA) attributes 30% of the variance to wintertime emissions, 13% to biomass burning, and 18% to the regional haze in the northern Indo-Gangetic Plain. Different analyses clearly demonstrates the contribution of biomass burning to pollution in Delhi and surrounding cities. Transboundary pollution to Kolkata is likely from the highly polluted region in Indo-Gangetic Plain. Coastal cities of Mumbai and Chennai has relatively lower pollution attributed to the influence of sea breeze dilution, with mostly local contribution and some potential transport from upwind industry clusters. Hyderabad also has local contribution due to high density of vehicular traffic and local small industries. This study shows that mitigation efforts targeting clusters of regions should be undertaken to curb the high PM2.5 pollution. Policy measures should be implemented both at local and the intra-state level to address shared sources and transport of pollution.

Hysplitbacktrajectories

Analyzing Multidimensional Image Data

Six computer programs perform histogram cluster analysis. Histogram Cluster Analysis Procedure (HICAP) developed to perform unsupervised classification of multidimensional image data. Clustering approach used in HICAP based on algorithm which uses multidimensional histogram to perform unsupervised classification of four-dimensional Landsat multispectral-scanner data. HICAP generalizes this procedure to process up to 32-bit data with arbitrary number of dimensions. Also incorporates efficiency improvements so classification requires less computation than original algorithm. Computational savings afforded by HICAP increase with number of dimensions in data. HICAP programs written in FORTRAN 77 for batch or interactive execution.

Wharton, S. W.