Search NASA⌕ Search

SEARCH · Search NASA

Results for “k mean”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Reducing Earth Topography Resolution for SMAP Mission Ground Tracks Using K-Means Clustering

The K-means clustering algorithm is used to reduce Earth topography resolution for the SMAP mission ground tracks. As SMAP propagates in orbit, knowledge of the radar antenna footprints on Earth is required for the antenna misalignment calibration. Each antenna footprint contains a latitude and longitude location pair on the Earth surface. There are 400 pairs in one data set for the calibration model. It is computationally expensive to calculate corresponding Earth elevation for these data pairs. Thus, the antenna footprint resolution is reduced. Similar topographical data pairs are grouped together with the K-means clustering algorithm. The resolution is reduced to the mean of each topographical cluster called the cluster centroid. The corresponding Earth elevation for each cluster centroid is assigned to the entire group. Results show that 400 data points are reduced to 60 while still maintaining algorithm performance and computational efficiency. In this work, sensitivity analysis is also performed to show a trade-off between algorithm performance versus computational efficiency as the number of cluster centroids and algorithm iterations are increased.

ground tracks↗

K-Means Cluster Study for Radiofrequency Propagation Characterization

The objective of this study is to design a simple method for mining radio frequency (RF) propagation data. The study explored the characteristics of a large dataset of propagation experiments conducted over the span of years and using several ground stations around the world. Furthermore, this study developed simple predictive models that can be used for link characterization and overall propagation behavior description, without the need for physical measurements on-site. It is understood that such statistical learning has several drawbacks in terms of accuracy and precision. K-means clustering was used to characterize the data set in a way never explored before in an attempt to create useful tools that reduce cost, time and risk. K-means clustering was used to characterize the data set. Cosine distance was used as a method to determine the optimal number for clustering each feature. Dependence and independence analysis was performed to explore intra and inter-sensitivity between the presented features, with respect to each other and time. Several predicative models were generated and evaluated with respect to a test set to assess a measure of prediction accuracy and precision. A simple method for data analysis was developed and tested as the basis for further studies and future refinement to produce optimal performing models.

Cognitive↗

Application of High-Dimensional Fuzzy K-Means Cluster Analysis to CALIOP/CALIPSO Version 4.1 Cloud-Aerosol Discrimination

This study applies fuzzy k-means (FKM) cluster analyses to a subset of the parameters reported in the CALIPSO lidar level 2 data products in order to classify the layers detected as either clouds or aerosols. The results obtained are used to assess the reliability of the cloud–aerosol discrimination (CAD) scores reported in the version 4.1 release of the CALIPSO data products. FKM is an unsupervised learning algorithm, whereas the CALIPSO operational CAD algorithm (COCA) takes a highly supervised approach. Despite these substantial computational and architectural differences, our statistical analyses show that the FKM classifications agree with the COCA classifications for more than 94 % of the cases in the troposphere. This high degree of similarity is achieved because the lidar-measured signatures of the majority of the clouds and the aerosols are naturally distinct, and hence objective methods can independently and effectively separate the two classes in most cases. Classification differences most often occur in complex scenes (e.g., evaporating water cloud filaments embedded in dense aerosol) or when observing diffuse features that occur only intermittently (e.g., volcanic ash in the tropical tropopause layer). The two methods examined in this study establish overall classification correctness boundaries due to their differing algorithm uncertainties. In addition to comparing the outputs from the two algorithms, analysis of sampling, data training, performance measurements, fuzzy linear discriminants, defuzzification, error propagation, and key parameters in feature type discrimination with the FKM method are further discussed in order to better understand the utility and limits of the application of clustering algorithms to space lidar measurements. In general, we find that both FKM and COCA classification uncertainties are only minimally affected by noise in the CALIPSO measurements, though both algorithms can be challenged by especially complex scenes containing mixtures of discrete layer types. Our analysis results show that attenuated backscatter and color ratio are the driving factors that separate water clouds from aerosols; backscatter intensity, depolarization, and mid-layer altitude are most useful in discriminating between aerosols and ice clouds; and the joint distribution of backscatter intensity and depolarization ratio is critically important for distinguishing ice clouds from water clouds.

Zeng, Shan↗

A Fast Implementation of the ISOCLUS Algorithm

Unsupervised clustering is a fundamental tool in numerous image processing and remote sensing applications. For example, unsupervised clustering is often used to obtain vegetation maps of an area of interest. This approach is useful when reliable training data are either scarce or expensive, and when relatively little a priori information about the data is available. Unsupervised clustering methods play a significant role in the pursuit of unsupervised classification. One of the most popular and widely used clustering schemes for remote sensing applications is the ISOCLUS algorithm, which is based on the ISODATA method. The algorithm is given a set of n data points (or samples) in d-dimensional space, an integer k indicating the initial number of clusters, and a number of additional parameters. The general goal is to compute a set of cluster centers in d-space. Although there is no specific optimization criterion, the algorithm is similar in spirit to the well known k-means clustering method in which the objective is to minimize the average squared distance of each point to its nearest center, called the average distortion. One significant feature of ISOCLUS over k-means is that clusters may be merged or split, and so the final number of clusters may be different from the number k supplied as part of the input. This algorithm will be described in later in this paper. The ISOCLUS algorithm can run very slowly, particularly on large data sets. Given its wide use in remote sensing, its efficient computation is an important goal. We have developed a fast implementation of the ISOCLUS algorithm. Our improvement is based on a recent acceleration to the k-means algorithm, the filtering algorithm, by Kanungo et al.. They showed that, by storing the data in a kd-tree, it was possible to significantly reduce the running time of k-means. We have adapted this method for the ISOCLUS algorithm. For technical reasons, which are explained later, it is necessary to make a minor modification to the ISOCLUS specification. We provide empirical evidence, on both synthetic and Landsat image data sets, that our algorithm's performance is essentially the same as that of ISOCLUS, but with significantly lower running times. We show that our algorithm runs from 3 to 30 times faster than a straightforward implementation of ISOCLUS. Our adaptation of the filtering algorithm involves the efficient computation of a number of cluster statistics that are needed for ISOCLUS, but not for k-means.

Memarsadeghi, Nargess↗

Fast Image Texture Classification Using Decision Trees

Texture analysis would permit improved autonomous, onboard science data interpretation for adaptive navigation, sampling, and downlink decisions. These analyses would assist with terrain analysis and instrument placement in both macroscopic and microscopic image data products. Unfortunately, most state-of-the-art texture analysis demands computationally expensive convolutions of filters involving many floating-point operations. This makes them infeasible for radiation- hardened computers and spaceflight hardware. A new method approximates traditional texture classification of each image pixel with a fast decision-tree classifier. The classifier uses image features derived from simple filtering operations involving integer arithmetic. The texture analysis method is therefore amenable to implementation on FPGA (field-programmable gate array) hardware. Image features based on the "integral image" transform produce descriptive and efficient texture descriptors. Training the decision tree on a set of training data yields a classification scheme that produces reasonable approximations of optimal "texton" analysis at a fraction of the computational cost. A decision-tree learning algorithm employing the traditional k-means criterion of inter-cluster variance is used to learn tree structure from training data. The result is an efficient and accurate summary of surface morphology in images. This work is an evolutionary advance that unites several previous algorithms (k-means clustering, integral images, decision trees) and applies them to a new problem domain (morphology analysis for autonomous science during remote exploration). Advantages include order-of-magnitude improvements in runtime, feasibility for FPGA hardware, and significant improvements in texture classification accuracy.

Thompson, David R.↗

Comparison of wheat classification accuracy using different classifiers of the image-100 system

Classification results using single-cell and multi-cell signature acquisition options, a point-by-point Gaussian maximum-likelihood classifier, and K-means clustering of the Image-100 system are presented. Conclusions reached are that: a better indication of correct classification can be provided by using a test area which contains various cover types of the study area; classification accuracy should be evaluated considering both the percentages of correct classification and error of commission; supervised classification approaches are better than K-means clustering; Gaussian distribution maximum likelihood classifier is better than Single-cell and Multi-cell Signature Acquisition Options of the Image-100 system; and in order to obtain a high classification accuracy in a large and heterogeneous crop area, using Gaussian maximum-likelihood classifier, homogeneous spectral subclasses of the study crop should be created to derive training statistics.

Dejesusparada, N.↗

Coronado Ecological Conservation: Assessing Vegetation Change Due to Border Wall Construction and Shifting Social Trails

Species monitoring is essential for mitigating the impacts of plant invasion, such as radical changes in an area’s ecosystem, degraded soil health, increased wildfire severity, landslides, and increased flooding. For this project, NASA DEVELOP partnered with the National Park Service (NPS) to investigate invasive species in disturbed lands: specifically, areas affected by off-trail travel and U.S.-Mexico border construction activities. The team assessed how construction has impacted the distribution of Lehmann’s lovegrass and Russian thistle invasives throughout Coronado National Memorial, AZ from 1986-2022. Using data from Landsat 5 and 8, Sentinel-2, NAIP, and PlanetScope, the team computed NDVI, NDMI, MSAVI2, EVI, and Tasseled Cap Wetness, Brightness, and Greenness transformations as vegetation health indicators to input into various machine learning algorithms. To minimize noise, the team conducted Principal Component Analysis on vegetation indices and spectral bands before running k-means clustering and random forest classification algorithms. Between all datasets, the team found that the median area fully overtaken by invasive plants was 5.37% of the park’s total area in 2022. The NPS will use end products to help increase restoration efforts in disturbed areas with high concentrations of invasive plants, and this project can serve as a jumping off point for future invasive species monitoring. The NPS’s collection of ground data for 2022-2023, in conjunction with future data collection, will notably improve the accuracy of classification models, leading to more precise monitoring of invasive species spread over time.

Coronado National Memorial↗

Separation of internal and external fields: A new technique of data screening

A method for eliminating transient variation from MAGSAT data is described. Instead of using the conventional Kp index, the Km index is used for the rejection of nonquiet data. To build the Km index, the surface of the Earth is divided in eight sectors. Each sector is defined by two or three magnetic observatories and its geographic limits. For each three hourly interval the MAGSAT tracks were drawn on a map showing the location of the station (the crosses), the limits of the sector, and the value of the mean K index inside each sector. No figure in a sector means that the activity level in this sector is lower than 5 nt; in a sector covered with number 1, the activity lies between 5 and 20 nt; 2 stands for any level greater than 20 nT.

Lemouel, J. L.↗

Passive remote sensing of ocean optical propagation parameters

A method is described for producing a global data of ocean optical properties through the exploitation of optical remote sensing techniques. The diffuse attenuation, K, and the radiance reflectance factor R sub L, of the ocean surface waters can be derived from the radiance data provided by the Coastal Zone Color Scanner and compiled into a computer based atlas of these properties. While these remotely sensed values can only be directly related to the water properties in the first or upper attenuation length, extensive in water measurements have demonstrated that a useful correlation exists between the K that applies to the upper attenuation length, for example, and the mean K over the upper 100 or 200 meters of the ocean. Such an empirical finding greatly enhances the usefulness of this remotely sensed propagation parameter. Examples of the type of information being produced and archived for the atlas are presented.

Austin, R. W.↗

Free convection in the Martian atmosphere

Researchers investigated the free convective regime for the Martian atmospheric boundary layer (ABL). Researchers generalized Schumann's model describing horizontal fluctuations and mean vertical gradients occurring during free convection to include convection driven by water vapor gradients and to include the effects of circulation above both aerodynamically smooth and rough surfaces. Applying the model to Mars, researchers found that nearly all the resistance to sensible and latent heat transfer in the ABL occurs within the thin interfacial sublayer at the surface. Free convection is found to readily occur at low pressures and high temperatures when surface ice is present. At 7 mb, the ABL should freely convect whenever the mean windspeed at the top of the surface layer drops below about 2.5 m s(-1) and surface temperatures exceed 250 K. Mean horizontal fluctuations within the surface layer are found to be as high as 3 m (-1) for windspeed, 0.5 K for temperature, and 10 (-4) kg m (-3) for water vapor density. Airflow over surfaces similar to the Antarctic Polar Plateau was found to be aerodynamically smooth on Mars during free convection for all pressures between 6 and 1000 mb, while surfaces with z sub o approx. equals 1 cm are aerodynamically rough over this pressure range.

Clow, G. D.↗

Operational Dynamic Configuration Analysis

Sectors may combine or split within areas of specialization in response to changing traffic patterns. This method of managing capacity and controller workload could be made more flexible by dynamically modifying sector boundaries. Much work has been done on methods for dynamically creating new sector boundaries [1-5]. Many assessments of dynamic configuration methods assume the current day baseline configuration remains fixed [6-7]. A challenging question is how to select a dynamic configuration baseline to assess potential benefits of proposed dynamic configuration concepts. Bloem used operational sector reconfigurations as a baseline [8]. The main difficulty is that operational reconfiguration data is noisy. Reconfigurations often occur frequently to accommodate staff training or breaks, or to complete a more complicated reconfiguration through a rapid sequence of simpler reconfigurations. Gupta quantified a few aspects of airspace boundary changes from this data [9]. Most of these metrics are unique to sector combining operations and not applicable to more flexible dynamic configuration concepts. To better understand what sort of reconfigurations are acceptable or beneficial, more configuration change metrics should be developed and their distribution in current practice should be computed. This paper proposes a method to select a simple sequence of configurations among operational configurations to serve as a dynamic configuration baseline for future dynamic configuration concept assessments. New configuration change metrics are applied to the operational data to establish current day thresholds for these metrics. These thresholds are then corroborated, refined, or dismissed based on airspace practitioner feedback. The dynamic configuration baseline selection method uses a k-means clustering algorithm to select the sequence of configurations and trigger times from a given day of operational sector combination data. The clustering algorithm selects a simplified schedule containing k configurations based on stability score of the sector combinations among the raw operational configurations. In addition, the number of the selected configurations is determined based on balance between accuracy and assessment complexity.

Lai, Chok Fung↗

Computational Inference of Vibratory System with Incomplete Modal Information Using Parallel, Interactive and Adaptive Markov Chains

Inverse analysis of vibratory system is an important subject in fault identification, model updating, and robust design and control. It is challenging subject because 1) the problem is oftentimes underdetermined while the measurements are limited and/or incomplete; 2) many combinations of parameters may yield results that are similar with respect to actual response measurements; and 3) uncertainties inevitably exist. The aim of this research is to leverage upon computational intelligence through statistical inference to facilitate an enhanced, probabilistic framework using incomplete modal response measurement. This new framework is built upon efficient inverse identification through optimization, whereas Bayesian inference is employed to account for the effect of uncertainties. To overcome the computational cost barrier, we adopt Markov chain Monte Carlo (MCMC) to characterize the target function/distribution. Instead of using single Markov chain in conventional Bayesian approach, we develop a new sampling theory with multiple parallel, interactive and adaptive Markov chains and incorporate it into Bayesian inference. This can harness the collective power of these Markov chains to realize the concurrent search of multiple local optima. The number of required Markov chains and their respective initial model parameters are automatically determined via Monte Carlo simulation-based sample pre-screening followed by K-means clustering analysis. These enhancements can effectively address the aforementioned challenges in finite element inverse analysis. The validity of this framework is systematically demonstrated through case studies.

K Zhou↗

Towards Accurate Predictions of Martensitic Transition Temperatures for Shape Memory Alloys from Ab Initio Simulations

Experimentally, NiTi undergoes a single martensitic phase transition around 341 K from the lowtemperature (T) monoclinic B19’ phase (P21/m) to the high-temperature cubic B2 phase (Pm3m). Theoretically, an orthorhombic B33 (Cmcm) has also been proposed as the T=0 ground state structure, although this phase has never been observed. Accurate predictions of martensitic transition temperatures (MTT) have remained elusive in part due to several well-known theoretical complexities of these systems including low temperature instabilities of the B2 phase. Recently, we proposed a rigorous thermodynamic integration approach based on ab initio simulations to resolve many of these difficulties [1,2]. However, an unsatisfying overprediction of the MTT relative to experiment (by ~100 K) means a fully quantitative theory is still lacking. In this work, we report several new developments to our method that bring first principles theory and experiment much closer into agreement. Our calculations indicate that phonon free energies at low temperature stabilizes B19’ over B33, rationalizing B19’ as the ground state down to T=0. We also find that accurate computations of the electronic free energy, i.e. the change in energy and the appearance of electronic configurational entropy due to finite temperature, is crucial to obtain accurate MTT. Incorporating these corrections results in an MTT prediction of 365 K for binary NiTi, which is in very close agreement with experiment. Our theoretical approach is expected to be a broadly applicable and predictive theory for MTT of SMAs.

Zhigang Wu↗

Design data collection with Skylab microwave radiometer-scatterometer S-193, volume 2

The author has identified the following significant results. Skylab S-193 radiometer/scatterometer produced terrain responses with various polarizations and observation angles for cells of 100 to 400 sq km area. Classification of the observations into natural categories was achieved by K-means and spatial clustering algorithms. Microwave data acquired over the Great Salt Lake Desert area by sensors aboard Skylab and Nimbus 5 indicate that the microwave emission and backscatter were strongly influenced by contributions from subsurface layers of sediment saturated with brine. Correlations were noted between microwave backscatter response at approximately 33 deg from scatterometer (operating at 13.9 GHz) and the configuration of ground targets in Brazil as discerned from coarse scale maps. With limited, available ground truth, these correlations were sufficient to permit the production of image-like displays which bear a marked resemblance to known terrain features in several instances.

Moore, R. K.↗

Temperature and Flow Measurements in Incompressible Heated Jets

An experimental study is conducted to perform time-resolved temperature and velocity measurements on a Mach 0.08 jet at total temperatures of 295 K and 353 K. Mean and rms temperature data acquired using two fine wire sensors with diameters 1.3 and 3.8 μm are compared. In order to extend the limited frequency response of the wires, the temperature data are post-processed using a frequency compensation technique available in the literature. Corresponding velocity measurements are performed at cold and heated conditions using single and parallel wire probes. Simultaneously measured temperature and velocity data obtained using the parallel wire probe are used to calculate correlations pertinent to axial turbulent heat flux. In addition to shedding some light on the aerothermal properties of heated turbulent jets, vis-`a-vis their cold counterparts, this study also provides a database for numerical prediction of these flows.

Temperature measurement, turbulence, Jets, Turbule↗

Unsupervised classification of remote multispectral sensing data

The new unsupervised classification technique for classifying multispectral remote sensing data which can be either from the multispectral scanner or digitized color-separation aerial photographs consists of two parts: (a) a sequential statistical clustering which is a one-pass sequential variance analysis and (b) a generalized K-means clustering. In this composite clustering technique, the output of (a) is a set of initial clusters which are input to (b) for further improvement by an iterative scheme. Applications of the technique using an IBM-7094 computer on multispectral data sets over Purdue's Flight Line C-1 and the Yellowstone National Park test site have been accomplished. Comparisons between the classification maps by the unsupervised technique and the supervised maximum liklihood technique indicate that the classification accuracies are in agreement.

Su, M. Y.↗

Mapping of terrain by computer clustering techniques using multispectral scanner data and using color aerial film

Two clustering techniques were used for terrain mapping by computer of test sites in Yellowstone National Park. One test was made with multispectral scanner data using a composite technique which consists of (1) a strictly sequential statistical clustering which is a sequential variance analysis, and (2) a generalized K-means clustering. In this composite technique, the output of (1) is a first approximation of the cluster centers. This is the input to (2) which consists of steps to improve the determination of cluster centers by iterative procedures. Another test was made using the three emulsion layers of color-infrared aerial film as a three-band spectrometer. Relative film densities were analyzed using a simple clustering technique in three-color space. Important advantages of the clustering technique over conventional supervised computer programs are (1) human intervention, preparation time, and manipulation of data are reduced, (2) the computer map, gives unbiased indication of where best to select the reference ground control data, (3) use of easy to obtain inexpensive film, and (4) the geometric distortions can be easily rectified by simple standard photogrammetric techniques.

Smedes, H. W.↗

The composite sequential clustering technique for analysis of multispectral scanner data

The clustering technique consists of two parts: (1) a sequential statistical clustering which is essentially a sequential variance analysis, and (2) a generalized K-means clustering. In this composite clustering technique, the output of (1) is a set of initial clusters which are input to (2) for further improvement by an iterative scheme. This unsupervised composite technique was employed for automatic classification of two sets of remote multispectral earth resource observations. The classification accuracy by the unsupervised technique is found to be comparable to that by traditional supervised maximum likelihood classification techniques. The mathematical algorithms for the composite sequential clustering program and a detailed computer program description with job setup are given.

Su, M. Y.↗