Search NASA⌕ Search

SEARCH · Search NASA

Results for “clustering statistics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Performance analysis of the Alliant FX/8 multiprocessor using statistical clustering

Results for two distinct, real, scientific workloads executed on an Alliant FX/8 are discussed. A combination of user concurrency and system overhead measurements was taken for both workloads. Preliminary analysis shows that the first sampled workload is comprised of consistently high user concurrency, low system overhead, and little paging. The second sample has much less user concurrency, but significant paging and system overhead. Statistical cluster analysis is used to extract a state transition model to jointly characterize user concurrency and system overhead. A skewness factor is introduced and used to bring out the effects of unbalanced clustering when determining states with important transitions. The results from the models show that during the collection of the first sample, the system was operating in states of high user concurrency approximately 75 percent of the time. The second workload sample shows the system in high user concurrency states only 26 percent of the time. In addition, it is ascertained that high system overhead is usually accompanied by low user concurrency. The analysis also shows a high predictability of system behavior for both workloads.

Dimpsey, Robert Tod↗

Discovery of Activities via Statistical Clustering of Fixation Patterns

Human behavior often consists of a series of distinct activities, each characterized by a unique pattern of interaction with the visual environment. This is true even in a restricted domain, such as a piloting an aircraft, where activities with distinct visual signatures might be things like communicating, navigating, and monitoring. We propose a novel analysis method for gaze-tracking data, to perform blind discovery of these hypothetical activities. The method is in some respects similar to recurrence analysis, but here we compare not individual fixations, but groups of fixations aggregated over a fixed time interval. The duration of this interval is a parameter that we will refer to as delta. We assume that the environment has been divided into a set of N different areas-of-interest (AOIs). For a given interval of time of duration delta, we compute the proportion of time spent fixating each AOI, resulting in an N-dimensional vector. These proportions can be converted to integer counts by multiplying by delta divided by the average fixation duration (another parameter that we fix at 280 milliseconds). We compare different intervals by computing the chi-square statistic. The p-value associated with the statistic is the likelihood of observing the data under the hypothesis that the data in the two intervals were generated by a single process with a single set of probabilities governing the fixation of each AOI. The method has been applied to approximately 100 hours of eye movement data collected from pilots in a high-fidelity B747 flight simulator, and the results have been compared to synthetic data in which the each activity is represented as first-order Markov process with random probabilities assigned to the AOIs. Randomly-generated synthetic activities can require thousands of fixations to be discriminated with statistical significance, while the human data can be clustered using averaging windows of some 10's of seconds, suggesting that the actual activities are much more narrowly focused than random Markov models.

activity analysis↗

Discovery of Activities via Statistical Clustering of Fixation Patterns

Human behavior often consists of a series of distinct activities, each characterized by a unique pattern of interaction with the visual environment. This is true even in a restricted domain, such as a pilot flying an airplane; in this case, activities with distinct visual signatures might be things like communicating, navigating, monitoring, etc. We propose a novel analysis method for gaze-tracking data, to perform blind discovery of these hypothetical activities. We compare, not individual fixations, but groups of fixations aggregated over a fixed time interval (Tau). We assume that the environment has been divided into a finite set of discrete areas-of-interest (AOIs). For a given time interval, we compute the proportion of time spent fixating each AOI, resulting in an N-dimensional vector, where N is the number of AOIs. These proportions can be converted to integer counts by multiplying by Tau divided by the average fixation duration, a parameter that we fix at 283 milliseconds. We compare different intervals by computing the chi-squared statistic. The p-value associated with the statistic is the likelihood of observing the data under the hypothesis that the data in the two intervals were generated by a single process with a single set of probabilities governing the fixation of each AOI. We cluster the intervals, first by merging adjacent intervals that are sufficiently similar, optionally shifting the boundary between non-merged intervals to maximize the difference. Then we compare and cluster non-adjacent intervals. The method is evaluated using synthetic data generated by a hand-crafted set of activities. While the method generally finds more activities than put into the simulation, we have obtained agreement as high as 80 percent between the inferred activity labels and ground truth.

Eye Movements↗

Discovery of Activities via Statistical Clustering of Fixation Patterns

Human behavior often consists of a series of distinct activities, each characterized by a unique signature of visual behavior. This is true even in a restricted domain, such as piloting an aircraft, where patterns of visual signatures might represent activities like communicating, navigating, and monitoring. We propose a novel analysis method for gaze-tracking data, to perform blind discovery of these activities based on their behavioral signatures. The method is in some respects similar to recurrence analysis, but here we compare not individual fixations, but groups of fixations aggregated over a fixed time interval. The duration of this interval is a parameter that we will refer to as τ. We assume that the environment has been divided into a set of N different areas-of-interest (AOIs). For a given interval of time of duration τ, we compute the proportion of time spent fixating each AOI, resulting in an N-dimensional vector. These proportions can be converted to counts by multiplying by τ divided by the average fixation duration (another parameter that we fix at 280 milliseconds). We compare different intervals by computing the chi-square statistic. The p-value associated with the statistic is the likelihood of observing the data under the hypothesis that the data in the two intervals were generated by a single process with a single set of probabilities governing the fixation of each AOI. We have investigated the method using a set of 10 synthetic "activities," that sample 4 AOIs. Four of these activities visit 3 of the 4 AOIs, with equal probability; as there are four different ways to leave-one- out, there are four such activities. Similarly, there are six different activities that leave-two-out. Sequences of simulated behavior were generated by running each activity for 40 seconds, in sequence, for a total of 6.7 minutes. The figure to the right shows the matrix of chi-square statistics, using a value of 2.8 seconds for τ, corresponding to 10 fixations. Low values (dark) indicate poor evidence for activity differences, while high values (bright) indicate strong evidence. The dark squares along the main diagonal each correspond to the forty second intervals in which the activity was held constant; the 4x4 block at the lower left corresponds to the four leave-one-out activities, while the 6x6 block in the upper right corresponds to the leave-two-out activities. (The anti-diagonal pattern of white squares indicates those activity pairs that share no AOIs.) The chi-square values can be binarized by choosing a particular significance level; we are interested in grouping bins that represent the same activity, effectively accepting the null hypothesis. Therefore, we may adopt a relatively lax criterion; for example, choosing a p-value of 0.2 means that two behaviors that have only a 1-in-5 chance of being produced by a single activity might nevertheless be clustered together. We have explored several methods to perform clustering on the data and solving for the activity probabilities. Greedy methods begin by selecting the time bin that is similar to the most (or least) other bins, and then forming a cluster from it and all other non-discriminable bins. These methods show mediocre performance, as they do not take into account temporal contiguity. Preliminary results indicate that methods that "grow" clusters in time from seed points perform better.

activity analysis↗

Self-organization of cosmic radiation pressure instability. II - One-dimensional simulations

The clustering of statistically uniform discrete absorbing particles moving solely under the influence of radiation pressure from uniformly distributed emitters is studied in a simple one-dimensional model. Radiation pressure tends to amplify statistical clustering in the absorbers; the absorbing material is swept into empty bubbles, the biggest bubbles grow bigger almost as they would in a uniform medium, and the smaller ones get crushed and disappear. Numerical simulations of a one-dimensional system are used to support the conjecture that the system is self-organizing. Simple statistics indicate that a wide range of initial conditions produce structure approaching the same self-similar statistical distribution, whose scaling properties follow those of the attractor solution for an isolated bubble. The importance of the process for large-scale structuring of the interstellar medium is briefly discussed.

Hogan, Craig J.↗

An unsupervised classification technique for multispectral remote sensing data.

Description of a two-part clustering technique consisting of (a) a sequential statistical clustering, which is essentially a sequential variance analysis, and (b) a generalized K-means clustering. In this composite clustering technique, the output of (a) is a set of initial clusters which are input to (b) for further improvement by an iterative scheme. This unsupervised composite technique was employed for automatic classification of two sets of remote multispectral earth resource observations. The classification accuracy by the unsupervised technique is found to be comparable to that by traditional supervised maximum-likelihood classification techniques.

Su, M. Y.↗

Using Clustering to Establish Climate Regimes from PCM Output

A multivariate statistical clustering technique--based on the k-means algorithm of Hartigan has been used to extract patterns of climatological significance from 200 years of general circulation model (GCM) output. Originally developed and implemented on a Beowulf-style parallel computer constructed by Hoffman and Hargrove from surplus commodity desktop PCs, the high performance parallel clustering algorithm was previously applied to the derivation of ecoregions from map stacks of 9 and 25 geophysical conditions or variables for the conterminous U.S. at a resolution of 1 sq km. Now applied both across space and through time, the clustering technique yields temporally-varying climate regimes predicted by transient runs of the Parallel Climate Model (PCM). Using a business-as-usual (BAU) scenario and clustering four fields of significance to the global water cycle (surface temperature, precipitation, soil moisture, and snow depth) from 1871 through 2098, the authors' analysis shows an increase in spatial area occupied by the cluster or climate regime which typifies desert regions (i.e., an increase in desertification) and a decrease in the spatial area occupied by the climate regime typifying winter-time high latitude perma-frost regions. The patterns of cluster changes have been analyzed to understand the predicted variability in the water cycle on global and continental scales. In addition, representative climate regimes were determined by taking three 10-year averages of the fields 100 years apart for northern hemisphere winter (December, January, and February) and summer (June, July, and August). The result is global maps of typical seasonal climate regimes for 100 years in the past, for the present, and for 100 years into the future. Using three-dimensional data or phase space representations of these climate regimes (i.e., the cluster centroids), the authors demonstrate the portion of this phase space occupied by the land surface at all points in space and time. Any single spot on the globe will exist in one of these climate regimes at any single point in time. By incrementing time, that same spot will trace out a trajectory or orbit between and among these climate regimes (or atmospheric states) in phase (or state) space. When a geographic region enters a state it never previously visited, a climatic change is said to have occurred. Tracing out the entire trajectory of a single spot on the globe yields a 'manifold' in state space representing the shape of its predicted climate occupancy. This sort of analysis enables a researcher to more easily grasp the multivariate behavior of the climate system.

Oglesby, Robert↗

Mimas: Preliminary Evidence For Amorphous Water Ice from VIMS

We have conducted a statistical clustering analysis (1,2) on a mosaic of VIMS data cubes obtained on February 13, 2010, for Saturn s satellite Mimas. Seven VIMS cubes were geometrically projected and re-sampled to a common spatial resolution. The clustering technique consists of a partitioning algorithm coupled to a criterion that prevents sub-optimal solutions and tests for the influence of random noise in the measurements. The clustering technique is agnostic about the meaning of the clusters, and scientific interpretation requires their a posteriori evaluation. The preliminary results yielded five clusters, demonstrating that spectral variability across Mimas surface is statistically significant. The ratios of the means calculated for each of the clusters show structure within the 1.6- micron water ice band, as well as the shape and the central wavelength of the strong ice band at 2 micron, that map spatially in patterns apparently related to the topography of Mimas, in particular certain regions in and around Herschel crater. The mean spectra of the five clusters, show similarities with laboratory spectra of amorphous and crystalline H2O ice (3) that are suggestive of the presence of an amorphous ice component in certain regions of Mimas, notably on the central peak of Herschel, on the crater floor, and in faults surrounding the crater. This may represent a mixture of both ice phases, or perhaps a layer of amorphous ice on a base of crystalline ice. Another possible occurrence of amorphous ice appears southwest of Herschel, close to the south pole.

Cruikshank, Dale P.↗

The composite sequential clustering technique for analysis of multispectral scanner data

The clustering technique consists of two parts: (1) a sequential statistical clustering which is essentially a sequential variance analysis, and (2) a generalized K-means clustering. In this composite clustering technique, the output of (1) is a set of initial clusters which are input to (2) for further improvement by an iterative scheme. This unsupervised composite technique was employed for automatic classification of two sets of remote multispectral earth resource observations. The classification accuracy by the unsupervised technique is found to be comparable to that by traditional supervised maximum likelihood classification techniques. The mathematical algorithms for the composite sequential clustering program and a detailed computer program description with job setup are given.

Su, M. Y.↗

Unsupervised classification of earth resources data.

A new clustering technique is presented. It consists of two parts: (a) a sequential statistical clustering which is essentially a sequential variance analysis and (b) a generalized K-means clustering. In this composite clustering technique, the output of (a) is a set of initial clusters which are input to (b) for further improvement by an iterative scheme. This unsupervised composite technique was employed for automatic classification of two sets of remote multispectral earth resource observations. The classification accuracy by the unsupervised technique is found to be comparable to that by existing supervised maximum liklihood classification technique.

Su, M. Y.↗

A search for extended halos of hot gas in the Perseus, Virgo, and Coma Clusters

Observations of the Perseus cluster by the HEAO 1 satellite have revealed a faint X-ray halo extending at least 2.5 deg from the center and contributing between 5% and 20% to the total luminosity. This may be of nonthermal origin, but it also may be explained in terms of hot gas bound by the gravitational field of the cluster. Statistical uncertainties made it impossible to detect any such halo in the Coma cluster. Observations of the Virgo cluster confirmed the detection by the Ariel 5 satellite of a broad region of faint X-ray emission (core radius 60 arcmin). If the very extended X-ray emission from Virgo is due to hot intracluster gas, the density of this gas is lower than expected from a consideration of gas and galaxy densities in the Perseus cluster.

Ulmer, M. P.↗

An observational view of large scale structure

A summary of recent observations of galaxy clustering is presented, including a brief review of redshift maps and galaxy clustering statistics. Simple arguments are presented that argue the underlying mass fluctuations are most likely associated with a clustering scale no larger than that of individual galaxies. The acceleration of the local group from the comoving frame of the universe and its connection to the microwave dipole anisotropy are also discussed. A final topic for consideration is the existence of large voids and clusters, and whether they are consistent with Gaussian initial conditions. The extreme size and depth of the Bootes void, if real, do present a puzzle. Finally, future directions for observational study of large scale structure are discussed.

Davis, Marc↗

Land cover stratification using Landsat Thematic Mapper data in Sahelian and Sudanian woodland and wooded grassland

A standard methodology for thematic mapping of natural vegetation using remotely sensed imagery and digital image processing was modified to account for the spatial and spectral properties of semi-arid landscapes, and tested in study areas in the Sahelian and Sudanian zones, Mali. A principal components transformation of registered wet and dry season Landsat TM images produced a set of synthetic spectral channels differentiating vegetation cover between seasons, and allowed areas with annual grass growth to be distinguished from areas with woody cover. The transformed data were statistically clustered and clusters were assigned to vegetation type and density categories. In a separate step, the images were manually interpreted to differentiate broad soil classes. Four statistics were compared to evaluate the accuracy of the maps based on sample points from air photos. For the relatively detailed categories initially defined, map accuracies were substandard; however, when vegetation density classes were aggregated, overall accuracy was around 90 percent, and class accuracy was greater than 80 percent for most classes. This method is suitable for stratification and inventory of woody biomass at a regional scale in semi-arid woodland and wooded grassland.

Franklin, J.↗

Disentangling Structures in the Cluster of Galaxies Abell 133

A dynamical analysis of the structure of the cluster of galaxies Abell 133 will be presented using multi-wavelength data combined from multiple space and earth based observations. New and familiar statistical clustering techniques are used in combination in an attempt to gain a fully consistent picture of this interesting nearby cluster of galaxies. The type of analysis presented should be typical of cluster studies in the future, especially those to come from the surveys like the Sloan Digital Sky Survey and the 2DF.

Way, Michael J.↗

Seasonal- and Beta-Angle-Dependent Latitude Bias Variations in Natural Decays

Prior work has demonstrated pronounced statistical clustering of natural decays of medium-to-high-inclination orbital objects peaking approximately 30 degrees in Argument of Latitude ahead of nodal crossings. This effect is caused by the physical bulge in the Earth and the overlying atmosphere, that cyclically modifies effective altitude (and therefore density) faster than the trajectory's decay itself. While prior work has averaged seasonal and RAAN effects over all non-uniform atmosphere possibilities to support long-term characterization of the clustering of final entries in generating a pre-mission Expectation of Casualty, the current study characterizes seasonal and beta angle effects as potential influences on the near-term statistical risks of specific tactical decay scenarios, relative to the average. Such effects on the density profile along an orbit may be important considerations in any scenario where small orbital adjustments are used to optimize the timing and location of final entry trajectories. I.E., two identical spacecraft entering in different seasons and/or beta angles may have different minimum-risk scenarios for identical control capabilities and space weather conditions. Further, the early heating history of shallow trajectories is explored, examining the influence of dramatically different density profiles over the final orbit as the spacecraft either skims over or dives into the atmosphere.

Bacon, John B.↗

Mapping of terrain by computer clustering techniques using multispectral scanner data and using color aerial film

Two clustering techniques were used for terrain mapping by computer of test sites in Yellowstone National Park. One test was made with multispectral scanner data using a composite technique which consists of (1) a strictly sequential statistical clustering which is a sequential variance analysis, and (2) a generalized K-means clustering. In this composite technique, the output of (1) is a first approximation of the cluster centers. This is the input to (2) which consists of steps to improve the determination of cluster centers by iterative procedures. Another test was made using the three emulsion layers of color-infrared aerial film as a three-band spectrometer. Relative film densities were analyzed using a simple clustering technique in three-color space. Important advantages of the clustering technique over conventional supervised computer programs are (1) human intervention, preparation time, and manipulation of data are reduced, (2) the computer map, gives unbiased indication of where best to select the reference ground control data, (3) use of easy to obtain inexpensive film, and (4) the geometric distortions can be easily rectified by simple standard photogrammetric techniques.

Smedes, H. W.↗