Search NASA⌕ Search

SEARCH · Search NASA

Results for “distributed clustering methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Efficient Agent-Based Cluster Ensembles

Numerous domains ranging from distributed data acquisition to knowledge reuse need to solve the cluster ensemble problem of combining multiple clusterings into a single unified clustering. Unfortunately current non-agent-based cluster combining methods do not work in a distributed environment, are not robust to corrupted clusterings and require centralized access to all original clusterings. Overcoming these issues will allow cluster ensembles to be used in fundamentally distributed and failure-prone domains such as data acquisition from satellite constellations, in addition to domains demanding confidentiality such as combining clusterings of user profiles. This paper proposes an efficient, distributed, agent-based clustering ensemble method that addresses these issues. In this approach each agent is assigned a small subset of the data and votes on which final cluster its data points should belong to. The final clustering is then evaluated by a global utility, computed in a distributed way. This clustering is also evaluated using an agent-specific utility that is shown to be easier for the agents to maximize. Results show that agents using the agent-specific utility can achieve better performance than traditional non-agent based methods and are effective even when up to 50% of the agents fail.

Agogino, Adrian↗

X-ray halos in galaxies and clusters of galaxies - Theory

X-ray measurements provide an excellent method to determine the amount and distribution of the dark matter in clusters. Unfortunately, accurate temperature profiles, necessary to this method, are currently not available. However, if the intracluster gas is assumed to have a monotonically decreasing temperature, one finds that the dark matter is strongly concentrated to the cluster center, and has a mass which only exceeds the known baryonic mass by a factor of about three. On a second topic, cooling flows are shown to be a very common feature of cluster central and normal elliptical galaxies. The cooling gas is probably ultimately converted into low mass stars.

Sarazin, Craig L.↗

Transient nucleation in condensed systems

Using classical nucleation theory we consider transient nucleation occurring in a one-component, condensed system under isothermal conditions. We obtain an exact closed-form expression for the time dependent cluster populations. In addition, a more versatile approach is developed: a numerical simulation technique which models directly the reactions by which clusters are produced. This simulation demonstrates the evolution of cluster populations and nucleation rate in the transient regime. Results from the simulation are verified by comparison with exact analytical solutions for the steady state. Experimental methods for measuring transient nucleation are assessed, and it is demonstrated that the observed behavior depends on the method used. The effect of preexisting cluster distributions is studied. Previous analytical and numerical treatments of transient nucleation are compared to the solutions obtained from the simulation. The simple expressions of Kashchiev are shown to give good descriptions of the nucleation behavior.

Kelton, K. F.↗

Technical support for creating an artificial intelligence system for feature extraction and experimental design

Techniques for classifying objects into groups or clases go under many different names including, most commonly, cluster analysis. Mathematically, the general problem is to find a best mapping of objects into an index set consisting of class identifiers. When an a priori grouping of objects exists, the process of deriving the classification rules from samples of classified objects is known as discrimination. When such rules are applied to objects of unknown class, the process is denoted classification. The specific problem addressed involves the group classification of a set of objects that are each associated with a series of measurements (ratio, interval, ordinal, or nominal levels of measurement). Each measurement produces one variable in a multidimensional variable space. Cluster analysis techniques are reviewed and methods for incuding geographic location, distance measures, and spatial pattern (distribution) as parameters in clustering are examined. For the case of patterning, measures of spatial autocorrelation are discussed in terms of the kind of data (nominal, ordinal, or interval scaled) to which they may be applied.

Glick, B. J.↗

Structured background grids for generation of unstructured grids by advancing front method

A new method of background grid construction is introduced for generation of unstructured tetrahedral grids using the advancing-front technique. Unlike the conventional triangular/tetrahedral background grids which are difficult to construct and usually inadequate in performance, the new method exploits the simplicity of uniform Cartesian meshes and provides grids of better quality. The approach is analogous to solving a steady-state heat conduction problem with discrete heat sources. The spacing parameters of grid points are distributed over the nodes of a Cartesian background grid by interpolating from a few prescribed sources and solving a Poisson equation. To increase the control over the grid point distribution, a directional clustering approach is used. The new method is convenient to use and provides better grid quality and flexibility. Sample results are presented to demonstrate the power of the method.

Pirzadeh, Shahyar↗

The use of unsupervised clustering as a classifier for LACIE MSS data

The author has identified the following significant results. This classification method appears to give accurate field center results and to give practical, statistically consistent and accurate estimates of crop proportions. The accuracy of this method is attributable to certain qualities of the particular clustering algorithm. These qualities are freedom from assumptions about Gaussian data, and the continual updating of distribution estimates, including updating the number of modes. This method is relatively tolerant of errors in the determination of crop type, as crop identity is used only for identifying clusters, and not for computing signatures.

Pentland, A. P.↗

Performance and Application of Parallel OVERFLOW Codes on Distributed and Shared Memory Platforms

The presentation discusses recent studies on the performance of the two parallel versions of the aerodynamics CFD code, OVERFLOW_MPI and _MLP. Developed at NASA Ames, the serial version, OVERFLOW, is a multidimensional Navier-Stokes flow solver based on overset (Chimera) grid technology. The code has recently been parallelized in two ways. One is based on the explicit message-passing interface (MPI) across processors and uses the _MPI communication package. This approach is primarily suited for distributed memory systems and workstation clusters. The second, termed the multi-level parallel (MLP) method, is simple and uses shared memory for all communications. The _MLP code is suitable on distributed-shared memory systems. For both methods, the message passing takes place across the processors or processes at the advancement of each time step. This procedure is, in effect, the Chimera boundary conditions update, which is done in an explicit "Jacobi" style. In contrast, the update in the serial code is done in more of the "Gauss-Sidel" fashion. The programming efforts for the _MPI code is more complicated than for the _MLP code; the former requires modification of the outer and some inner shells of the serial code, whereas the latter focuses only on the outer shell of the code. The _MPI version offers a great deal of flexibility in distributing grid zones across a specified number of processors in order to achieve load balancing. The approach is capable of partitioning zones across multiple processors or sending each zone and/or cluster of several zones into a single processor. The message passing across the processors consists of Chimera boundary and/or an overlap of "halo" boundary points for each partitioned zone. The MLP version is a new coarse-grain parallel concept at the zonal and intra-zonal levels. A grouping strategy is used to distribute zones into several groups forming sub-processes which will run in parallel. The total volume of grid points in each group are approximately balanced. A proper number of threads are initially allocated to each group, and in subsequent iterations during the run-time, the number of threads are adjusted to achieve load balancing across the processes. Each process exploits the multitasking directives already established in Overflow.

Djomehri, M. Jahed↗

Quantitative characterization of spatial distribution of particles in materials: Application to materials processing

Most engineering materials contain second phase particles or fibers which serve to reinforce the matrix phase. The effect of reinforcements on material properties is usually analyzed in terms of the average volume fraction and spacing of reinforcements, quantities which are global microstructural characteristics. However, material properties can also depend on local microstructural characteristics; for example, on how uniformly the reinforcing phase is distributed in the material. The analysis method will then be applied to a materials processing problem to discover how processing parameters can be selected to maximize redistribution of the reinforcing phase during processing. Several mathematical analysis methods could be adapted to the problem of characterizing the distribution of particles in materials. A tessellation-based method was selected. In the first phase of the investigation, a software package was written to automate the analysis. Typical results are shown. The analysis technique allows the degree to which particles are clustered together, the size and spacing of particle clusters, and the particle density in clusters to be found. The analysis methods were applied to computer-generated distributions and to a few real particle-containing materials. Methods for analyzing a nonuniform particle distribution in a material can be applied to two broad classes of materials science problems: understanding how the resulting particle distribution affects properties. The analysis method described is applied to a materials processing problem: how to select extrusion conditions to maximize the redistribution of reinforcing particles that are initially nonuniformly distributed. In addition, the tessellation-based method to analyze star distributions in spiral galaxies was adapted, illustrating the diverse types of problems to which the analysis method can be applied.

Parse, J. B.↗

Convergence rate enhancement of navier-stokes codes on clustered grids

Our Sensitivity-Based Minimal Residual (SBMR) method which is based on our earlier Distributed Minimal Residual (DMR) method allows each component of the solution vector in a system of equations to have its own convergence speed. Our global SBMR method was found to consistently outperform the DMR method while requiring considerably less computer memory. Recently, we have developed and tested a new Line SBMR or LSBMR method and a Time-Step-Scaling (TSS) method that are even more robust and computationally efficient than our global SBMR method, especially on highly clustered computational grids in laminar and turbulent flow computations.

Choi, Kwang-Yoon↗

An analysis of fracture trace patterns in areas of flat-lying sedimentary rocks for the detection of buried geologic structure

Two study areas in a cratonic platform underlain by flat-lying sedimentary rocks were analyzed to determine if a quantitative relationship exists between fracture trace patterns and their frequency distributions and subsurface structural closures which might contain petroleum. Fracture trace lengths and frequency (number of fracture traces per unit area) were analyzed by trend surface analysis and length frequency distributions also were compared to a standard Gaussian distribution. Composite rose diagrams of fracture traces were analyzed using a multivariate analysis method which grouped or clustered the rose diagrams and their respective areas on the basis of the behavior of the rays of the rose diagram. Analysis indicates that the lengths of fracture traces are log-normally distributed according to the mapping technique used. Fracture trace frequency appeared higher on the flanks of active structures and lower around passive reef structures. Fracture trace log-mean lengths were shorter over several types of structures, perhaps due to increased fracturing and subsequent erosion. Analysis of rose diagrams using a multivariate technique indicated lithology as the primary control for the lower grouping levels. Groupings at higher levels indicated that areas overlying active structures may be isolated from their neighbors by this technique while passive structures showed no differences which could be isolated.

Podwysocki, M. H.↗

CLASSY: An adaptive maximum likelihood clustering algorithm

The CLASSY clustering method alternates maximum likelihood iterative techniques for estimating the parameters of a mixture distribution with an adaptive procedure for splitting, combining, and eliminating the resultant components of the mixture. The adaptive procedure is based on maximizing the fit of a mixture of multivariate normal distributions to the observed data using its first through fourth central moments. It generates estimates of the number of multivariate normal components in the mixture as well as the proportion, mean vector, and covariance matrix for each component. The basic mathematical model for CLASSY and the actual operation of the algorithm as currently implemented are described. Results of applying CLASSY to real and simulated LANDSAT data are presented and compared with those generated by the iterative self-organizing clustering system algorithm on the same data sets.

Lennington, R. K.↗

Rotation of Low-mass Stars in Upper Centaurus-Lupus and Lower Centaurus-Crux with TESS

We present stellar rotation rates derived from Transiting Exoplanet Survey Satellite (TESS) light curves for stars in Upper Centaurus–Lupus (UCL; ∼136 pc, ∼16 Myr) and Lower Centaurus–Crux (LCC; ∼115 pc, ∼17 Myr). We find spot-modulated periods (P) for ∼90% of members. The range of light-curve and periodogram shapes echoes that found for other clusters with K2, but fewer multiperiod stars may be an indication of the different noise characteristics of TESS, or a result of the source selection methods here. The distribution of P as a function of color as a proxy for mass fits nicely in between that for both older and younger clusters observed by K2, with fast rotators being found among both the highest and lowest masses probed here, and a well-organized distribution of M-star rotation rates. About 13% of the stars have an infrared excess, suggesting a circumstellar disk; this is well matched to expectations, given the age of the stars. There is an obvious pileup of disked M stars at P ∼ 2 days, and the pileup may move to shorter P as the mass decreases. There is also a strong concentration of disk-free M stars at P ∼ 2 days, hinting that perhaps these stars have recently freed themselves from their disks. Exploring the rotation rates of stars in UCL/LCC has the potential to help us understand the beginning of the end of the influence of disks on rotation, and the timescale on which stars respond to unlocking.

L M Rebull↗

An X-ray method for detecting substructure in galaxy clusters - Application to Perseus, A2256, Centaurus, Coma, and Sersic 40/6

We use the moments of the X-ray surface brightness distribution to constrain the dynamical state of a galaxy cluster. Using X-ray observations from the Einstein Observatory IPC, we measure the first moment FM, the ellipsoidal orientation angle, and the axial ratio at a sequence of radii in the cluster. We argue that a significant variation in the image centroid FM as a function of radius is evidence for a nonequilibrium feature in the intracluster medium (ICM) density distribution. In simple terms, centroid shifts indicate that the center of mass of the ICM varies with radius. This variation is a tracer of continuing dynamical evolution. For each cluster, we evaluate the significance of variations in the centroid of the IPC image by computing the same statistics on an ensemble of simulated cluster images. In producing these simulated images we include X-ray point source emission, telescope vignetting, Poisson noise, and characteristics of the IPC. Application of this new method to five Abell clusters reveals that the core of each one has significant substructure. In addition, we find significant variations in the orientation angle and the axial ratio for several of the clusters.

Mohr, Joseph J.↗

The estimation of masses of individual galaxies in clusters of galaxies.

Three different methods of estimating masses are discussed. The 'density method' is based on the analysis of the density distribution of galaxies around the object whose mass is to be found. The 'bound-galaxy method' gives estimates of the mass of a double, triple, or quadruple system from analysis of the orbital motion of the components. The 'virial method' utilizes the formulas derived for the second method to obtain estimates of the virial-theorem masses of whole clusters, and thus to obtain upper limits on the mass of an individual galaxy in a cluster. The analytic formulas are developed and compared with computer experiments, and some applications are given.

Wolf, R. A.↗

A simple method for obtaining a three-dimensional proton distribution function from Voyager plasma data

The main sensor of the Vogager plasma experiment consists of a cluster of three, modulated-grid Faraday cups whose normals are arranged symmetrically about the symmetry axis of the cluster at an angle of 20 degrees to that axis. In interplanetary space, each cup explores the positive ion distribution by accepting particles from contiguous slices in velocity space. The slices are narrow in the direction of the normal to the modulating grid but are broad in planes parallel to that grid. The resulting three sets of measurements can be combined to yield the three-dimensional distribution function in the following way: the distribution function is assumed to be gyrotropic. For each value of speed in a frame of reference moving with the bulk velocity of the solar wind, the variation of the distribution function with angle from the field direction is represented by a series of Legendre polynomials. Effects such as double-streaming and heat flow can be well represented by using only the first three terms of the series which are fully specified by the measurements. Examples of the use of this method in the analysis of Voyager data are shown.

Olbert, S.↗

Improved Test Planning and Analysis Through the Use of Advanced Statistical Methods

The goal of this work is, through computational simulations, to provide statistically-based evidence to convince the testing community that a distributed testing approach is superior to a clustered testing approach for most situations. For clustered testing, numerous, repeated test points are acquired at a limited number of test conditions. For distributed testing, only one or a few test points are requested at many different conditions. The statistical techniques of Analysis of Variance (ANOVA), Design of Experiments (DOE) and Response Surface Methods (RSM) are applied to enable distributed test planning, data analysis and test augmentation. The D-Optimal class of DOE is used to plan an optimally efficient single- and multi-factor test. The resulting simulated test data are analyzed via ANOVA and a parametric model is constructed using RSM. Finally, ANOVA can be used to plan a second round of testing to augment the existing data set with new data points. The use of these techniques is demonstrated through several illustrative examples. To date, many thousands of comparisons have been performed and the results strongly support the conclusion that the distributed testing approach outperforms the clustered testing approach.

Green, Lawrence L.↗

Collaborative Simulation Grid: Multiscale Quantum-Mechanical/Classical Atomistic Simulations on Distributed PC Clusters in the US and Japan

A multidisciplinary, collaborative simulation has been performed on a Grid of geographically distributed PC clusters. The multiscale simulation approach seamlessly combines i) atomistic simulation backed on the molecular dynamics (MD) method and ii) quantum mechanical (QM) calculation based on the density functional theory (DFT), so that accurate but less scalable computations are performed only where they are needed. The multiscale MD/QM simulation code has been Grid-enabled using i) a modular, additive hybridization scheme, ii) multiple QM clustering, and iii) computation/communication overlapping. The Gridified MD/QM simulation code has been used to study environmental effects of water molecules on fracture in silicon. A preliminary run of the code has achieved a parallel efficiency of 94% on 25 PCs distributed over 3 PC clusters in the US and Japan, and a larger test involving 154 processors on 5 distributed PC clusters is in progress.

Kikuchi, Hideaki↗