Search NASA⌕ Search

SEARCH · Search NASA

Results for “cluster computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Computational ranking identifies Plexin-B2 in circulating tumor cell clustering with monocytes in breast cancer metastasis

Abstract Multicellular circulating tumor cell (CTC) clusters can be up to 50 times more efficient than single CTCs in mediating viable metastasis. Here, combining computational ranking and functional determination, we identify the transmembrane protein Plexin-B2 (PLXNB2) as one of the top molecular targets associated with unfavorable distant metastasis-free survival, showing enriched expression in CTC clusters versus single CTCs from patients with advanced breast cancer (mostly female). Loss of PLXNB2 (Plxnb2) reduces the formation of homotypic tumor cell clusters and heterotypic tumor-myeloid cell clusters, reducing spontaneous metastases in female mice bearing human (mouse) breast cancer. Interactions of PLXNB2 with its ligands SEMA4C on tumor cells and SEMA4A on myeloid cells (monocytes) promote homotypic and heterotypic CTC cluster formation, respectively, thereby driving lung metastasis. Global proteomic analysis reveals downstream effectors of the PLXNB2 pathway associated with tumor cell clustering. Thus, PLXNB2 is a therapeutic target for preventing new metastasis in breast cancer.

Science & Technology - Other Topics↗

Computer program documentation: ISOCLS iterative self-organizing clustering program, program C094

The author has identified the following significant results. This program implements an algorithm which, ideally, sorts a given set of multivariate data points into similar groups or clusters. The program is intended for use in the evaluation of multispectral scanner data; however, the algorithm could be used for other data types as well. The user may specify a set of initial estimated cluster means to begin the procedure, or he may begin with the assumption that all the data belongs to one cluster. The procedure is initiatized by assigning each data point to the nearest (in absolute distance) cluster mean. If no initial cluster means were input, all of the data is assigned to cluster 1. The means and standard deviations are calculated for each cluster.

Minter, R. T.↗

Understanding the Scalability of Bayesian Network Inference using Clique Tree Growth Curves

Bayesian networks (BNs) are used to represent and efficiently compute with multi-variate probability distributions in a wide range of disciplines. One of the main approaches to perform computation in BNs is clique tree clustering and propagation. In this approach, BN computation consists of propagation in a clique tree compiled from a Bayesian network. There is a lack of understanding of how clique tree computation time, and BN computation time in more general, depends on variations in BN size and structure. On the one hand, complexity results tell us that many interesting BN queries are NP-hard or worse to answer, and it is not hard to find application BNs where the clique tree approach in practice cannot be used. On the other hand, it is well-known that tree-structured BNs can be used to answer probabilistic queries in polynomial time. In this article, we develop an approach to characterizing clique tree growth as a function of parameters that can be computed in polynomial time from BNs, specifically: (i) the ratio of the number of a BN's non-root nodes to the number of root nodes, or (ii) the expected number of moral edges in their moral graphs. Our approach is based on combining analytical and experimental results. Analytically, we partition the set of cliques in a clique tree into different sets, and introduce a growth curve for each set. For the special case of bipartite BNs, we consequently have two growth curves, a mixed clique growth curve and a root clique growth curve. In experiments, we systematically increase the degree of the root nodes in bipartite Bayesian networks, and find that root clique growth is well-approximated by Gompertz growth curves. It is believed that this research improves the understanding of the scaling behavior of clique tree clustering, provides a foundation for benchmarking and developing improved BN inference and machine learning algorithms, and presents an aid for analytical trade-off studies of clique tree clustering using growth curves.

Mengshoel, Ole Jakob↗

Forest and range mapping in the Houston area with ERTS-1

ERTS-1 data acquired over the Houston area has been analyzed for applications to forest and range mapping. In the field of forestry the Sam Houston National Forest (Texas) was chosen as a test site, (Scene ID 1037-16244). Conventional imagery interpretation as well as computer processing methods were used to make classification maps of timber species, condition and land-use. The results were compared with timber stand maps which were obtained from aircraft imagery and checked in the field. The preliminary investigations show that conventional interpretation techniques indicated an accuracy in classification of 63 percent. The computer-aided interpretations made by a clustering technique gave 70 percent accuracy. Computer-aided and conventional multispectral analysis techniques were applied to range vegetation type mapping in the gulf coast marsh. Two species of salt marsh grasses were mapped.

Heath, G. R.↗

Possibilistic clustering for shape recognition

Clustering methods have been used extensively in computer vision and pattern recognition. Fuzzy clustering has been shown to be advantageous over crisp (or traditional) clustering in that total commitment of a vector to a given class is not required at each iteration. Recently fuzzy clustering methods have shown spectacular ability to detect not only hypervolume clusters, but also clusters which are actually 'thin shells', i.e., curves and surfaces. Most analytic fuzzy clustering approaches are derived from Bezdek's Fuzzy C-Means (FCM) algorithm. The FCM uses the probabilistic constraint that the memberships of a data point across classes sum to one. This constraint was used to generate the membership update equations for an iterative algorithm. Unfortunately, the memberships resulting from FCM and its derivatives do not correspond to the intuitive concept of degree of belonging, and moreover, the algorithms have considerable trouble in noisy environments. Recently, we cast the clustering problem into the framework of possibility theory. Our approach was radically different from the existing clustering methods in that the resulting partition of the data can be interpreted as a possibilistic partition, and the membership values may be interpreted as degrees of possibility of the points belonging to the classes. We constructed an appropriate objective function whose minimum will characterize a good possibilistic partition of the data, and we derived the membership and prototype update equations from necessary conditions for minimization of our criterion function. In this paper, we show the ability of this approach to detect linear and quartic curves in the presence of considerable noise.

Keller, James M.↗

Possibilistic clustering for shape recognition

Clustering methods have been used extensively in computer vision and pattern recognition. Fuzzy clustering has been shown to be advantageous over crisp (or traditional) clustering in that total commitment of a vector to a given class is not required at each iteration. Recently fuzzy clustering methods have shown spectacular ability to detect not only hypervolume clusters, but also clusters which are actually 'thin shells', i.e., curves and surfaces. Most analytic fuzzy clustering approaches are derived from Bezdek's Fuzzy C-Means (FCM) algorithm. The FCM uses the probabilistic constraint that the memberships of a data point across classes sum to one. This constraint was used to generate the membership update equations for an iterative algorithm. Unfortunately, the memberships resulting from FCM and its derivatives do not correspond to the intuitive concept of degree of belonging, and moreover, the algorithms have considerable trouble in noisy environments. Recently, the clustering problem was cast into the framework of possibility theory. Our approach was radically different from the existing clustering methods in that the resulting partition of the data can be interpreted as a possibilistic partition, and the membership values may be interpreted as degrees of possibility of the points belonging to the classes. An appropriate objective function whose minimum will characterize a good possibilistic partition of the data was constructed, and the membership and prototype update equations from necessary conditions for minimization of our criterion function were derived. The ability of this approach to detect linear and quartic curves in the presence of considerable noise is shown.

Keller, James M.↗

Visualization of Unsteady Computational Fluid Dynamics

The current compute environment that most researchers are using for the calculation of 3D unsteady Computational Fluid Dynamic (CFD) results is a super-computer class machine. The Massively Parallel Processors (MPP's) such as the 160 node IBM SP2 at NAS and clusters of workstations acting as a single MPP (like NAS's SGI Power-Challenge array and the J90 cluster) provide the required computation bandwidth for CFD calculations of transient problems. If we follow the traditional computational analysis steps for CFD (and we wish to construct an interactive visualizer) we need to be aware of the following: (1) Disk space requirements. A single snap-shot must contain at least the values (primitive variables) stored at the appropriate locations within the mesh. For most simple 3D Euler solvers that means 5 floating point words. Navier-Stokes solutions with turbulence models may contain 7 state-variables. (2) Disk speed vs. Computational speeds. The time required to read the complete solution of a saved time frame from disk is now longer than the compute time for a set number of iterations from an explicit solver. Depending, on the hardware and solver an iteration of an implicit code may also take less time than reading the solution from disk. If one examines the performance improvements in the last decade or two, it is easy to see that depending on disk performance (vs. CPU improvement) may not be the best method for enhancing interactivity. (3) Cluster and Parallel Machine I/O problems. Disk access time is much worse within current parallel machines and cluster of workstations that are acting in concert to solve a single problem. In this case we are not trying to read the volume of data, but are running the solver and the solver outputs the solution. These traditional network interfaces must be used for the file system. (4) Numerics of particle traces. Most visualization tools can work upon a single snap shot of the data but some visualization tools for transient problems require dealing with time.

Haimes, Robert↗

Radio Sources toward Galaxy Clusters at 30 GHz

Extragalactic radio sources are a significant contaminant in cosmic microwave background and Sunyaev-Zel'dovich effect experiments. Deep interferometric observations with the BIMA and OVRO arrays are used to characterize the spatial, spectral, and flux distributions of radio sources toward massive galaxy clusters at 28.5 GHz. We compute counts of millijansky source fluxes from 89 fields centered on known massive galaxy clusters and 8 noncluster fields. We find that source counts in the inner regions of the cluster fields (within 0.5' of the cluster center) are a factor of 8.9 (sup +4.3)(sub -2.8) times higher than counts in the outer regions of the cluster fields (radius greater than 0.5'). Counts in the outer regions of the cluster fields are, in turn, a factor of 3.3 (sup +4.1) (sub -1.8) greater than those in the noncluster fields. Counts in the noncluster fields are consistent with extrapolations from the results of other surveys. We compute the spectral indices of millijansky sources in the cluster fields between 1.4 and 28.5 GHz and find a mean spectral index of alpha = 0.66 with an rms dispersion of 0.36, where flux S proportional to nu(sup -alpha). The distribution is skewed, with a median spectral index of 0.72 and 25th and 75th percentiles of 0.51 and 0.92, respectively. This is steeper than the spectral indices of stronger field sources measured by other surveys.

Coble, K.↗

Accurate In Bond Energies

InXn atomization energies are computed for n = 1-3 and X = H, Cl, and CH3. The geometries and frequencies are determined using density functional theory. The atomization energies are computed at the coupled cluster level of theory. The complete basis set limit is obtained by extrapolation. The scalar relativistic effect is computed using the Douglas-Kroll approach. While the heats of formation for InH, InCl and InCl3 are in good agreement with experiment, the current results show that the experimental value for In(CH3)3 must be wrong.

Bauschlicher, Charles W., Jr.↗

A Fast Implementation of the ISOCLUS Algorithm

Unsupervised clustering is a fundamental building block in numerous image processing applications. One of the most popular and widely used clustering schemes for remote sensing applications is the ISOCLUS algorithm, which is based on the ISODATA method. The algorithm is given a set of n data points in d-dimensional space, an integer k indicating the initial number of clusters, and a number of additional parameters. The general goal is to compute the coordinates of a set of cluster centers in d-space, such that those centers minimize the mean squared distance from each data point to its nearest center. This clustering algorithm is similar to another well-known clustering method, called k-means. One significant feature of ISOCLUS over k-means is that the actual number of clusters reported might be fewer or more than the number supplied as part of the input. The algorithm uses different heuristics to determine whether to merge lor split clusters. As ISOCLUS can run very slowly, particularly on large data sets, there has been a growing .interest in the remote sensing community in computing it efficiently. We have developed a faster implementation of the ISOCLUS algorithm. Our improvement is based on a recent acceleration to the k-means algorithm of Kanungo, et al. They showed that, by using a kd-tree data structure for storing the data, it is possible to reduce the running time of k-means. We have adapted this method for the ISOCLUS algorithm, and we show that it is possible to achieve essentially the same results as ISOCLUS on large data sets, but with significantly lower running times. This adaptation involves computing a number of cluster statistics that are needed for ISOCLUS but not for k-means. Both the k-means and ISOCLUS algorithms are based on iterative schemes, in which nearest neighbors are calculated until some convergence criterion is satisfied. Each iteration requires that the nearest center for each data point be computed. Naively, this requires O(kn) time, where k denotes the current number of centers. Traditional techniques for accelerating nearest neighbor searching involve storing the k centers in a data structure. However, because of the iterative nature of the algorithm, this data structure would need to be rebuilt with each new iteration. Our approach is to store the data points in a kd-tree data structure. The assignment of points to nearest neighbors is carried out by a filtering process, which successively eliminates centers that can not possibly be the nearest neighbor for a given region of space. This algorithm is significantly faster, because large groups of data points can be assigned to their nearest center in a single operation. Preliminary results on a number of real Landsat datasets show that our revised ISOCLUS-like scheme runs about twice as fast.

Memarsadeghi, Nargess↗

Arcminute fluctuations in the microwave background from clusters of galaxies

A method for computing arcmin microwave fluctuations produced by Compton scattering of the cosmic background photons by hot electrons in clusters of galaxies is described. Microwave images of the sky for a range of Omega and primordial fluctuation spectral index n are generated which are then 'observed' to determine Delta T/T in precisely the same manner as actual observations to determine if the cluster-induced fluctuations are consistent with the measured upper limit. The geometry used by Uson and Wilkinson (1984) in the NRAO experiment and Readhead et al. (1989) in the OVRO experiment are applied to the simulated images. The 95 percent confidence lower limit for Omega is found to be about 1/10 for n = -1 (which approximates the CDM mass spectrum for clusters), while for n = 0 it is 1/7; for n = +1 the limit is 1/5 if the gas density profile extends to five core radii.

Markevitch, M.↗

Data Management as a Cluster Middleware Centerpiece

Through earth and space modeling and the ongoing launches of satellites to gather data, NASA has become one of the largest producers of data in the world. These large data sets necessitated the creation of a Data Management System (DMS) to assist both the users and the administrators of the data. Halcyon Systems Inc. was contracted by the NASA Center for Computational Sciences (NCCS) to produce a Data Management System. The prototype of the DMS was produced by Halcyon Systems Inc. (Halcyon) for the Global Modeling and Assimilation Office (GMAO). The system, which was implemented and deployed within a relatively short period of time, has proven to be highly reliable and deployable. Following the prototype deployment, Halcyon was contacted by the NCCS to produce a production DMS version for their user community. The system is composed of several existing open source or government-sponsored components such as the San Diego Supercomputer Center s (SDSC) Storage Resource Broker (SRB), the Distributed Oceanographic Data System (DODS), and other components. Since Data Management is one of the foremost problems in cluster computing, the final package not only extends its capabilities as a Data Management System, but also to a cluster management system. This Cluster/Data Management System (CDMS) can be envisioned as the integration of existing packages.

Zero, Jose↗

Accurate ab initio quartic force fields for the ions HCO(+) and HOC(+)

The quartic force fields of HCO(+) and HOC(+) have been computed using augmented coupled cluster methods and basis sets of spdf and spdfg quality. Calculations on HCN, CO, and N2 have been performed to assist in calibrating the computed results. Going from an spdf to an spdfg basis shortens triple bonds by about 0.004 A, and increases the corresponding harmonic frequency by 10-20/cm, leaving bond distances about 0.003 A too long and triple bond stretching frequencies about 5/cm too low. Accurate estimates for the bond distances, fundamental frequencies, and thermochemical quantities are given. HOC(+) lies 37.8 +/- 0.5 kcal/mol (0 K) above HCO(+); the classical barrier height for proton exchange is 76.7 +/- 1.0 kcal/mol.

Martin, J. M. L.↗

Efficient Agent-Based Cluster Ensembles

Numerous domains ranging from distributed data acquisition to knowledge reuse need to solve the cluster ensemble problem of combining multiple clusterings into a single unified clustering. Unfortunately current non-agent-based cluster combining methods do not work in a distributed environment, are not robust to corrupted clusterings and require centralized access to all original clusterings. Overcoming these issues will allow cluster ensembles to be used in fundamentally distributed and failure-prone domains such as data acquisition from satellite constellations, in addition to domains demanding confidentiality such as combining clusterings of user profiles. This paper proposes an efficient, distributed, agent-based clustering ensemble method that addresses these issues. In this approach each agent is assigned a small subset of the data and votes on which final cluster its data points should belong to. The final clustering is then evaluated by a global utility, computed in a distributed way. This clustering is also evaluated using an agent-specific utility that is shown to be easier for the agents to maximize. Results show that agents using the agent-specific utility can achieve better performance than traditional non-agent based methods and are effective even when up to 50% of the agents fail.

Agogino, Adrian↗

Investigating the crust of neutron stars with neural-network quantum states

An accurate description of low-density nuclear matter is crucial for explaining the physics of neutron star crusts. In the density range between approximately 0.01 fm −3 and 0.1 fm −3 , matter transitions from neutron-rich nuclei to various higher-density pasta shapes, before ultimately reaching a uniform liquid. In this work, we introduce a variational Monte Carlo method based on a neural Pfaffian-Jastrow quantum state, which allows us to model the transition from the liquid phase to neutron-rich nuclei microscopically. At low densities, nuclear clusters dynamically emerge from the microscopic interactions among protons and neutrons, which we model based on pionless effective field theory. Our variational Monte Carlo approach represents a significant improvement over the state-of-the-art auxiliary-field diffusion Monte Carlo method, which is severely hindered by the fermion-sign problem in this low-density regime and cannot capture the onset of clusters. In addition to computing the energy per particle of symmetric nuclear matter and pure neutron matter, we analyze an intermediate isospin-asymmetry configuration to elucidate the formation of nuclear clusters. We also provide evidence that the presence of such nuclear clusters influences the amount of protons in the crust compared to protons in beta-equilibrated, neutrino-transparent matter.

Nuclear astrophysics↗

The Dissociation Energies of AlH2 and AlAr

The D(sub 0) values for AlH2 and AlAr are computed using the coupled cluster approach in conjunction with large basis sets. Basis set superposition and spin-orbit effects are accounted for as they are sizeable due to the small binding energy. The computed dissociation energy for AlAr is 101 /cm , which is 83% of the experimental value (122.4/ cm). Our best estimate for the H2 binding energy in AlH2 is 40 +/- 28 /cm.

Ricca, Alessandra↗

Photoelectron Spectroscopy and Computational Study on Microsolvated [B 10 H 10 ] 2– Clusters and Comparisons to Their [B 12 H 12 ] 2– Analogues

Microhydrated closo-Boranes have attracted great interests due to their superchaotropic activity related to well-known Hofmeister effect and important applications in biomedical and battery fields. In this work, we report a combined negative ion photoelectron spectroscopy and quantum chemical investigation on hydrated closo-decaborate clusters [B 10 H 10 ] 2- ·nH 2 O (n = 1 – 7) with a direct comparison to their analogues [B 12 H 12 ] 2- ·nH 2 O and free water clusters. A single H 2 O molecule is found sufficient to stabilize the intrinsically unstable [B 10 H 10 ] 2- dianion. The first two water molecules strongly interact with the solute forming B-H···H-O dihydrogen bonds while additional water molecules show substantially reduced binding energies. Unlike [B 12 H 12 ] 2- ·nH 2 O possessing highly structured water network with the attached H 2 O molecules arranged in a unified pattern by maximizing B-H···H-O dihydrogen bonding, distinct structural arrangements of the water clusters within [B 10 H 10 ] 2– ·nH 2 O are achieved with the water cluster networks from trimer to heptamer resembling free water clusters. Such a distinct difference arises from the variations in size, symmetry, and charge distributions between these two dianions. Finally, the present finding again confirms the structural diversity of hydrogen-bonding networks in microhydrated closo-boranes and enrich our understanding of aqueous borate chemistry.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Microcarbonation of Naphthalene: An Experimental and Computational Study of Photoionization in Naphthalene-Carbon Dioxide Clusters

The photoionization of naphthalene (N)-carbon dioxide (CO 2 ) clusters was studied using tunable vacuum ultraviolet (VUV) radiation from a synchrotron in the photon range of 8.0 to 13.7 eV, in combination with time-of-flight mass spectrometry. Clusters of monomer, dimer, and trimer naphthalene with CO 2 (N­(CO 2 ) 0–6 , N 2 ­(CO 2 ) 0–3 , N3) were observed. The lowest-energy conformers were obtained via a conformer search, followed by geometry optimizations at the ωB97X-V2/aug-cc-pVTZ (monomer) and ωB97X-V2/aug-cc-pVDZ (dimer) levels of theory. Carbon dioxide was found to preferentially cluster on top of the naphthalene molecule (in an out-of-plane configuration). From the mass spectra, photoionization intensity curves (PICs) were constructed, and appearance energies (AEs) were determined. No substantial trend in AE was observed with increasing size of the naphthalene-carbon dioxide clusters; rather, AE oscillations around the value for pure naphthalene were observed. These AE oscillations are also observed in a recently studied naphthalene-water cluster system, though in this system a slight downward trend in AE for pure naphthalene clusters (N1/2/3/4) was observed. The differences between the two systems are attributed to differing interaction strengths. In conclusion, understanding these differences may aid in determining how photoprocessing can proceed differently depending on the dominant matrix component in interstellar ices.

Wannenmacher, Anna [Lawrence Berkeley National Lab↗