Search NASASearch

SEARCH · Search NASA

Results for “clustering algorithm heterogeneous”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

LANDSAT-4 image data quality analysis

Seven heterogeneous areas within the entire Des Moines, Iowa test site were selected to define candidate spectral training classes using a clustering algorithm. In addition to the 91 cluster nonsupervised classes, three supervised training classes were defined. The original candidate training classes were reduced to 42 spectrally separable training classes. The minimum and average transformed divergence values for the 42 spectral classes and for the best subsets of Y TM spectral bands are shown in a table. The best spectral band for any combination of 1 through 7 bands is the first middle IR band. The next best band is the near IR, followed by the red band and than the thermal IR. The best combination of four bands includes one from each of the four regions of the spectrum (visible, near IR, middle IR, and thermal IR).

Anuta, P. E.

LANDSAT-4 image data quality analysis

Seven heterogeneous areas within the Des Moines, Iowa area test site were selected to define candidate spectral training classes using a clustering algorithm. In addition to the 91 cluster (nonsupervised) classes, three supervised training classes were defined and subsequently included in the training statistics file. The identity of all 94 candidate classes were determined using available reference data. Through analysis of the interclass separabilities, the original 94 candidate training classes were reduced to 42 spectrally separable final classes. The minimum average transformed divergence values for the 42 spectral classes and for the best subsets of TM spectral bands are shown in a table.

Anuta, P. E.

Interpretable Categorization of Heterogeneous Time Series Data

We analyze data from simulated aircraft encounters to validate and inform the development of a prototype aircraft collision avoidance system. The high-dimensional and heterogeneous time series dataset is analyzed to discover properties of near mid-air collisions (NMACs) and categorize the NMAC encounters. Domain experts use these properties to better organize and understand NMAC occurrences. Existing solutions either are not capable of handling high-dimensional and heterogeneous time series datasets or do not provide explanations that are interpretable by a domain expert. The latter is critical to the acceptance and deployment of safety-critical systems. To address this gap, we propose grammar-based decision trees along with a learning algorithm. Our approach extends decision trees with a grammar framework for classifying heterogeneous time series data. A context-free grammar is used to derive decision expressions that are interpretable, application-specific, and support heterogeneous data types. In addition to classification, we show how grammar-based decision trees can also be used for categorization, which is a combination of clustering and generating interpretable explanations for each cluster. We apply grammar-based decision trees to a simulated aircraft encounter dataset and evaluate the performance of four variants of our learning algorithm. The best algorithm is used to analyze and categorize near mid-air collisions in the aircraft encounter dataset. We describe each discovered category in detail and discuss its relevance to aircraft collision avoidance.

Drones

Analysis methods for Thematic Mapper data of urban regions

Studies have indicated the difficulty in deriving a detailed land-use/land-cover classification for heterogeneous metropolitan areas with Landsat MSS and TM data. The major methodological issues of digital analysis which possibly have effected the results of classification are examined. In response to these methodological issues, a multichannel hierarchical clustering algorithm has been developed and tested for a more complete analysis of the data for urban areas.

Wang, S. C.

A Parallel Particle Swarm Optimization Algorithm Accelerated by Asynchronous Evaluations

A parallel Particle Swarm Optimization (PSO) algorithm is presented. Particle swarm optimization is a fairly recent addition to the family of non-gradient based, probabilistic search algorithms that is based on a simplified social model and is closely tied to swarming theory. Although PSO algorithms present several attractive properties to the designer, they are plagued by high computational cost as measured by elapsed time. One approach to reduce the elapsed time is to make use of coarse-grained parallelization to evaluate the design points. Previous parallel PSO algorithms were mostly implemented in a synchronous manner, where all design points within a design iteration are evaluated before the next iteration is started. This approach leads to poor parallel speedup in cases where a heterogeneous parallel environment is used and/or where the analysis time depends on the design point being analyzed. This paper introduces an asynchronous parallel PSO algorithm that greatly improves the parallel e ciency. The asynchronous algorithm is benchmarked on a cluster assembled of Apple Macintosh G5 desktop computers, using the multi-disciplinary optimization of a typical transport aircraft wing as an example.

Venter, Gerhard

BOREAS AFM-12 1-km AVHRR Seasonal Land Cover Classification

The Boreal Ecosystem-Atmosphere Study (BOREAS) Airborne Fluxes and Meteorology (AFM)-12 team's efforts focused on regional scale Surface Vegetation and Atmosphere (SVAT) modeling to improve parameterization of the heterogeneous BOREAS landscape for use in larger scale Global Circulation Models (GCMs). This regional land cover data set was developed as part of a multitemporal one-kilometer Advanced Very High Resolution Radiometer (AVHRR) land cover analysis approach that was used as the basis for regional land cover mapping, fire disturbance-regeneration, and multiresolution land cover scaling studies in the boreal forest ecosystem of central Canada. This land cover classification was derived by using regional field observations from ground and low-level aircraft transits to analyze spectral-temporal clusters that were derived from an unsupervised cluster analysis of monthly Normalized Difference Vegetation Index (NDVI) image composites (April-September 1992). This regional data set was developed for use by BOREAS investigators, especially those involved in simulation modeling, remote sensing algorithm development, and aircraft flux studies. Based on regional field data verification, this multitemporal one-kilometer AVHRR land cover mapping approach was effective in characterizing the biome-level land cover structure, embedded spatially heterogeneous landscape patterns, and other types of key land cover information of interest to BOREAS modelers.The land cover mosaics in this classification include: (1) wet conifer mosaic (low, medium, and high tree stand density), (2) mixed coniferous-deciduous forest (80% coniferous, codominant, and 80% deciduous), (3) recent visible bum, vegetation regeneration, or rock outcrops-bare ground-sparsely vegetated slow regeneration bum (four classes), (4) open water and grassland marshes, and (5) general agricultural land use/ grasslands (three classes). This land cover mapping approach did not detect small subpixel-scale landscape features such as fens, bogs, and small water bodies. Field observations and comparisons with Landsat Thematic Mapper (TM) suggest a minimum effective resolution of these land cover classes in the range of three to four kilometers, in part, because of the daily to monthly compositing process. In general, potential accuracy limitations are mitigated by the use of conservative parameterization rules such as aggregation of predominant land cover classes within minimum horizontal grid cell sizes of ten kilometers. The AFM-12 one-kilometer AVHRR seasonal land cover classification data are available from the Earth Observing System Data and Information System (EOSDIS) Oak Ridge National Laboratory (ORNL) Distributed Active Archive Center (DAAC). The data files are available on a CD-ROM (see document number 20010000884).

Steyaert, Lou

Visual Computing Environment

The Visual Computing Environment (VCE) is a NASA Lewis Research Center project to develop a framework for intercomponent and multidisciplinary computational simulations. Many current engineering analysis codes simulate various aspects of aircraft engine operation. For example, existing computational fluid dynamics (CFD) codes can model the airflow through individual engine components such as the inlet, compressor, combustor, turbine, or nozzle. Currently, these codes are run in isolation, making intercomponent and complete system simulations very difficult to perform. In addition, management and utilization of these engineering codes for coupled component simulations is a complex, laborious task, requiring substantial experience and effort. To facilitate multicomponent aircraft engine analysis, the CFD Research Corporation (CFDRC) is developing the VCE system. This system, which is part of NASA's Numerical Propulsion Simulation System (NPSS) program, can couple various engineering disciplines, such as CFD, structural analysis, and thermal analysis. The objectives of VCE are to (1) develop a visual computing environment for controlling the execution of individual simulation codes that are running in parallel and are distributed on heterogeneous host machines in a networked environment, (2) develop numerical coupling algorithms for interchanging boundary conditions between codes with arbitrary grid matching and different levels of dimensionality, (3) provide a graphical interface for simulation setup and control, and (4) provide tools for online visualization and plotting. VCE was designed to provide a distributed, object-oriented environment. Mechanisms are provided for creating and manipulating objects, such as grids, boundary conditions, and solution data. This environment includes parallel virtual machine (PVM) for distributed processing. Users can interactively select and couple any set of codes that have been modified to run in a parallel distributed fashion on a cluster of heterogeneous workstations. A scripting facility allows users to dictate the sequence of events that make up the particular simulation.

Lawrence, Charles

Verification and Validation of Elastodynamic Simulation Software for Aerospace Research

Physics-based simulation of nondestructive evaluation (NDE) inspection can help to advance the inspectability and reliability of mechanical systems. However, NDE simulations applicable to non-idealized mechanical components often require large compute domains and long run times. This has prompted development of custom NDE simulation software tailored to high performance computing (HPC) hardware. Verification and validation (V&V) is an integral part of developing this software to ensure implementations are robust and applicable to inspection problems, producing tools and simulations suitable for computational NDE research. This presentation addresses factors common to V&V of several elastodynamic simulation codes applicable to ultrasonic NDE. Examples are drawn from in-house simulation software at NASA Langley Research Center, ranging from ensuring reliability in a 1D heterogeneous media wave equation solver to the V&V needs of 3D cluster-parallel elastodynamic software. Factors specific to a research environment are addressed, where individual simulation results can be as relevant as the software product itself. Distinct facets of V&V are discussed including testing to establish software reliability, employing systematic approaches for consistency with fundamental conservation laws, establishing the numerical stability of algorithms, and demonstrating concurrence with empirical data. This talk also addresses V&V practices for small groups of researchers. This includes establishing resources (e.g. time and personnel) for V&V during project planning to mitigate and control the risk of setbacks. Similarly, we identify ways for individual researchers to use V&V during simulation software development itself to both speed up the development process and reduce incurred technical debt.

NDE

Scheduling Operations for Massive Heterogeneous Clusters

High-performance computing (HPC) programming has become increasingly difficult with the advent of hybrid supercomputers consisting of multicore CPUs and accelerator boards such as the GPU. Manual tuning of software to achieve high performance on this type of machine has been performed by programmers. This is needlessly difficult and prone to being invalidated by new hardware, new software, or changes in the underlying code. A system was developed for task-based representation of programs, which when coupled with a scheduler and runtime system, allows for many benefits, including higher performance and utilization of computational resources, easier programming and porting, and adaptations of code during runtime. The system consists of a method of representing computer algorithms as a series of data-dependent tasks. The series forms a graph, which can be scheduled for execution on many nodes of a supercomputer efficiently by a computer algorithm. The schedule is executed by a dispatch component, which is tailored to understand all of the hardware types that may be available within the system. The scheduler is informed by a cluster mapping tool, which generates a topology of available resources and their strengths and communication costs. Software is decoupled from its hardware, which aids in porting to future architectures. A computer algorithm schedules all operations, which for systems of high complexity (i.e., most NASA codes), cannot be performed optimally by a human. The system aids in reducing repetitive code, such as communication code, and aids in the reduction of redundant code across projects. It adds new features to code automatically, such as recovering from a lost node or the ability to modify the code while running. In this project, the innovators at the time of this reporting intend to develop two distinct technologies that build upon each other and both of which serve as building blocks for more efficient HPC usage. First is the scheduling and dynamic execution framework, and the second is scalable linear algebra libraries that are built directly on the former.

Humphrey, John

The Cumulus and Stratocumulus CloudSat-CALIPSO Dataset (CASCCAD)

Low clouds continue to contribute greatly to the uncertainty in cloud feedback estimates. Depending on whether a region is dominated by cumulus (Cu) or stratocumulus (Sc) clouds, the interannual low-cloud feedback is somewhat different in both spaceborne and large-eddy simulation studies. Therefore, simulating the correct amount and variation of the Cu and Sc cloud distributions could be crucial to predict future cloud feedbacks. Here we document spatial distributions and profiles of Sc and Cu clouds derived from Cloud-Aerosol Lidar and Infrared Pathfinder Satellite Observations (CALIPSO) and CloudSat measurements. For this purpose, we create a new dataset called the Cumulus And Stratocumulus CloudSat-CALIPSO Dataset (CASCCAD), which identifies Sc, broken Sc, Cu under Sc, Cu with stratiform outflow and Cu. To separate the Cu from Sc, we design an original method based on the cloud height, horizontal extent, vertical variability and horizontal continuity, which is separately applied to both CALIPSO and combined CloudSat–CALIPSO observations. First, the choice of parameters used in the discrimination algorithm is investigated and validated in selected Cu, Sc and Sc–Cu transition case studies. Then, the global statistics are compared against those from existing passive- and active-sensor satellite observations. Our results indicate that the cloud optical thickness – as used in passive-sensor observations – is not a sufficient parameter to discriminate Cu from Sc clouds, in agreement with previous literature. Using clustering-derived datasets shows better results although one cannot completely separate cloud types with such an approach. On the contrary, classifying Cu and Sc clouds and the transition between them based on their geometrical shape and spatial heterogeneity leads to spatial distributions consistent with prior knowledge of these clouds, from ground-based, ship-based and field campaigns. Furthermore, we show that our method improves existing Sc–Cu classifications by using additional information on cloud height and vertical cloud fraction variation. Finally, the CASCCAD datasets provide a basis to evaluate shallow convection and stratocumulus clouds on a global scale in climate models and potentially improve our understanding of low-level cloud feedbacks. The CASCCAD dataset (Cesana, 2019, https://doi.org/10.5281/zenodo.2667637) is available on the Goddard Institute for Space Studies (GISS) website at https://data.giss.nasa.gov/clouds/casccad/ (last access: 5 November 2019) and on the zenodo website at https://zenodo.org/record/2667637 (last access: 5 November 2019).

Cesana, Gregory V.

Enabling Computational Nanotechnology through JavaGenes in a Cycle Scavenging Environment

A genetic algorithm procedure is developed and implemented for fitting parameters for many-body inter-atomic force field functions for simulating nanotechnology atomistic applications using portable Java on cycle-scavenged heterogeneous workstations. Given a physics based analytic functional form for the force field, correlated parameters in a multi-dimensional environment are typically chosen to fit properties given either by experiments and/or by higher accuracy quantum mechanical simulations. The implementation automates this tedious procedure using an evolutionary computing algorithm operating on hundreds of cycle-scavenged computers. As a proof of concept, we demonstrate the procedure for evaluating the Stillinger-Weber (S-W) potential by (a) reproducing the published parameters for Si using S-W energies in the fitness function, and (b) evolving a "new" set of parameters using semi-empirical tightbinding energies in the fitness function. The "new" parameters are significantly better suited for Si cluster energies and forces as compared to even the published S-W potential.

Globus, Al

Dynamic Load-Balancing for Distributed Heterogeneous Computing of Parallel CFD Problems

The developed methodology is aimed at improving the efficiency of executing block-structured algorithms on parallel, distributed, heterogeneous computers. The basic approach of these algorithms is to divide the flow domain into many sub- domains called blocks, and solve the governing equations over these blocks. Dynamic load balancing problem is defined as the efficient distribution of the blocks among the available processors over a period of several hours of computations. In environments with computers of different architecture, operating systems, CPU speed, memory size, load, and network speed, balancing the loads and managing the communication between processors becomes crucial. Load balancing software tools for mutually dependent parallel processes have been created to efficiently utilize an advanced computation environment and algorithms. These tools are dynamic in nature because of the chances in the computer environment during execution time. More recently, these tools were extended to a second operating system: NT. In this paper, the problems associated with this application will be discussed. Also, the developed algorithms were combined with the load sharing capability of LSF to efficiently utilize workstation clusters for parallel computing. Finally, results will be presented on running a NASA based code ADPAC to demonstrate the developed tools for dynamic load balancing.

Ecer, A.

Performance Assessment of OVERFLOW on Distributed Computing Environment

The aerodynamic computer code, OVERFLOW, with a multi-zone overset grid feature, has been parallelized to enhance its performance on distributed and shared memory paradigms. Practical application benchmarks have been set to assess the efficiency of code's parallelism on high-performance architectures. The code's performance has also been experimented with in the context of the distributed computing paradigm on distant computer resources using the Information Power Grid (IPG) toolkit, Globus. Two parallel versions of the code, namely OVERFLOW-MPI and -MLP, have developed around the natural coarse grained parallelism inherent in a multi-zonal domain decomposition paradigm. The algorithm invokes a strategy that forms a number of groups, each consisting of a zone, a cluster of zones and/or a partition of a large zone. Each group can be thought of as a process with one or multithreads assigned to it and that all groups run in parallel. The -MPI version of the code uses explicit message-passing based on the standard MPI library for sending and receiving interzonal boundary data across processors. The -MLP version employs no message-passing paradigm; the boundary data is transferred through the shared memory. The -MPI code is suited for both distributed and shared memory architectures, while the -MLP code can only be used on shared memory platforms. The IPG applications are implemented by the -MPI code using the Globus toolkit. While a computational task is distributed across multiple computer resources, the parallelism can be explored on each resource alone. Performance studies are achieved with some practical aerodynamic problems with complex geometries, consisting of 2.5 up to 33 million grid points and a large number of zonal blocks. The computations were executed primarily on SGI Origin 2000 multiprocessors and on the Cray T3E. OVERFLOW's IPG applications are carried out on NASA homogeneous metacomputing machines located at three sites, Ames, Langley and Glenn. Plans for the future will exploit the distributed parallel computing capability on various homogeneous and heterogeneous resources and large scale benchmarks. Alternative IPG toolkits will be used along with sophisticated zonal grouping strategies to minimize the communication time across the computer resources.

Djomehri, M. Jahed

Distributed Prognostics and Health Management with a Wireless Network Architecture

A heterogeneous set of system components monitored by a varied suite of sensors and a particle-filtering (PF) framework, with the power and the flexibility to adapt to the different diagnostic and prognostic needs, has been developed. Both the diagnostic and prognostic tasks are formulated as a particle-filtering problem in order to explicitly represent and manage uncertainties in state estimation and remaining life estimation. Current state-of-the-art prognostic health management (PHM) systems are mostly centralized in nature, where all the processing is reliant on a single processor. This can lead to a loss in functionality in case of a crash of the central processor or monitor. Furthermore, with increases in the volume of sensor data as well as the complexity of algorithms, traditional centralized systems become for a number of reasons somewhat ungainly for successful deployment, and efficient distributed architectures can be more beneficial. The distributed health management architecture is comprised of a network of smart sensor devices. These devices monitor the health of various subsystems or modules. They perform diagnostics operations and trigger prognostics operations based on user-defined thresholds and rules. The sensor devices, called computing elements (CEs), consist of a sensor, or set of sensors, and a communication device (i.e., a wireless transceiver beside an embedded processing element). The CE runs in either a diagnostic or prognostic operating mode. The diagnostic mode is the default mode where a CE monitors a given subsystem or component through a low-weight diagnostic algorithm. If a CE detects a critical condition during monitoring, it raises a flag. Depending on availability of resources, a networked local cluster of CEs is formed that then carries out prognostics and fault mitigation by efficient distribution of the tasks. It should be noted that the CEs are expected not to suspend their previous tasks in the prognostic mode. When the prognostics task is over, and after appropriate actions have been taken, all CEs return to their original default configuration. Wireless technology-based implementation would ensure more flexibility in terms of sensor placement. It would also allow more sensors to be deployed because the overhead related to weights of wired systems is not present. Distributed architectures are furthermore generally robust with regard to recovery from node failures.

Goebel, Kai