Search NASA⌕ Search

SEARCH · Search NASA

Results for “distributed parallel computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40

Network issues for large mass storage requirements

File Servers and Supercomputing environments need high performance networks to balance the I/O requirements seen in today's demanding computing scenarios. UltraNet is one solution which permits both high aggregate transfer rates and high task-to-task transfer rates as demonstrated in actual tests. UltraNet provides this capability as both a Server-to-Server and Server-to-Client access network giving the supercomputing center the following advantages highest performance Transport Level connections (to 40 MBytes/sec effective rates); matches the throughput of the emerging high performance disk technologies, such as RAID, parallel head transfer devices and software striping; supports standard network and file system applications using SOCKET's based application program interface such as FTP, rcp, rdump, etc.; supports access to the Network File System (NFS) and LARGE aggregate bandwidth for large NFS usage; provides access to a distributed, hierarchical data server capability using DISCOS UniTree product; supports file server solutions available from multiple vendors, including Cray, Convex, Alliant, FPS, IBM, and others.

Perdue, James↗

Time-domain analysis of planar microstrip devices using a generalized Yee-algorithm based on unstructured grids

The generalized Yee-algorithm is presented for the temporal full-wave analysis of planar microstrip devices. This algorithm has the significant advantage over the traditional Yee-algorithm in that it is based on unstructured and irregular grids. The robustness of the generalized Yee-algorithm is that structures that contain curved conductors or complex three-dimensional geometries can be more accurately, and much more conveniently modeled using standard automatic grid generation techniques. This generalized Yee-algorithm is based on the the time-marching solution of the discrete form of Maxwell's equations in their integral form. To this end, the electric and magnetic fields are discretized over a dual, irregular, and unstructured grid. The primary grid is assumed to be composed of general fitted polyhedra distributed throughout the volume. The secondary grid (or dual grid) is built up of the closed polyhedra whose edges connect the centroid's of adjacent primary cells, penetrating shared faces. Faraday's law and Ampere's law are used to update the fields normal to the primary and secondary grid faces, respectively. Subsequently, a correction scheme is introduced to project the normal fields onto the grid edges. It is shown that this scheme is stable, maintains second-order accuracy, and preserves the divergenceless nature of the flux densities. Finally, for computational efficiency the algorithm is structured as a series of sparse matrix-vector multiplications. Based on this scheme, the generalized Yee-algorithm has been implemented on vector and parallel high performance computers in a highly efficient manner.

Gedney, Stephen D.↗

Structural relations between collagen and mineral in bone as determined by high voltage electron microscopic tomography

Aspects of the ultrastructural interaction between collagen and mineral crystals in embryonic chick bone have been examined by the novel technique of high voltage electron microscopic tomography to obtain three-dimensional information concerning extracellular calcification in this tissue. Newly mineralizing osteoid along periosteal surfaces of mid-diaphyseal regions from normal chick tibiae was embedded, cut into 0.25 microns thick sections, and documented at 1.0 MV in the Albany AEI-EM7 high voltage electron microscope. The areas of the tissue studied contained electron dense mineral crystals associated with collagen fibrils, some marked by crystals disposed along their cylindrically shaped lengths. Tomographic reconstructions of one site with two mineralizing fibrils were computed from a 5 degrees tilt series of micrographs over a +/- 60 degrees range. Reconstructions showed that the mineral crystals were platelets of irregular shape. Their sizes were variable, measured here up to 80 x 30 x 8 nm in length, width, and thickness, respectively. The longest crystal dimension, corresponding to the c-axis crystallographically, was generally parallel to the collagen fibril long axis. Individual crystals were oriented parallel to one another in each fibril examined. They were also parallel in the neighboring but apparently spatially separate fibrils. Crystals were periodically (approximately 67 nm repeat distance) arranged along the fibrils and their location appeared to correspond to collagen hole and overlap zones defined by geometrical imaging techniques. The crystals appeared to be continuously distributed along a fibril, their size and number increasing in a tapered fashion from a relatively narrow tip containing smaller and infrequent crystals to wider regions having more densely packed and larger crystals. Defined for the first time by direct visual 3D imaging, these data describe the size, shape, location, orientation, and development of early crystals in normal bone collagen. The results suggest that platelet-shaped crystals are arranged in channels or grooves which are formed by collagen hole zones in register and that crystal sizes may exceed the dimensions of hole zones. Such data agree with those from mineral-matrix interaction in normally calcifying avian tendon obtained by similar high voltage tomographic means, but in addition they indicate a possible gradual and continuous deposition of crystals in collagen of bone unlike tendon and imply that individual collagen fibrils in local regions of osteoid are organized such that they all may be aligned in a coherent manner.

NASA Discipline Number 40-40↗

DIMPLES: Distributed Influence Maximization for Pandemic pLanning on Exascale Systems

We study exascale parallel algorithms for the selection of intervention or monitoring strategies in massive realistic socio-technical networks through scalable Influence Maximization (InfMax) algorithms. We employ novel techniques to enable efficient scaling on up to 8k nodes of OLCF Frontier, with 65k AMD GPUs and 458k AMD CPU cores. Current state-of-the-art InfMax tools are limited to networks with only a few million actors (vertices) and a few hundred million interactions (edges). By overcoming these limitations, we show that our approach is capable of processing a realistic social contact network of the United States with 285 million nodes and about 8 billion edges. This two orders-of-magnitude improvement over the previous state-of-the-art is obtained by leveraging algorithmic advancements for the InfMax problem and designing several problem-specific approaches to overlap communication with computation, improve GPU efficiency, and lower the application’s memory requirements. We evaluate strong scaling for computing 10k most influential seeds using up to 8k nodes of an exascale system, and weak scaling from 128 to 8k system nodes for seed sets ranging from 625 to 40k seeds. We achieve the fastest-known runtime of 25 minutes while performing 48 million diffusion simulations totaling 2.31 petabytes to identify 40k influential seeds using 8k nodes, and take 5.75 minutes to identify 10k seeds while using 4k nodes.

Minutoli, Marco [Pacific Northwest National Labora↗

In situ multi-tier auto-ignition detection applied to dual-fuel combustion simulations

Here we use an anomaly detection methodology that is centered on analyzing fourth-order joint moments (co-kurtosis), particularly focusing on its application in auto-ignition of combustion problems with large numbers of species. Unsupervised anomaly detection is challenging to generalize across problem types and domains. A recent technique, centered on analyzing information in the fourth-order joint moment co-kurtosis, has shown promise, especially for high-dimensional scientific data. In this work we present developments to the co-kurtosis based anomaly detection method needed to make it effective and scalable for large-scale distributed scientific data, such as those generated by massively parallel simulations. An in situ co-kurtosis algorithm is employed as the anomaly detection method for identifying ignition kernels in simulations of turbulent combustion. Here, we extend an existing methodology which identifies regions of the domain where anomalies are present, and add another tier of anomaly detection where the individual samples contributing to the anomaly are identified. We apply this algorithm on-the-fly to a variety of turbulent reacting flow problems and compare it to the widely used (but significantly more expensive) chemical explosive mode analysis (CEMA). We demonstrate the ability of the method to detect and identify the onset of low and high temperature ignition which can be used for computational steering, as chemical and combustion anomalies occur intermittently at spatio-temporal locations unknown a priori. Finally, we apply our lightweight in situ algorithm to an exascale high-fidelity simulation with a total of 2.4 Trillion degrees of freedom, performed using an adaptive mesh refinement solver. Furthermore, through a scalability analysis, we show that the relative computational cost of this in-situ anomaly detection algorithm compared to an iteration of the reacting flow solver is negligible.

97 MATHEMATICS AND COMPUTING↗

MPI, HPF or OpenMP: A Study with the NAS Benchmarks

Porting applications to new high performance parallel and distributed platforms is a challenging task. Writing parallel code by hand is time consuming and costly, but this task can be simplified by high level languages and would even better be automated by parallelizing tools and compilers. The definition of HPF (High Performance Fortran, based on data parallel model) and OpenMP (based on shared memory parallel model) standards has offered great opportunity in this respect. Both provide simple and clear interfaces to language like FORTRAN and simplify many tedious tasks encountered in writing message passing programs. In our study, we implemented the parallel versions of the NAS Benchmarks with HPF and OpenMP directives. Comparison of their performance with the MPI implementation and pros and cons of different approaches will be discussed along with experience of using computer-aided tools to help parallelize these benchmarks. Based on the study, potentials of applying some of the techniques to realistic aerospace applications will be presented.

Jin, H.↗

MPI, HPF or OpenMP: A Study with the NAS Benchmarks

Porting applications to new high performance parallel and distributed platforms is a challenging task. Writing parallel code by hand is time consuming and costly, but the task can be simplified by high level languages and would even better be automated by parallelizing tools and compilers. The definition of HPF (High Performance Fortran, based on data parallel model) and OpenMP (based on shared memory parallel model) standards has offered great opportunity in this respect. Both provide simple and clear interfaces to language like FORTRAN and simplify many tedious tasks encountered in writing message passing programs. In our study we implemented the parallel versions of the NAS Benchmarks with HPF and OpenMP directives. Comparison of their performance with the MPI implementation and pros and cons of different approaches will be discussed along with experience of using computer-aided tools to help parallelize these benchmarks. Based on the study,potentials of applying some of the techniques to realistic aerospace applications will be presented

Jin, Hao-Qiang↗

Electrical Capacitance Volume Tomography with High-Contrast Dielectrics

The Electrical Capacitance Volume Tomography (ECVT) system has been designed to complement the tools created to sense the presence of water in nonconductive spacecraft materials, by helping to not only find the approximate location of moisture but also its quantity and depth. The ECVT system has been created for use with a new image reconstruction algorithm capable of imaging high-contrast dielectric distributions. Rather than relying solely on mutual capacitance readings as is done in traditional electrical capacitance tomography applications, this method reconstructs high-resolution images using only the self-capacitance measurements. The image reconstruction method assumes that the material under inspection consists of a binary dielectric distribution, with either a high relative dielectric value representing the water or a low dielectric value for the background material. By constraining the unknown dielectric material to one of two values, the inverse math problem that must be solved to generate the image is no longer ill-determined. The image resolution becomes limited only by the accuracy and resolution of the measurement circuitry. Images were reconstructed using this method with both synthetic and real data acquired using an aluminum structure inserted at different positions within the sensing region. The cuboid geometry of the system has two parallel planes of 16 conductors arranged in a 4 4 pattern. The electrode geometry consists of parallel planes of copper conductors, connected through custom-built switch electronics, to a commercially available capacitance to digital converter. The figure shows two 4 4 arrays of electrodes milled from square sections of copper-clad circuit-board material and mounted on two pieces of glass-filled plastic backing, which were cut to approximately square shapes, 10 cm on a side. Each electrode is placed on 2.0-cm centers. The parallel arrays were mounted with the electrode arrays approximately 3 cm apart. The open ends were surrounded by a metal guard to reduce the sensitivity of the electrodes to outside interference and to help maintain the spacing between the arrays. Other uses for this innovation potentially include quantifying the amount of commodity remaining in the fuel and oxidizer tanks while on-orbit without having to fire spacecraft engines. Another orbit application is moisture sensing in plant-growth experiments because microgravity causes moisture in soil to distribute itself in unusual ways. At the moment, the hardware and image reconstruction technique may only be of interest to people involved in nondestructive evaluation. The reconstructed image takes almost a full week to reproduce with existing computer power. However, because computer power and speeds follows Moore s Law, execution times are likely to become acceptable within the next five to eight years. The code was written in Mathematica for dedicated use with the ECVT system. In its present form, it is not suitable to be used directly as a consumer product. However, the code could be likely improved by rewriting it in a compiled language such as C or Fortran.

Nurge, Mark↗

Parallel Runtime Interface for Fortran (PRIF) Specification (Rev. 0.3)

This document specifies an interface to support the parallel features of Fortran, named the Parallel Runtime Interface for Fortran (PRIF). PRIF is a proposed solution in which the runtime library is responsible for coarray allocation, deallocation and accesses, image synchronization, atomic operations, events, and teams. In this interface, the compiler is responsible for transforming the invocation of Fortran-level parallel features into procedure calls to the necessary PRIF procedures. The interface is designed for portability across shared- and distributed-memory machines, different operating systems, and multiple architectures. Implementations of this interface are intended as an augmentation for the compiler's own runtime library. With an implementation-agnostic interface, alternative parallel runtime libraries may be developed that support the same interface. One benefit of this approach is the ability to vary the communication substrate. A central aim of this document is to define a parallel runtime interface in standard Fortran syntax, which enables us to leverage Fortran to succinctly express various properties of the procedure interfaces, including argument attributes.

97 MATHEMATICS AND COMPUTING↗

Parallel Runtime Interface for Fortran (PRIF) Specification (Rev. 0.4)

This document specifies an interface to support the parallel features of Fortran, named the Parallel Runtime Interface for Fortran (PRIF). PRIF is a proposed solution in which the runtime library is responsible for coarray allocation, deallocation and accesses, image synchronization, atomic operations, events, and teams. In this interface, the compiler is responsible for transforming the invocation of Fortran-level parallel features into procedure calls to the necessary PRIF procedures. The interface is designed for portability across shared- and distributed-memory machines, different operating systems, and multiple architectures. Implementations of this interface are intended as an augmentation for the compiler's own runtime library. With an implementation-agnostic interface, alternative parallel runtime libraries may be developed that support the same interface. One benefit of this approach is the ability to vary the communication substrate. A central aim of this document is to define a parallel runtime interface in standard Fortran syntax, which enables us to leverage Fortran to succinctly express various properties of the procedure interfaces, including argument attributes.

97 MATHEMATICS AND COMPUTING↗

ARM-IRL: Adaptive Resilience Metric Quantification Using Inverse Reinforcement Learning

The resilience of safety-critical systems is gaining importance due to the rise in cyber and physical threats, especially within critical infrastructure. Traditional static resilience metrics may not capture dynamic system states, leading to inaccurate assessments and ineffective responses to cyber threats. This work aims to develop a data-driven, adaptive method for resilience metric learning. We propose a data-driven approach using inverse reinforcement learning (IRL) to learn a single, adaptive resilience metric. The method infers a reward function from expert control actions. Unlike previous approaches using static weights or fuzzy logic, this work applies adversarial inverse reinforcement learning (AIRL), training a generator and discriminator in parallel to learn the reward structure and derive an optimal policy. The proposed approach is evaluated on multiple scenarios: optimal communication network rerouting, power distribution network reconfiguration, and cyber–physical restoration of critical loads using the IEEE 123-bus system. The adaptive, learned resilience metric enables faster critical load restoration in comparison to conventional RL approaches.

97 MATHEMATICS AND COMPUTING↗

Sparse distributed memory: Principles and operation

Sparse distributed memory is a generalized random access memory (RAM) for long (1000 bit) binary words. Such words can be written into and read from the memory, and they can also be used to address the memory. The main attribute of the memory is sensitivity to similarity, meaning that a word can be read back not only by giving the original write address but also by giving one close to it as measured by the Hamming distance between addresses. Large memories of this kind are expected to have wide use in speech recognition and scene analysis, in signal detection and verification, and in adaptive control of automated equipment, in general, in dealing with real world information in real time. The memory can be realized as a simple, massively parallel computer. Digital technology has reached a point where building large memories is becoming practical. Major design issues were resolved which were faced in building the memories. The design is described of a prototype memory with 256 bit addresses and from 8 to 128 K locations for 256 bit words. A key aspect of the design is extensive use of dynamic RAM and other standard components.

Flynn, M. J.↗

Coarse-grained simulation of colloidal self-assembly, cation exchange, and rheology in Na/Ca smectite clay gels

Knowledge Gap: The aggregation of clay minerals—layered silicate nanoparticles—strongly impacts fluid flow, solute migration, and solid mechanics in soils, sediments, and sedimentary rocks. Experimental and computational characterization of clay aggregation is inhibited by the delicate water-mediated nature of clay colloidal interactions and by the range of spatial scales involved, from 1 nm thick platelets to flocs with dimensions up to micrometers or more. Simulations: Using a new coarse-grained molecular dynamics (CGMD) approach, we predicted the microstructure, dynamics, and rheology of hydrated smectite (more precisely, montmorillonite) clay gels containing up to 2,000 clay platelets on length scales up to 0.1 μm. Further, simulations investigated the impact of simulation time, platelet diameters (6 to 25nm), and the ratio of Na to Ca exchangeable cations on the assembly of tactoids (i.e., stacks of parallel clay platelets) and larger aggregates (i.e., assemblages of tactoids). We analyzed structural features including tactoid size and size distribution, basal spacing, counterion distribution in the electrical double layer, clay association modes, and the rheological properties of smectite gels. Findings: Our results demonstrate new potential to characterize and understand clay aggregation in dilute suspensions and gels on a scale of thousands of particles with explicit representation of counterion clouds and with accuracy approaching that of all-atom molecular dynamics (MD) simulations. For example, our simulations predict the strong impact of Na/Ca ratio on clay tactoid formation and the shear-thinning rheology of clay gels.

42 ENGINEERING↗

Two-Fluid and Discrete Element Modeling of a Parallel Plate Fluidized Bed Heat Exchanger for Concentrating Solar Power

A novel high-temperature particle solar receiver is developed using a light trapping planar cavity configuration. As particles fall through the cavity, the concentrated solar radiation warms the boundaries of the receiver and in turn heats the particles. Particles flow through the system, forming a fluidized bed at the lower section, leaving the system from the bottom at a constant flowrate. Air is introduced to the system as the fluidizing medium to improve particle heat transfer and mixing. A laboratory scale cavity receiver is built by collaborators at the Colorado School of Mines and their data are used for model validation. In this experimental setup, near IR quartz lamp is used to provide flux to the vertical wall of the heat exchanger. The system is modeled using the discrete element method and a continuum two-fluid method. The computational model matches the experimental system size and the particle size distribution is assumed monodisperse. A new continuum conduction model that accounts for the effects of solid concentration is implemented, and the heat flux boundary condition matches the experimental setup. Radiative heat transfer is estimated using a widely used correlation during the post-processing step to determine an overall heat transfer coefficient. The model is validated against testing data and achieves less than 30% discrepancy and a heat transfer coefficient greater than 1000 W/m2 K.

concentrating solar power↗

Comparative Results of Tests on Several Different Types of Nozzles

This paper presents the results of tests conducted to determine the effect of the constructional elements of a Laval nozzle on the velocity and pressure distribution and the magnitude of the reaction force of the jet. The effect was studied of the shapes of the entrance section of the nozzle and three types of divergent sections: namely, straight cone, conoidal with cylindrical and piece and diffuser obtained computationally by a graphical method due to Professor F. I. Frankle. The effect of the divergence angle of the nozzle on the jet reaction was also investigated. The results of the investigation showed that the shape of the generator of the inner surface of the entrance part of the nozzle essentially has no effect on the character of the flow and on the reaction. The nozzle that was obtained by graphical computation assured the possibility of obtaining a flow for which the velocity of all the gas particles is parallel to the axis of symmetry of the nozzle, the reaction being on the average 2 to 3 percent greater than for the usual conical nozzle under the same conditions, For the conical nozzle the maximum reaction was obtained for a cone angle of 25deg to 27deg. At the end of this paper a sample computation is given by the graphical method. The tests were started at the beginning of 1936 and this paper was written at the same time.

Kisenko, M. S.↗

Applying Orbital Multi-Angle Photopolarimetric Observations to Study Properties of Aerosols in the Earth's Atmosphere: Implications of Measurements in the 1.378 µm Spectral Channel to Retrieve Microphysical Characteristics and Composition of Stratospheric Aerosols

We analyze the possibilities of orbital photopolarimetric measurements to study properties of aerosols in the Earth's atmosphere. As an example, we consider the case when such measurements are performed within a narrow spectral channel centered at 1.378 µm that allows to retrieve microphysical characteristics of stratospheric aerosols separately from those of tropospheric aerosols. We consider the case of stratospheric aerosols caused by volcanic eruption, and adopt the model of the stratosphere in the form of a homogeneous plane-parallel layer composed of polydisperse spherical particles. We use numerically exact solutions of the vector radiative transfer equation to theoretically simulate measurements carried out at various numbers of scattering angles, including: (i) radiance measurements alone; (ii) polarization measurements alone; and (iii) radiance and polarization measurements together. The results of computations show that the simultaneous use of radiance and polarization measurements at a sufficiently large number of scattering angles enables one to retrieve the optical thickness, effective radius, and refractive index of aerosols with adequate accuracy. We demonstrate how the accuracy of the derived values of the optical parameters of aerosols depends on the accuracy of measurements of the intensity and polarization of the reflected light, optical thickness of aerosol layer itself, effective radius of aerosols, width of the particle size distribution, and number of viewing angles.

Aerosols↗

Elevating SolTrace's Capabilities for the Next Generation of Concentrating Solar Analysis

SolTrace is an open-source Monte Carlo ray tracing software developed at NREL. SolTrace can characterize concentrating solar thermal (CST) collector optical performance and is CST technology agnostic. Shown in Fig. 1, SolTrace is a foundational tool in NREL's CST system and component modeling suite. SolTrace's generic surface elements can flexibly model novel collector and receiver designs to predict spatial and temporal flux distributions - critical to understand for CST component design, performance prediction, and system integration. Since its initial development, SolTrace has over 1,650 references on Google Scholar, over 9,800 downloads since 2017, and has served the CST research and development community as a benchmark of 3rd party verification. SolTrace provides users with many options for defining surface shape and boundaries. However, SolTrace provides limited documentation which can result in a steep learning curve for new users. Additionally, SolTrace lacks the computational performance required to evaluate optical performance of a CST system over the course of a year and/or iteratively over design parameters in a timely manner. To address this, we are working towards a new release of SolTrace that enables increased computational throughput by implementing ray tracing acceleration structures and enabling GPU parallelization. Additionally, we are working to improve SolTrace's usability, accessibility, and maintainability by (1) automating solar position time-dependent simulation processes, (2) creating general CST collector templates of grouped elements, (3) updating the user interface to better visualize model inputs and outputs, and (4) creating a user support network through forums, "how to" videos, and documentation.

14 SOLAR ENERGY↗

On the Sensitivity of Piezoceramics and Piezopolymers in Structural Integrity Monitoring of Large Trusses

An analytical assessment has been made of the reliability of using integrated microactuators and sensors in the form of piezoceramics and piezopolymers as joint integrity monitors in trussed systems. The concept is first implemented for a simple structure which consists of two truss members with a 45 deg lift angle joined at the apex. A piezoceramic patch (or piezopolymer film) bonded on the surface of one of the members at a location near the joint is used as a collocated actuator/sensor. The overall structural dynamic response under an excitation was modeled by finite element method. Different degrees of nodal constraints at the joints representing various degrees of joint integrity are employed. The resulting dynamic response showed distinct responses for varying joint stiffnesses. Parallel experimental work on a truss model using a multichannel data acquisition system and a digital signal analyzer confirms the results from analysis. We further studied the sensitivity of the micro-sensors to the behavior of joints of large arch truss structure. Results obtained for large trusses with many degrees of freedom indicate optimum locations of sensors for which the dynamic response signatures are distinct and distinguishable for relatively small changes in joint integrity and/or structural geometry. Computations based on finite element modeling show that locating the single actuator/sensor at the joint corresponding to the first loss of static stability appear optimal. Hence, static stability analysis of complex trusses can give us a good indication of the optimum placement of sensors for maximum response. This observation is important if few distributed sensors and actuators are available for placement in constructed facilities made from large trusses with many degrees of freedom. As an extension of this work a dynamic response signature identification technique to monitor in-service degradation of joints is under development for application to the monitoring of the integrity of adhesive joints in composite structures.

Abatan, A. O.↗