Search NASA⌕ Search

SEARCH · Search NASA

Results for “Parallel algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,063 records · Page 59

Enhancements to Program LAURA for computation of three-dimensional hypersonic flow

Changes to Program Laura (Langley Aerothermodynamic Upwind Relaxation Algorithm) are presented which enhance both stability and accuracy of the algorithm. A discussion of iteration/sweeping strategies and their relation to computer architectures is included to best exploit the capabilities of serial, vector, and parallel processor machines. Test cases for Mach 10 perfect gas flow and Mach 32 real gas flow in chemical nonequilibrium over a blunt, raked elliptic cone using the thin-layer Navier-Stokes equations are presented in order to demonstrate the current improved capabilities. Algorithm changes include the use of volume averaging, application of a symmetric total variation diminishing (TVD) scheme, and stronger interaction between the grid/shock alignment routine and the relaxation algorithm. Good comparisons with heat transfer and pitching moment data at three different angles of attack for the Mach 10 tests serve to further validate the present algorithm. Parameters are defined which control the coupling of the specie continuity equations with the solution of the mixture conservation equations. A discussion of the consequences involved in the choice of strong versus weak coupling is presented, and a sample nonequilibrium calculation on a fine grid over a full scale model of the Aeroassist Flight Experiment (AFE) demonstrates current capabilities.

Gnoffo, Peter A.↗

Maximizing TDRS Command Load Lifetime

The GNC software onboard ISS utilizes TORS command loads, and a simplistic model of TORS orbital motion to generate onboard TORS state vectors. Each TORS command load contains five "invariant" orbital elements which serve as inputs to the onboard propagation algorithm. These elements include semi-major axis, inclination, time of last ascending node crossing, right ascension of ascending node, and mean motion. Running parallel to the onboard software is the TORS Command Builder Tool application, located in the JSC Mission Control Center. The TORS Command Builder Tool is responsible for building the TORS command loads using a ground TORS state vector, mirroring the onboard propagation algorithm, and assessing the fidelity of current TORS command loads onboard ISS. The tool works by extracting a ground state vector at a given time from a current TORS ephemeris, and then calculating the corresponding "onboard" TORS state vector at the same time using the current onboard TORS command load. The tool then performs a comparison between these two vectors and displays the relative differences in the command builder tool GUI. If the RSS position difference between these two vectors exceeds the tolerable lim its, a new command load is built using the ground state vector and uplinked to ISS. A command load's lifetime is therefore defined as the time from when a command load is built to the time the RSS position difference exceeds the tolerable limit. From the outset of TORS command load operations (STS-98), command load lifetime was limited to approximately one week due to the simplicity of both the onboard propagation algorithm, and the algorithm used by the command builder tool to generate the invariant orbital elements. It was soon desired to extend command load lifetime in order to minimize potential risk due to frequent ISS commanding. Initial studies indicated that command load lifetime was most sensitive to changes in mean motion. Finding a suitable value for mean motion was therefore the key to achieving this goal. This goal was eventually realized through development of an Excel spreadsheet tool called EMMIE (Excel Mean Motion Interactive Estimation). EMMIE utilizes ground ephemeris nodal data to perform a least-squares fit to inferred mean anomaly as a function of time, thus generating an initial estimate for mean motion. This mean motion in turn drives a plot of estimated downtrack position difference versus time. The user can then manually iterate the mean motion, and determine an optimal value that will maximize command load lifetime. Once this optimal value is determined, the mean motion initially calculated by the command builder tool is overwritten with the new optimal value, and the command load is built for uplink to ISS. EMMIE also provides the capability for command load lifetime to be tracked through multiple TORS ephemeris updates. Using EMMIE, TORS command load lifetimes of approximately 30 days have been achieved.

Brown, Aaron J.↗

Real time identification of large space structures

Identification of frequencies, damping ratios, and mode shapes of large space structures (LSSs) are examined in real time. Real time processing allows for quick updates of model processing after a reconfiguration of structural failure. Recursive lattice least squares (RLLS) was selected as the baseline algorithm for the identification. Simulation results on a one dimensional LSS demonstrated that it provides good estimates, was not ill-conditioned in the presence of under-excited modes, allowed activity by a supervisory control system which prevented damage to the LSS or excessive drift, and was capable of real-time processing for typical LSS models. A suboptimal version of RLLS, which is equivalent to simulated parallel processing, was derived. A NASTRAN model of the dual keel U.S. space station was used to demonstrate the input/identification algorithm package in a more realistic simulation. Because the first eight flexible modes were very close together, the identification was much more difficult than in the simple examples. Even so, the model was accurately identified in real time.

Voss, Janice E.↗

Modal identification using single-mode projection filters and comparison with ERA and MLE results

The Single-Mode Projection Filter (SPF) is a newly developed algorithm for eigensystem parameter identification from both analytical results and test data. The SPF is formulated with a single mode only and practical for parallel processing implementation. Explicit formulations of SPF are derived for the multi-input multi-output (MIMO) system by using the orthogonal matrices of the controllability and observability matrices in the general sense. The modal parameters of SPF are initially obtained from an analytical model in modal space. The experimental data are then processed through SPF to update its modal parameters and to minimize a cost function defined by the norm of an error matrix. The updated modal parameters represent the characteristics of the test data. A two-dimensional global minimum optimization algorithm is developed and applied for the filter update by using the interval analysis method. The SPF is developed based on a single-mode subsystem and identifies only one modal frequency and one modal damping within a specified region. For an n-modes structure, n SPF can be implemented for parallel processing to reduce the computational burden. The SPF is applied to analyze the simulated data for the MAST beam structure. The estimated modal parameters are comparable to those from the Eigensystem Realization Algorithm (ERA) and repeated modal frequencies are identified. The modal analysis of the Spacecraft Control Laboratory Experiment (SCOLE) data is also performed by using the ERA and the Maximum Likelihood Estimate (MLE). The result shows that the first five modal frequencies are very close from ERA and MLE. However, there are slight disparities in the damping rates and the computational burdens are quite different among these two algorithms.

Huang, Jen-Kuang↗

Evaluation of a dual processor implementation for a fault inferring nonlinear detection system

The design of a modified fault inferring nonlinear detection system (FINDS) algorithm for a dual-processor configured flight computer is described. The algorithm was changed in order to divide it into its translational dynamics and rotational kinematics and to use it for parallel execution on the flight computer. The FINDS consists of: (1) a no-fail filter (NFF), (2) a set of test-of-mean detection tests, (3) a bank of first order filters to estimate failure levels in individual sensors, and (4) a decision function. NFF filter performance using flight recorded sensor data is analyzed using a filter autoinitialization routine. The failure detection and isolation capability of the partitioned algorithm is evaluated. A multirate implementation for the bias-free and bias filter gain and covariance matrices is discussed.

Godiwala, P. M.↗

Implementation and Characterization of Three-Dimensional Particle-in-Cell Codes on Multiple-Instruction-Multiple-Data Massively Parallel Supercomputers

A three-dimensional electrostatic particle-in-cell (PIC) plasma simulation code has been developed on coarse-grain distributed-memory massively parallel computers with message passing communications. Our implementation is the generalization to three-dimensions of the general concurrent particle-in-cell (GCPIC) algorithm. In the GCPIC algorithm, the particle computation is divided among the processors using a domain decomposition of the simulation domain. In a three-dimensional simulation, the domain can be partitioned into one-, two-, or three-dimensional subdomains ("slabs," "rods," or "cubes") and we investigate the efficiency of the parallel implementation of the push for all three choices. The present implementation runs on the Intel Touchstone Delta machine at Caltech; a multiple-instruction-multiple-data (MIMD) parallel computer with 512 nodes. We find that the parallel efficiency of the push is very high, with the ratio of communication to computation time in the range 0.3%-10.0%. The highest efficiency (> 99%) occurs for a large, scaled problem with 64(sup 3) particles per processing node (approximately 134 million particles of 512 nodes) which has a push time of about 250 ns per particle per time step. We have also developed expressions for the timing of the code which are a function of both code parameters (number of grid points, particles, etc.) and machine-dependent parameters (effective FLOP rate, and the effective interprocessor bandwidths for the communication of particles and grid points). These expressions can be used to estimate the performance of scaled problems--including those with inhomogeneous plasmas--to other parallel machines once the machine-dependent parameters are known.

Lyster, P. M.↗

I/O Parallelization for the Goddard Earth Observing System Data Assimilation System (GEOS DAS)

The National Aeronautics and Space Administration (NASA) Data Assimilation Office (DAO) at the Goddard Space Flight Center (GSFC) has developed the GEOS DAS, a data assimilation system that provides production support for NASA missions and will support NASA's Earth Observing System (EOS) in the coming years. The GEOS DAS will be used to provide background fields of meteorological quantities to EOS satellite instrument teams for use in their data algorithms as well as providing assimilated data sets for climate studies on decadal time scales. The DAO has been involved in prototyping parallel implementations of the GEOS DAS for a number of years and is now embarking on an effort to convert the production version from shared-memory parallelism to distributed-memory parallelism using the portable Message-Passing Interface (MPI). The GEOS DAS consists of two main components, an atmospheric General Circulation Model (GCM) and a Physical-space Statistical Analysis System (PSAS). The GCM operates on data that are stored on a regular grid while PSAS works with observational data that are scattered irregularly throughout the atmosphere. As a result, the two components have different data decompositions. The GCM is decomposed horizontally as a checkerboard with all vertical levels of each box existing on the same processing element(PE). The dynamical core of the GCM can also operate on a rotated grid, which requires communication-intensive grid transformations during GCM integration. PSAS groups observations on PEs in a more irregular and dynamic fashion.

Lucchesi, Rob↗

Analytical study of mixing and reacting three-dimensional supersonic combustor flow fields

An analytical investigation is presented of mixing and reacting hydrogen jets injected from multiple orifices transverse and parallel to a supersonic airstream. The COMOC computer program, based upon a finite-element solution algorithm, was developed to solve the governing equations for three-dimensional, turbulent, reacting, boundary-region, and confined flow fields. The computational results provide a three-dimensional description of the velocity, temperature, and species-concentration fields downstream of hydrogen injection. Detailed comparisons between cold-flow data and results of the computational analysis have established validity of the turbulent-mixing model based on the elementary mixing-length hypothesis. A method is established to initiate computations for reacting flow fields based upon cold-flow correlations and the appropriate experimental parameters of Mach number, injector spacing, and pressure ratio. Key analytical observations on mixing and combustion efficiency for reacting flows are presented and discussed.

Baker, A. J.↗

Evaluation of ERTS multispectral signatures in relation to ground control signatures using a nested-sampling approach

The author has identified the following significant results. Ground measured spectral signatures of wavelength bands matching ERTS MSS were collected using a radiometer at several Californian and Nevadan sites, and directly compared with similar data from ERTS CCTs. The comparison was tested at the highest possible spatial resolution for ERTS, using deconvoluted MSS data, and contrasted with that of ground measured spectra, originally from 1 meter squares. In the mobile traverses of the grassland sites, these one meter fields of view were integrated into eighty meter transects along the five km track across four major rock/soil types. Suitable software was developed to read the MSS CCT tapes, to shadeprint individual bands with user-determined greyscale stretching. Four new algorithms for unsupervised and supervised, normalized and unnormalized clustering were developed, into a program termed STANSORT. Parallel software allowed the field data to be calibrated, and by using concurrently continuously collected, upward- and downward-viewing, 4 band radiometers, bidirectional reflectances could be calculated.

Lyon, R. J. P.↗

Advanced detection, isolation, and accommodation of sensor failures in turbofan engines: Real-time microcomputer implementation

The objective of the Advanced Detection, Isolation, and Accommodation Program is to improve the overall demonstrated reliability of digital electronic control systems for turbine engines. For this purpose, an algorithm was developed which detects, isolates, and accommodates sensor failures by using analytical redundancy. The performance of this algorithm was evaluated on a real time engine simulation and was demonstrated on a full scale F100 turbofan engine. The real time implementation of the algorithm is described. The implementation used state-of-the-art microprocessor hardware and software, including parallel processing and high order language programming.

Delaat, John C.↗

Advanced techniques in reliability model representation and solution

The current tendency of flight control system designs is towards increased integration of applications and increased distribution of computational elements. The reliability analysis of such systems is difficult because subsystem interactions are increasingly interdependent. Researchers at NASA Langley Research Center have been working for several years to extend the capability of Markov modeling techniques to address these problems. This effort has been focused in the areas of increased model abstraction and increased computational capability. The reliability model generator (RMG) is a software tool that uses as input a graphical object-oriented block diagram of the system. RMG uses a failure-effects algorithm to produce the reliability model from the graphical description. The ASSURE software tool is a parallel processing program that uses the semi-Markov unreliability range evaluator (SURE) solution technique and the abstract semi-Markov specification interface to the SURE tool (ASSIST) modeling language. A failure modes-effects simulation is used by ASSURE. These tools were used to analyze a significant portion of a complex flight control system. The successful combination of the power of graphical representation, automated model generation, and parallel computation leads to the conclusion that distributed fault-tolerant system architectures can now be analyzed.

Palumbo, Daniel L.↗

Cascaded VLSI neural network architecture for on-line learning

High-speed, analog, fully-parallel, and asynchronous building blocks are cascaded for larger sizes and enhanced resolution. A hardware compatible algorithm permits hardware-in-the-loop learning despite limited weight resolution. A computation intensive feature classification application was demonstrated with this flexible hardware and new algorithm at high speed. This result indicates that these building block chips can be embedded as an application specific coprocessor for solving real world problems at extremely high data rates.

Thakoor, Anilkumar P.↗

Algorithmic trends in computational fluid dynamics; The Institute for Computer Applications in Science and Engineering (ICASE)/LaRC Workshop, NASA Langley Research Center, Hampton, VA, US, Sep. 15-17, 1991

The purpose here is to assess the state of the art in the areas of numerical analysis that are particularly relevant to computational fluid dynamics (CFD), to identify promising new developments in various areas of numerical analysis that will impact CFD, and to establish a long-term perspective focusing on opportunities and needs. Overviews are given of discretization schemes, computational fluid dynamics, algorithmic trends in CFD for aerospace flow field calculations, simulation of compressible viscous flow, and massively parallel computation. Also discussed are accerelation methods, spectral and high-order methods, multi-resolution and subcell resolution schemes, and inherently multidimensional schemes.

Hussaini, M. Y.↗

Daily and 3-hourly Variability in Global Fire Emissions and Consequences for Atmospheric Model Predictions of Carbon Monoxide

Attribution of the causes of atmospheric trace gas and aerosol variability often requires the use of high resolution time series of anthropogenic and natural emissions inventories. Here we developed an approach for representing synoptic- and diurnal-scale temporal variability in fire emissions for the Global Fire Emissions Database version 3 (GFED3). We disaggregated monthly GFED3 emissions during 2003.2009 to a daily time step using Moderate Resolution Imaging Spectroradiometer (MODIS) ]derived measurements of active fires from Terra and Aqua satellites. In parallel, mean diurnal cycles were constructed from Geostationary Operational Environmental Satellite (GOES) Wildfire Automated Biomass Burning Algorithm (WF_ABBA) active fire observations. Daily variability in fires varied considerably across different biomes, with short but intense periods of daily emissions in boreal ecosystems and lower intensity (but more continuous) periods of burning in savannas. These patterns were consistent with earlier field and modeling work characterizing fire behavior dynamics in different ecosystems. On diurnal timescales, our analysis of the GOES WF_ABBA active fires indicated that fires in savannas, grasslands, and croplands occurred earlier in the day as compared to fires in nearby forests. Comparison with Total Carbon Column Observing Network (TCCON) and Measurements of Pollution in the Troposphere (MOPITT) column CO observations provided evidence that including daily variability in emissions moderately improved atmospheric model simulations, particularly during the fire season and near regions with high levels of biomass burning. The high temporal resolution estimates of fire emissions developed here may ultimately reduce uncertainties related to fire contributions to atmospheric trace gases and aerosols. Important future directions include reconciling top ]down and bottom up estimates of fire radiative power and integrating burned area and active fire time series from multiple satellite sensors to improve daily emissions estimates.

Mu, M.↗

Northern Hemisphere Surface Freeze-Thaw Product from Aquarius L-Band Radiometers

In the Northern Hemisphere, seasonal changes in surface freeze–thaw (FT) cycles are an important component of surface energy, hydrological and eco-biogeochemical processes that must be accurately monitored. This paper presents the weekly polar-gridded Aquarius passive L-band surface freeze–thaw product (FT-AP) distributed on the Equal-Area Scalable Earth Grid version 2.0, above the parallel 50 N, with a spatial resolution of 36 km36 km. The FT-AP classification algorithm is based on a seasonal threshold approach using the normalized polarization ratio, references for frozen and thawed conditions and optimized thresholds. To evaluate the uncertainties of the product, we compared it with another satellite FT product also derived from passive microwave observations but at higher frequency: the resampled 37 GHz FT Earth Science Data Record (FTESDR). The assessment was carried out during the overlapping period between 2011 and 2014. Results show that 77.1% of their common grid cells have an agreement better than 80 %. Their differences vary with land cover type (tundra, forest and open land) and freezing and thawing periods. The best agreement is obtained during the thawing transition and over forest areas, with differences between product mean freeze or thaw onsets of under 0.4 weeks. Over tundra, FT-AP tends to detect freeze onset 2–5 weeks earlier than FT-ESDR, likely due to FT sensitivity to the different frequencies used. Analysis with mean surface air temperature time series from six in situ meteorological stations shows that the main discrepancies between FT-AP and FT-ESDR are related to false frozen retrievals in summer for some regions with FT-AP. The Aquarius product is distributed by the U.S. National Snow and Ice Data Center (NSIDC) at https://nsidc.org/data/aq3_ft/versions/5 with the DOI https://doi.org/10.5067/OV4R18NL3BQR.

Prince, Michael↗

Parallel interior-point solver for block-structured nonlinear programs on SIMD/GPU architectures

Here, we investigate how to port the standard interior-point method to new exascale architectures for block-structured nonlinear programs with state equations. Computationally, we decompose the interior-point algorithm into two successive operations: the evaluation of the derivatives and the solution of the associated Karush-Kuhn-Tucker (KKT) linear system. Our method accelerates both operations using two levels of parallelism. First, we distribute the computations on multiple processes using coarse parallelism. Second, each process uses SIMD/GPU accelerators locally to accelerate the operations using fine-grained parallelism. The KKT system is reduced by eliminating the inequalities and the state variables from the corresponding equations. We demonstrate our method's capability on the supercomputer Polaris, a testbed for the future exascale Aurora system. Each node is equipped with four GPUs, a setup amenable to our two-level approach. Our experiments on the stochastic optimal power flow problem show that the reduction method is 50x faster than the sparse linear solver HSL MA57 running in serial on the CPU, and 6x faster than Pardiso running in parallel on CPU on the same number of processes.

97 MATHEMATICS AND COMPUTING↗

Impacts of floating-point non-associativity on reproducibility for HPC and deep learning applications

Run to run variability in parallel programs caused by floating-point non-associativity has been known to significantly affect reproducibility in iterative algorithms, due to accumulating errors. Non-reproducibility can critically affect the efficiency and effectiveness of correctness testing for stochastic programs. Recently, the sensitivity of deep learning training and inference pipelines to floating-point non-associativity has been found to sometimes be extreme. It can prevent certification for commercial applications, accurate assessment of robustness and sensitivity, and bug detection. New approaches in scientific computing applications have coupled deep learning models with high-performance computing, leading to an aggravation of debugging and testing challenges. Here we perform an investigation of the statistical properties of floating-point non-associativity within modern parallel programming models, and analyze performance and productivity impacts of replacing atomic operations with deterministic alternatives on GPUs. We examine the recently-added deterministic options in PyTorch within the context of GPU deployment for deep learning, uncovering and quantifying the impacts of input parameters triggering run to run variability and reporting on the reliability and completeness of the documentation. Finally, we evaluate the strategy of exploiting automatic determinism that could be provided by deterministic hardware, using the Groq LPUTM accelerator for inference portions of the deep learning pipeline. We demonstrate the benefits that a hardware-based strategy can provide within reproducibility and correctness efforts.

Shanmugavelu, Sanjif↗