Search NASA⌕ Search

SEARCH · Search NASA

Results for “Performance Tuning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Constructing Space-Time Views from Fixed Size Statistical Data: Getting the Best of Both Worlds

Many performance monitoring tools are currently available to the super-computing community. The performance data gathered and analyzed by these tools fall under two categories: statistics and event traces. Statistical data is much more compact but lacks the probative power event traces offer. Event traces, on the other hand, can easily fill up the entire file system during execution such that the instrumented execution may have to be terminated half way through. In this paper, we propose an innovative methodology for performance data gathering and representation that offers a middle ground. The user can trade-off tracing overhead, trace data size vs. data quality incrementally. In other words, the user will be able to limit the amount of trace collected and, at the same time, carry out some of the analysis event traces offer using spacetime views for the entire execution. Two basic ideas are employed: the use of averages to replace recording data for each instance and "formulae" to represent sequences associated with communication and control flow. With the help of a few simple examples, we illustrate the use of these techniques in performance tuning and compare the quality of the traces we collected vs. event traces. We found that the trace files thus obtained are, in deed, small, bounded and predictable before program execution and that the quality of the space time views generated from these statistical data are excellent. Furthermore, experimental results showed that the formulae proposed were able to capture 100% of all the sequences associated with 11 of the 15 applications tested. The performance of the formulae can be incrementally improved by allocating more memory at run-time to learn longer sequences.

Schmidt, Melisa↗

Constructing Space-Time Views from Fixed Size Statistical Data: Getting the Best of both Worlds

Many performance monitoring tools are currently available to the super-computing community. The performance data gathered and analyzed by these tools fall under two categories: statistics and event traces. Statistical data is much more compact but lacks the probative power event traces offer. Event traces, on the other hand, can easily fill up the entire file system during execution such that the instrumented execution may have to be terminated half way through. In this paper, we propose an innovative methodology for performance data gathering and representation that offers a middle ground. The user can trade-off tracing overhead, trace data size vs. data quality incrementally. In other words, the user will be able to limit the amount of trace collected and, at the same time, carry out some of the analysis event traces offer using space-time views for the entire execution. Two basic ideas arc employed: the use of averages to replace recording data for each instance and formulae to represent sequences associated with communication and control flow. With the help of a few simple examples, we illustrate the use of these techniques in performance tuning and compare the quality of the traces we collected vs. event traces. We found that the trace files thus obtained are, in deed, small, bounded and predictable before program execution and that the quality of the space time views generated from these statistical data are excellent. Furthermore, experimental results showed that the formulae proposed were able to capture 100% of all the sequences associated with 11 of the 15 applications tested. The performance of the formulae can be incrementally improved by allocating more memory at run-time to learn longer sequences.

Schmidt, Melisa↗

Mobile Thread Task Manager

The Mobile Thread Task Manager (MTTM) is being applied to parallelizing existing flight software to understand the benefits and to develop new techniques and architectural concepts for adapting software to multicore architectures. It allocates and load-balances tasks for a group of threads that migrate across processors to improve cache performance. In order to balance-load across threads, the MTTM augments a basic map-reduce strategy to draw jobs from a global queue. In a multicore processor, memory may be "homed" to the cache of a specific processor and must be accessed from that processor. The MTTB architecture wraps access to data with thread management to move threads to the home processor for that data so that the computation follows the data in an attempt to avoid L2 cache misses. Cache homing is also handled by a memory manager that translates identifiers to processor IDs where the data will be homed (according to rules defined by the user). The user can also specify the number of threads and processors separately, which is important for tuning performance for different patterns of computation and memory access. MTTM efficiently processes tasks in parallel on a multiprocessor computer. It also provides an interface to make it easier to adapt existing software to a multiprocessor environment.

Clement, Bradley J.↗

Integrating Cache Performance Modeling and Tuning Support in Parallelization Tools

With the resurgence of distributed shared memory (DSM) systems based on cache-coherent Non Uniform Memory Access (ccNUMA) architectures and increasing disparity between memory and processors speeds, data locality overheads are becoming the greatest bottlenecks in the way of realizing potential high performance of these systems. While parallelization tools and compilers facilitate the users in porting their sequential applications to a DSM system, a lot of time and effort is needed to tune the memory performance of these applications to achieve reasonable speedup. In this paper, we show that integrating cache performance modeling and tuning support within a parallelization environment can alleviate this problem. The Cache Performance Modeling and Prediction Tool (CPMP), employs trace-driven simulation techniques without the overhead of generating and managing detailed address traces. CPMP predicts the cache performance impact of source code level "what-if" modifications in a program to assist a user in the tuning process. CPMP is built on top of a customized version of the Computer Aided Parallelization Tools (CAPTools) environment. Finally, we demonstrate how CPMP can be applied to tune a real Computational Fluid Dynamics (CFD) application.

Waheed, Abdul↗

Programming Tools: Status, Evaluation, and Comparison

In this tutorial I will first describe the characteristics of scientific applications and their developers, and describe the computing environment in a typical high-performance computing center. I will define the user requirements for tools that support application portability and present the difficulties to satisfy them. These form the basis of the evaluation and comparison of the tools. I will then describe the tools available in the market and the tools available in the public domain. Specifically, I will describe the tools for converting sequential programs, tools for developing portable new programs, tools for debugging and performance tuning, tools for partitioning and mapping, and tools for managing network of resources. I will introduce the main goals and approaches of the tools, and show main features of a few tools in each category. Meanwhile, I will compare tool usability for real-world application development and compare their different technological approaches. Finally, I will indicate the future directions of the tools in each category.

Cheng, Doreen Y.↗

Are Event Traces Really That Necessary?

Many performance monitoring tools are currently available to the super-computing community. The performance data gathered and analyzed by these tools fall under two categories: statistics and event traces. Statistical data is much more compact but lack the probative power event traces offer. Event traces, on the other hand, can easily fill up the entire file system during execution such that the instrumented execution have to be terminated. In this paper, we propose an innovative methodology for monitoring and trace representation that offers a middle ground. The user can trace-off trace data size vs. quality incrementally. Specifically, the user will be able to limit the amount of trace collected and, at the same time, carry out some of the analysis event traces offer for the entire execution. With the help of a few CFD examples, we illustrate the use of our technique in performance tuning. We also compare quantitatively, the quality of the traces we collected vs. event traces.

Schmidt, Melisa↗

The Automated Instrumentation and Monitoring System (AIMS): Design and Architecture

Whether a researcher is designing the 'next parallel programming paradigm', another 'scalable multiprocessor' or investigating resource allocation algorithms for multiprocessors, a facility that enables parallel program execution to be captured and displayed is invaluable. Careful analysis of such information can help computer and software architects to capture, and therefore, exploit behavioral variations among/within various parallel programs to take advantage of specific hardware characteristics. A software tool-set that facilitates performance evaluation of parallel applications on multiprocessors has been put together at NASA Ames Research Center under the sponsorship of NASA's High Performance Computing and Communications Program over the past five years. The Automated Instrumentation and Monitoring Systematic has three major software components: a source code instrumentor which automatically inserts active event recorders into program source code before compilation; a run-time performance monitoring library which collects performance data; and a visualization tool-set which reconstructs program execution based on the data collected. Besides being used as a prototype for developing new techniques for instrumenting, monitoring and presenting parallel program execution, AIMS is also being incorporated into the run-time environments of various hardware testbeds to evaluate their impact on user productivity. Currently, the execution of FORTRAN and C programs on the Intel Paragon and PALM workstations can be automatically instrumented and monitored. Performance data thus collected can be displayed graphically on various workstations. The process of performance tuning with AIMS will be illustrated using various NAB Parallel Benchmarks. This report includes a description of the internal architecture of AIMS and a listing of the source code.

Yan, Jerry C.↗

Charon Toolkit for Parallel, Implicit Structured-Grid Computations: Functional Design

In a previous report the design concepts of Charon were presented. Charon is a toolkit that aids engineers in developing scientific programs for structured-grid applications to be run on MIMD parallel computers. It constitutes an augmentation of the general-purpose MPI-based message-passing layer, and provides the user with a hierarchy of tools for rapid prototyping and validation of parallel programs, and subsequent piecemeal performance tuning. Here we describe the implementation of the domain decomposition tools used for creating data distributions across sets of processors. We also present the hierarchy of parallelization tools that allows smooth translation of legacy code (or a serial design) into a parallel program. Along with the actual tool descriptions, we will present the considerations that led to the particular design choices. Many of these are motivated by the requirement that Charon must be useful within the traditional computational environments of Fortran 77 and C. Only the Fortran 77 syntax will be presented in this report.

VanderWijngaart, Rob F.↗

On-Board Deployment Event Verification for GOES-R Spacecraft

As is common with many spacecraft designs, the GOES-R vehicles require a series of deployment events to transition from the launch configuration to the operational configuration. Rather than implementing additional sensors to verify various deployments, the GOES-R program developed an alternate approach that uses existing gyro rate sensing. The approach includes two pieces: the first is a new onboard shock detection capability to confirm initiation of individual deployment events, and the second is a ground-based dynamics verification step to confirm completion of deployment events. We first present the algorithm that detects shock events for various deployment devices along with the parameter tuning performed during ground tests. We then show inflight performance of the shock detection algorithm. While shock detection is useful for observing initiation of deployment events, completion of some deployment events cannot be determined by shock detection alone. For these events, such as solar panel latch-up and deployable boom extension, the program developed dynamics models for the deployment transient responses. High-rate gyro data were recorded for these events, which allow the ground team to verify that the appendages were fully deployed. We show the predictive models for these events and corresponding flight results.

GeoXO↗

Performance Evaluation Methodologies and Tools for Massively Parallel Programs

The need for computing power has forced a migration from serial computation on a single processor to parallel processing on multiprocessors. However, without effective means to monitor (and analyze) program execution, tuning the performance of parallel programs becomes exponentially difficult as program complexity and machine size increase. The recent introduction of performance tuning tools from various supercomputer vendors (Intel's ParAide, TMC's PRISM, CSI'S Apprentice, and Convex's CXtrace) seems to indicate the maturity of performance tool technologies and vendors'/customers' recognition of their importance. However, a few important questions remain: What kind of performance bottlenecks can these tools detect (or correct)? How time consuming is the performance tuning process? What are some important technical issues that remain to be tackled in this area? This workshop reviews the fundamental concepts involved in analyzing and improving the performance of parallel and heterogeneous message-passing programs. Several alternative strategies will be contrasted, and for each we will describe how currently available tuning tools (e.g., AIMS, ParAide, PRISM, Apprentice, CXtrace, ATExpert, Pablo, IPS-2)) can be used to facilitate the process. We will characterize the effectiveness of the tools and methodologies based on actual user experiences at NASA Ames Research Center. Finally, we will discuss their limitations and outline recent approaches taken by vendors and the research community to address them.

Yan, Jerry C.↗

Description and performance analysis of a generalized optimal algorithm for aerobraking guidance

A practical real-time guidance algorithm has been developed for aerobraking vehicles which nearly minimizes the maximum heating rate, the maximum structural loads, and the post-aeropass delta V requirement for orbit insertion. The algorithm is general and reusable in the sense that a minimum of assumptions are made, thus greatly reducing the number of parameters that must be determined prior to a given mission. A particularly interesting feature is that in-plane guidance performance is tuned by adjusting one mission-dependent, the bank margin; similarly, the out-of-plane guidance performance is tuned by adjusting a plane controller time constant. Other features of the algorithm are simplicity, efficiency and ease of use. The trimmed vehicle with bank angle modulation as the method of trajectory control. Performance of this guidance algorithm is examined by its use in an aerobraking testbed program. The performance inquiry extends to a wide range of entry speeds covering a number of potential mission applications. Favorable results have been obtained with a minimum of development effort, and directions for improvement of performance are indicated.

Evans, Steven W.↗

Examination of a Practical Aerobraking Guidance Algorithm

A practical real time guidance algorithm has been developed for aerobraking vehicles that minimizes the post-aeropass Delta V requirements for orbit insertion while nearly minimizing the maximum heating rate and the maximum structural loads. The algorithm is general in the sense that a minimum of assumptions is made, thus greatly reducing the number of parameters that must be determined prior to a given mission. An interesting feature is that in-plane guidance performance is tuned by adjusting one mission-dependent parameter, the bank margin; similarly, the out-of-plane guidance performance is tuned by adjusting a plane controller time constant. Other features of the algorithm are simplicity, efficiency, and ease of use. The algorithm is designed for, but not restricted to, a trimmed vehicle with bank angle modulation as the method of trajectory control. Performance of this guidance algorithm during flight in Earth's atmosphere is examined by its use in an aerobraking testbed program. The performance inquiry extends to a wide range of entry speeds covering a number of potential mission applications. Favorable results have been obtained with a minimum of development effort, and directions for improvement of performance are indicated.

Evans, Steven W.↗

Performance advantages of dynamically tuned gyroscopes in high accuracy spacecraft pointing and stabilization applications

The paper compares and describes the advantages of dry tuned gyros over floated gyros for space applications. Attention is given to describing the Teledyne SDG-5 gyro and the second-generation NASA Standard Dry Rotor Inertial Reference Unit (DRIRU II). Certain tests which were conducted to evaluate the SDG-5 and DRIRU II for specific mission requirements are outlined, and their results are compared with published test results on other gyro types. Performance advantages are highlighted.

Irvine, R.↗

Legacy Code Modernization

Over the past decade, high performance computing has evolved rapidly; systems based on commodity microprocessors have been introduced in quick succession from at least seven vendors/families. Porting codes to every new architecture is a difficult problem; in particular, here at NASA, there are many large CFD applications that are very costly to port to new machines by hand. The LCM ("Legacy Code Modernization") Project is the development of an integrated parallelization environment (IPE) which performs the automated mapping of legacy CFD (Fortran) applications to state-of-the-art high performance computers. While most projects to port codes focus on the parallelization of the code, we consider porting to be an iterative process consisting of several steps: 1) code cleanup, 2) serial optimization,3) parallelization, 4) performance monitoring and visualization, 5) intelligent tools for automated tuning using performance prediction and 6) machine specific optimization. The approach for building this parallelization environment is to build the components for each of the steps simultaneously and then integrate them together. The demonstration will exhibit our latest research in building this environment: 1. Parallelizing tools and compiler evaluation. 2. Code cleanup and serial optimization using automated scripts 3. Development of a code generator for performance prediction 4. Automated partitioning 5. Automated insertion of directives. These demonstrations will exhibit the effectiveness of an automated approach for all the steps involved with porting and tuning a legacy code application for a new architecture.

Hribar, Michelle R.↗

How Difficult is it to Reduce Low-Level Cloud Biases With the Higher-Order Turbulence Closure Approach in Climate Models?

Low-level clouds cover nearly half of the Earth and play a critical role in regulating the energy and hydrological cycle. Despite the fact that a great effort has been put to advance the modeling and observational capability in recent years, low-level clouds remains one of the largest uncertainties in the projection of future climate change. Low-level cloud feedbacks dominate the uncertainty in the total cloud feedback in climate sensitivity and projection studies. These clouds are notoriously difficult to simulate in climate models due to its complicated interactions with aerosols, cloud microphysics, boundary-layer turbulence and cloud dynamics. The biases in both low cloud coverage/water content and cloud radiative effects (CREs) remain large. A simultaneous reduction in both cloud and CRE biases remains elusive. This presentation first reviews the effort of implementing the higher-order turbulence closure (HOC) approach to representing subgrid-scale turbulence and low-level cloud processes in climate models. There are two HOCs that have been implemented in climate models. They differ in how many three-order moments are used. The CLUBB are implemented in both CAM5 and GDFL models, which are compared with IPHOC that is implemented in CAM5 by our group. IPHOC uses three third-order moments while CLUBB only uses one third-order moment while both use a joint double-Gaussian distribution to represent the subgrid-scale variability. Despite that HOC is more physically consistent and produces more realistic low-cloud geographic distributions and transitions between cumulus and stratocumulus regimes, GCMs with traditional cloud parameterizations outperform in CREs because tuning of this type of models is more extensively performed than those with HOCs. We perform several tuning experiments with CAM5 implemented with IPHOC in an attempt to produce the nearly balanced global radiative budgets without deteriorating the low-cloud simulation. One of the issues in CAM5-IPHOC is that cloud water content is much higher than in CAM5, which is combined with higher low-cloud coverage to produce larger shortwave CREs in some low-cloud prevailing regions. Thus, the cloud-radiative feedbacks are exaggerated there. The turning exercise is focused on microphysical parameters, which are also commonly used for tuning in climate models. The results will be discussed in this presentation.

Xu, Kuan-Man↗

Extremum-seeking control for an Ultrasonic/Sonic Driller/Corer (USDC) driven at high-power

Future NASA exploration missions will increasingly require sampling, in-situ analysis and possibly the return of material to Earth for further tests. One of the challenges to addressing this need is the ability to drill using for low axial loading while operating from light weight platforms (e.g., lander, rover, etc.) as well as operate at planets with low gravity. For this purpose, the authors developed the Ultrasonic/Sonic Driller/Corer (USDC) jointly with Cybersonics Inc. Studies of the operation of the USDC at high power have shown there is a critical need to self-tune to maintain the operation of the piezoelectric actuator at resonance. Performing such tuning is encountered with difficulties and to address them an extremum-seeking control algorithm is being investigated. This algorithm is designed to tune the driving frequency of a time-varying resonating actuator subjected to both random and high-power impulsive noise disturbances. Using this algorithm the performance of the actuator is monitored on a time-scale that is compatible with its slowly time-varying physical characteristics. The algorithm includes a parameter estimator, which estimates the coefficients of a function that characterizes the quality factor of the USDC. Since the parameter estimator converges sufficiently faster than the time-varying drift of the USDC's physical parameters, the proposed extremum-seeking estimation and control algorithm is potentially applicable for use as a closed-loop health monitoring system. Specifically, this system may be programmed to automatically adjust the duty-cycle of the sinusoidal driver signal to guarantee that the quality factor of the USDC does not fall below a user-defined set-point. Such fault-tolerant functionality is especially important in automated drilling applications where it is essential not to inadvertently drive the piezoelectric ceramic crystals of the USDC beyond their capacities. The details of the algorithm and experimental results will be described and discussed in this paper.

Peek-seeking estimation and control↗

Autonomous Performance Monitoring System: Monitoring and Self-Tuning (MAST)

Maintaining the long-term performance of software onboard a spacecraft can be a major factor in the cost of operations. In particular, the task of controlling and maintaining a future mission of distributed spacecraft will undoubtedly pose a great challenge, since the complexity of multiple spacecraft flying in formation grows rapidly as the number of spacecraft in the formation increases. Eventually, new approaches will be required in developing viable control systems that can handle the complexity of the data and that are flexible, reliable and efficient. In this paper we propose a methodology that aims to maintain the accuracy of flight software, while reducing the computational complexity of software tuning tasks. The proposed Monitoring and Self-Tuning (MAST) method consists of two parts: a flight software monitoring algorithm and a tuning algorithm. The dependency on the software being monitored is mostly contained in the monitoring process, while the tuning process is a generic algorithm independent of the detailed knowledge on the software. This architecture will enable MAST to be applicable to different onboard software controlling various dynamics of the spacecraft, such as attitude self-calibration, and formation control. An advantage of MAST over conventional techniques such as filter or batch least square is that the tuning algorithm uses machine learning approach to handle uncertainty in the problem domain, resulting in reducing over all computational complexity. The underlying concept of this technique is a reinforcement learning scheme based on cumulative probability generated by the historical performance of the system. The success of MAST will depend heavily on the reinforcement scheme used in the tuning algorithm, which guarantees the tuning solutions exist.

Peterson, Chariya↗