Search NASASearch

SEARCH · Search NASA

Results for “performance optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Dynamic remapping decisions in multi-phase parallel computations

The effectiveness of any given mapping of workload to processors in a parallel system is dependent on the stochastic behavior of the workload. Program behavior is often characterized by a sequence of phases, with phase changes occurring unpredictably. During a phase, the behavior is fairly stable, but may become quite different during the next phase. Thus a workload assignment generated for one phase may hinder performance during the next phase. We consider the problem of deciding whether to remap a paralled computation in the face of uncertainty in remapping's utility. Fundamentally, it is necessary to balance the expected remapping performance gain against the delay cost of remapping. This paper treats this problem formally by constructing a probabilistic model of a computation with at most two phases. We use stochastic dynamic programming to show that the remapping decision policy which minimizes the expected running time of the computation has an extremely simple structure: the optimal decision at any step is followed by comparing the probability of remapping gain against a threshold. This theoretical result stresses the importance of detecting a phase change, and assessing the possibility of gain from remapping. We also empirically study the sensitivity of optimal performance to imprecise decision threshold. Under a wide range of model parameter values, we find nearly optimal performance if remapping is chosen simply when the gain probability is high. These results strongly suggest that except in extreme cases, the remapping decision problem is essentially that of dynamically determining whether gain can be achieved by remapping after a phase change; precise quantification of the decision model parameters is not necessary.

Nicol, D. M.

Performance and Portability of a Linear Solver Across Emerging Architectures

A linear solver algorithm used by a large-scale unstructured-grid computational fluid dynamics application is examined for a broad range of familiar and emerging architectures. Efficient implementation of a linear solver is challenging on recent CPUs offering vector architectures. Vector loads and stores are essential to effectively utilize available memory bandwidth on CPUs, and maintaining performance across different CPUs can be difficult in the face of varying vector lengths offered by each. A similar challenge occurs on GPU architectures, where it is essential to have coalesced memory accesses to utilize memory bandwidth effectively. In this work, we demonstrate that restructuring a computation, and possibly data layout, with regard to architecture is essential to achieve optimal performance by establishing a performance benchmark for each target architecture in a low level language such as vector intrinsics or CUDA. In doing so, we demonstrate how a linear solver kernel can be mapped to Intel® Xeon™ and Xeon Phi™, Marvell® ThunderX2®, NEC® SX-Aurora™ TSUBASA Vector Engine, and NVIDIA® and AMD® GPUs. We further demonstrate that the required code restructuring can be achieved in higher level programming environments such as OpenACC, OCCA, and Intel® OneAPI™/SYCL, and that each generally results in optimal performance on the target architecture. Relative performance metrics for all implementations are shown, and subjective ratings for ease of implementation and optimization are suggested.

Programming models

Performance and optimization of a derated ion thruster for auxiliary propulsion

The characteristics and implications of use of a derated ion thruster for north-south stationkeeping (NSSK) propulsion are discussed. A derated thruster is a 30 cm diameter primary propulsion ion thruster operated at highly throttled conditions appropriate to NSSK functions. The performance characteristics of a 30 cm ion thruster are presented, emphasizing throttled operation at low specific impulse and high thrust-to-power ratio. Performance data and component erosion are compared to other NSSK ion thrusters. Operations benefits derived from the performance advantages of the derated approach are examined assuming an INTELSAt 7-type spacecraft. Minimum ground test facility pumping capabilities required to maintain facility enhanced accelerator grid erosion at acceptable levels in a lifetest are quantified as a function of thruster operating condition. Approaches to reducing the derated thruster mass and volume are also discussed.

Patterson, Michael J.

Performance and optimization of a 'derated' ion thruster for auxiliary propulsion

This paper discusses the characteristics and implications of use of a derated ion thruster for north-south-stationkeeping (NSSK) propulsion. A derated thruster is a 30-cm diameter primary propulsion ion thruster operated at highly throttled conditions appropriate to NSSK functions. The performance characteristics of a 30-cm ion thruster are presented, emphasizing throttled operation at low specific impulse and high thrust-to-power ratio. Performance data and component erosion are compared to other NSSK ion thrusters. Operations benefits derived from the performance advantages of the derated approach are examined assuming an Intelsat VII-type spacecraft. Minimum ground test facility pumping capabilities required to maintain facility enhanced accelerator grid erosion at acceptable levels in a lifetest are quantified as a function of thruster operating condition. Novel approaches to reducing the derated thruster mass and volume are also discussed.

Patterson, Michael J.

Optimization and performance calculation of dual-rotation propellers

An analysis is given which enables the design of dual-rotation propellers. It relies on the use of a new tip loss factor deduced from T. Theodorsen's measurements coupled with the general methodology of C. N. H. Lock. In addition, it includes the effect of drag in optimizing. Some values for the tip loss factor are calculated for one advance ratio.

Davidson, R. E.

On the use of controls for subsonic transport performance improvement: Overview and future directions

Increasing competition among airline manufacturers and operators has highlighted the issue of aircraft efficiency. Fewer aircraft orders have led to an all-out efficiency improvement effort among the manufacturers to maintain if not increase their share of the shrinking number of aircraft sales. Aircraft efficiency is important in airline profitability and is key if fuel prices increase from their current low. In a continuing effort to improve aircraft efficiency and develop an optimal performance technology base, NASA Dryden Flight Research Center developed and flight tested an adaptive performance seeking control system to optimize the quasi-steady-state performance of the F-15 aircraft. The demonstrated technology is equally applicable to transport aircraft although with less improvement. NASA Dryden, in transitioning this technology to transport aircraft, is specifically exploring the feasibility of applying adaptive optimal control techniques to performance optimization of redundant control effectors. A simulation evaluation of a preliminary control law optimizes wing-aileron camber for minimum net aircraft drag. Two submodes are evaluated: one to minimize fuel and the other to maximize velocity. This paper covers the status of performance optimization of the current fleet of subsonic transports. Available integrated controls technologies are reviewed to define approaches using active controls. A candidate control law for adaptive performance optimization is presented along with examples of algorithm operation.

Gilyard, Glenn

Automation of POST Cases via External Optimizer and "Artificial p2" Calculation

During early conceptual design of complex systems, speed and accuracy are often at odds with one another. While many characteristics of the design are fluctuating rapidly during this phase there is nonetheless a need to acquire accurate data from which to down-select designs as these decisions will have a large impact upon program life-cycle cost. Therefore enabling the conceptual designer to produce accurate data in a timely manner is tantamount to program viability. For conceptual design of launch vehicles, trajectory analysis and optimization is a large hurdle. Tools such as the industry standard Program to Optimize Simulated Trajectories (POST) have traditionally required an expert in the loop for setting up inputs, running the program, and analyzing the output. The solution space for trajectory analysis is in general non-linear and multi-modal requiring an experienced analyst to weed out sub-optimal designs in pursuit of the global optimum. While an experienced analyst presented with a vehicle similar to one which they have already worked on can likely produce optimal performance figures in a timely manner, as soon as the "experienced" or "similar" adjectives are invalid the process can become lengthy. In addition, an experienced analyst working on a similar vehicle may go into the analysis with preconceived ideas about what the vehicle's trajectory should look like which can result in sub-optimal performance being recorded. Thus, in any case but the ideal either time or accuracy can be sacrificed. In the authors' previous work a tool called multiPOST was created which captures the heuristics of a human analyst over the process of executing trajectory analysis with POST. However without the instincts of a human in the loop, this method relied upon Monte Carlo simulation to find successful trajectories. Overall the method has mixed results, and in the context of optimizing multiple vehicles it is inefficient in comparison to the method presented POST's internal optimizer functions like any other gradient-based optimizer. It has a specified variable to optimize whose value is represented as optval, a set of dependent constraints to meet with associated forms and tolerances whose value is represented as p2, and a set of independent variables known as the u-vector to modify in pursuit of optimality. Each of these quantities are calculated or manipulated at a certain phase within the trajectory. The optimizer is further constrained by the requirement that the input u-vector must result in a trajectory which proceeds through each of the prescribed events in the input file. For example, if the input u-vector causes the vehicle to crash before it can achieve the orbital parameters required for a parking orbit, then the run will fail without engaging the optimizer, and a p2 value of exactly zero is returned. This poses a problem, as this "non-connecting" region of the u-vector space is far larger than the "connecting" region which returns a non-zero value of p2 and can be worked on by the internal optimizer. Finding this connecting region and more specifically the global optimum within this region has traditionally required the use of an expert analyst.

Dees, Patrick D.

Agent Reward Shaping for Alleviating Traffic Congestion

Traffic congestion problems provide a unique environment to study how multi-agent systems promote desired system level behavior. What is particularly interesting in this class of problems is that no individual action is intrinsically "bad" for the system but that combinations of actions among agents lead to undesirable outcomes, As a consequence, agents need to learn how to coordinate their actions with those of other agents, rather than learn a particular set of "good" actions. This problem is ubiquitous in various traffic problems, including selecting departure times for commuters, routes for airlines, and paths for data routers. In this paper we present a multi-agent approach to two traffic problems, where far each driver, an agent selects the most suitable action using reinforcement learning. The agent rewards are based on concepts from collectives and aim to provide the agents with rewards that are both easy to learn and that if learned, lead to good system level behavior. In the first problem, we study how agents learn the best departure times of drivers in a daily commuting environment and how following those departure times alleviates congestion. In the second problem, we study how agents learn to select desirable routes to improve traffic flow and minimize delays for. all drivers.. In both sets of experiments,. agents using collective-based rewards produced near optimal performance (93-96% of optimal) whereas agents using system rewards (63-68%) barely outperformed random action selection (62-64%) and agents using local rewards (48-72%) performed worse than random in some instances.

Tumer, Kagan

Parametric study of critical constraints for a canard configured medium range transport using conceptual design optimization

Constrained parameter optimization was used to perform optimal conceptual design of both canard and conventional configurations of a medium range transport. A number of design constants and design constraints were systematically varied to compare the sensitivities of canard and conventional configurations to a variety of technology assumptions. Main landing gear location and horizontal stabilizer high-lift performance were identified as critical design parameters for a statically stable, subsonic canard transport.

Arbuckle, P. D.

Parametric study of a canard-configured transport using conceptual design optimization

Constrained-parameter optimization is used to perform optimal conceptual design of both canard and conventional configurations of a medium-range transport. A number of design constants and design constraints are systematically varied to compare the sensitivities of canard and conventional configurations to a variety of technology assumptions. Main-landing-gear location and canard surface high-lift performance are identified as critical design parameters for a statically stable, subsonic, canard-configured transport.

Arbuckle, P. D.

Development of management technology for large power systems

Autonomous power management has been proposed as a method to perform optimization of power subsystem performance in connection with the management of multikilowatt space platforms. A concept for a 250-kW utility-type power subsystem was developed. A Cassegrain concentrator solar array primary source is conditioned by a solar array switching unit which supplies seventeen 220 +20 Vdc power channels. A power management subsystem provides the monitoring and control of the overall electrical power subsystem. The discussed system concept for autonomous management of high power space platforms utilizes on-board microprocessors in a decentralized data management architecture. A data bus protocol and a data bus contention resolution scheme were selected in conjunction with the dencentralized management architecture.

Decker, D. K.

Selection for optimal crew performance - Relative impact of selection and training

An empirical study supporting Helmreich's (1986) theoretical work on the distinct manner in which training and selection impact crew coordination is presented. Training is capable of changing attitudes, while selection screens for stable personality characteristics. Training appears least effective for leadership, an area strongly influenced by personality. Selection is least effective for influencing attitudes about personal vulnerability to stress, which appear to be trained in resource management programs. Because personality correlates with attitudes before and after training, it is felt that selection may be necessary even with a leadership-oriented training cirriculum.

Chidester, Thomas R.

Optimizing raid performance with cache

We live in a world of increasingly complex applications and operating systems. Information is increasing at a mind-boggling rate. The consolidation of text, voice, and imaging represents an even greater challenge for our information systems. Which forced us to address three important questions: Where do we store all this information? How do we access it? And, how do we protect it against the threat of loss or damage? Introduced in the 1980s, RAID (Redundant Arrays of Independent Disks) represents a cost-effective solution to the needs of the information age. While fulfilling expectations for high storage, and reliability, RAID is sometimes subject to criticisms in the area of performance. However, there are design elements that can significantly enhance performance. They can be subdivided into two areas: (1) RAID levels or basic architecture. And, (2) enhancement schemes such as intelligent caching, support of tagged command queuing, and use of SCSI-2 Fast and Wide features.

Bouzari, Alex

The Efficiency of Various Computers and Optimizations in Performing Finite Element Computations

With the advent of computers with many processors, it becomes unclear how to best exploit this advantage. For example, matrices can be inverted by applying several processors to each vector operation, or one processor can be applied to each matrix. The former approach has diminishing returns beyond a handful of processors, but how many processors depends on the computer architecture. Applying one processor to each matrix is feasible with enough ram memory and scratch disk space, but the speed at which this is done is found to vary by a factor of three depending on how it is done. The cost of the computer must also be taken into account. A computer with many processors and fast interprocessor communication is much more expensive than the same computer and processors with slow interprocessor communication. Consequently, for problems that require several matrices to be inverted, the best speed per dollar for computers is found to be several small workstations that are networked together, such as in a Beowulf cluster. Since these machines typically have two processors per node, each matrix is most efficiently inverted with no more than two processors assigned to it.

Marcus, Martin H.

Particle Swarm Optimization

The purpose of this paper is to show how the search algorithm known as particle swarm optimization performs. Here, particle swarm optimization is applied to structural design problems, but the method has a much wider range of possible applications. The paper's new contributions are improvements to the particle swarm optimization algorithm and conclusions and recommendations as to the utility of the algorithm, Results of numerical experiments for both continuous and discrete applications are presented in the paper. The results indicate that the particle swarm optimization algorithm does locate the constrained minimum design in continuous applications with very good precision, albeit at a much higher computational cost than that of a typical gradient based optimizer. However, the true potential of particle swarm optimization is primarily in applications with discrete and/or discontinuous functions and variables. Additionally, particle swarm optimization has the potential of efficient computation with very large numbers of concurrently operating processors.

Venter, Gerhard