Search NASASearch

SEARCH · Search NASA

Results for “Job scheduling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Combining Quick-Turnaround and Batch Workloads at Scale

NAS uses PBS Professional to schedule and manage the workload on Pleiades, an 11,000+ node 1B cluster. At this scale the user experience for quick-turnaround jobs can degrade, which led NAS initially to set up two separate PBS servers, each dedicated to a particular workload. Recently we have employed PBS hooks and scheduler modifications to merge these workloads together under one PBS server, delivering sub-1-minute start times for the quick-turnaround workload, and enabling dynamic management of the resources set aside for that workload.

Matthews, Gregory A.

Prototyping an Onboard Scheduler for the Mars 2020 Rover

Efficiently operating a rover on the surface of Mars is challenging. Two factors combine to make this job particularly difficult: 1) communication opportunities are limited, 2) certain aspects of rover performance are difficult to predict. With limited communications, the rover must be given instructions on what to do for one or more Martian days at a time. In addition, the duration of many rover activities can be hard to predict, which leads to unpredictable energy use. Traditionally, conservatism is used to keep the rover safe and healthy. This approach, however can lead to a measurable loss in rover productivity. To regain some of this productivity, the Mars 2020 mission is prototyping the use of onboard scheduling software. The primary objective of this software is to identify and utilize opportunities that arise when actual rover performance is more efficient than the original, conservative prediction.

Benowitz, Ed

Space shuttle descent design: From development to operations

The descent guidance system, the descent trajectories design, and generating of the associated flight products are discussed. The programs which allow the successful transitions from development to STS operations, resulting in reduced manpower requirements and compressed schedules for flight design cycles are addressed. The topics include: (1) continually upgraded tools for the job, i.e., consolidating tools via electronic data transfers, tailoring general purpose software for needs, easy access to tools through an interactive approach, and appropriate flexibility to allow design changes and provide growth capability; (2) stabilizing the flight profile designs (I-loads) in an uncertain environment; and (3) standardizing external interfaces within performance and subsystems constraints of the Orbiter.

Crull, T. J.

Self-Scheduling Parallel Methods for Multiple Serial Codes with Application to WOPWOP

This paper presents a scheme for efficiently running a large number of serial jobs on parallel computers. Two examples are given of computer programs that run relatively quickly, but often they must be run numerous times to obtain all the results needed. It is very common in science and engineering to have codes that are not massive computing challenges in themselves, but due to the number of instances that must be run, they do become large-scale computing problems. The two examples given here represent common problems in aerospace engineering: aerodynamic panel methods and aeroacoustic integral methods. The first example simply solves many systems of linear equations. This is representative of an aerodynamic panel code where someone would like to solve for numerous angles of attack. The complete code for this first example is included in the appendix so that it can be readily used by others as a template. The second example is an aeroacoustics code (WOPWOP) that solves the Ffowcs Williams Hawkings equation to predict the far-field sound due to rotating blades. In this example, one quite often needs to compute the sound at numerous observer locations, hence parallelization is utilized to automate the noise computation for a large number of observers.

Long, Lyle N.

So Easy a Greybeard Can Do It: Mobile Paperless Engineering

Picture a NASA Engineer out at the launch pad - some new construction has been completed, and he has been tasked to inspect the job. While working, he repeatedly travels between the site and his desk, spending time seeking out information, printing out drawings for reference, and attempting to align schedules to meet with the team. This type of scenario is not just specific to Kennedy Space Center though, many industries partake in similar circumstances most realizing that there is always a need for more information, and often at a moments notice. While this approach will eventually get the job done, NASA KSC-ESC has questioned how to use readily available resources to streamline work processes, become more efficient, and produce less waste.

Greybeard

The SGI/Cray T3E: Experiences and Insights

The NASA Goddard Space Flight Center is home to the fifth most powerful supercomputer in the world, a 1024 processor SGI/Cray T3E-600. The original 512 processor system was placed at Goddard in March, 1997 as part of a cooperative agreement between the High Performance Computing and Communications Program's Earth and Space Sciences Project (ESS) and SGI/Cray Research. The goal of this system is to facilitate achievement of the Project milestones of 10, 50 and 100 GFLOPS sustained performance on selected Earth and space science application codes. The additional 512 processors were purchased in March, 1998 by the NASA Earth Science Enterprise for the NASA Seasonal to Interannual Prediction Project (NSIPP). These two "halves" still operate as a single system, and must satisfy the unique requirements of both aforementioned groups, as well as guest researchers from the Earth, space, microgravity, manned space flight and aeronautics communities. Few large scalable parallel systems are configured for capability computing, so models are hard to find. This unique environment has created a challenging system administration task, and has yielded some insights into the supercomputing needs of the various NASA Enterprises, as well as insights into the strengths and weaknesses of the T3E architecture and software. The T3E is a distributed memory system in which the processing elements (PE's) are connected by a low latency, high bandwidth bidirectional 3-D torus. Due to the focus on high speed communication between PE's, the T3E requires PE's to be allocated contiguously per job. Further, jobs will only execute on the user specified number of PE's and PE timesharing is possible but impractical. With a highly varied job mix in both size and runtime of jobs, the resulting scenario is PE fragmentation and an inability to achieve near 100% utilization. SGI/Cray has provided several scheduling and configuration tools to minimize the impact of fragmentation. These tools include PScheD (the political scheduler), GRM (the global resource manager) and NQE (the Network Queuing Environment). Features and impact of these tools will be discussed, as will resulting performance and utilization data. As a distributed memory system, the T3E is designed to be programmed through explicit message passing. Consequently, certain assumptions related to code design are made by the operating system (UNICOS/mk) and its scheduling tools. With the exception of HPF, which does run on the T3E, however poorly, alternative programming styles have the potential to impact the T3E in unexpected and undesirable ways. Several examples will be presented (preceeded with the disclaimer, "Don't try this at home! Violators will be prosecuted!")

Bernard, Lisa Hamet

Scheduling for Parallel Supercomputing: A Historical Perspective of Achievable Utilization

The NAS facility has operated parallel supercomputers for the past 11 years, including the Intel iPSC/860, Intel Paragon, Thinking Machines CM-5, IBM SP-2, and Cray Origin 2000. Across this wide variety of machine architectures, across a span of 10 years, across a large number of different users, and through thousands of minor configuration and policy changes, the utilization of these machines shows three general trends: (1) scheduling using a naive FIFO first-fit policy results in 40-60% utilization, (2) switching to the more sophisticated dynamic backfilling scheduling algorithm improves utilization by about 15 percentage points (yielding about 70% utilization), and (3) reducing the maximum allowable job size further increases utilization. Most surprising is the consistency of these trends. Over the lifetime of the NAS parallel systems, we made hundreds, perhaps thousands, of small changes to hardware, software, and policy, yet, utilization was affected little. In particular these results show that the goal of achieving near 100% utilization while supporting a real parallel supercomputing workload is unrealistic.

Jones, James Patton

Affordability Approaches for Human Space Exploration

The design and development of historical NASA Programs (Apollo, Shuttle and International Space Station), have been based on pre-agreed missions which included specific pre-defined destinations (e.g., the Moon and low Earth orbit). Due to more constrained budget profiles, and the desire to have a more flexible architecture for Mission capture as it is affordable, NASA is working toward a set of Programs that are capability based, rather than mission and/or destination specific. This means designing for a performance capability that can be applied to a specific human exploration mission/destination later (sometime years later). This approach does support developing systems to flatter budgets over time, however, it also poses the challenge of how to accomplish this effectively while maintaining a trained workforce, extensive manufacturing, test and launch facilities, and ensuring mission success ranging from Low Earth Orbit to asteroid destinations. NASA Marshall Space Flight Center (MSFC) in support of Exploration Systems Directorate (ESD) in Washington, DC has been developing approaches to track affordability across multiple Programs. The first step is to ensure a common definition of affordability: the discipline to bear cost in meeting a budget with margin over the life of the program. The second step is to infuse responsibility and accountability for affordability into all levels of the implementing organization since affordability is no single person s job; it is everyone s job. The third step is to use existing data to identify common affordability elements organized by configuration (vehicle/facility), cost, schedule, and risk. The fourth step is to analyze and trend this affordability data using an affordability dashboard to provide status, measures, and trends for ESD and Program level of affordability tracking. This paper will provide examples of how regular application of this approach supports affordable and therefore sustainable human space exploration architecture.

Holladay, Jon

Deep Space Network (DSN), Network Operations Control Center (NOCC) computer-human interfaces

The Network Operations Control Center (NOCC) of the DSN is responsible for scheduling the resources of DSN, and monitoring all multi-mission spacecraft tracking activities in real-time. Operations performs this job with computer systems at JPL connected to over 100 computers at Goldstone, Australia and Spain. The old computer system became obsolete, and the first version of the new system was installed in 1991. Significant improvements for the computer-human interfaces became the dominant theme for the replacement project. Major issues required innovating problem solving. Among these issues were: How to present several thousand data elements on displays without overloading the operator? What is the best graphical representation of DSN end-to-end data flow? How to operate the system without memorizing mnemonics of hundreds of operator directives? Which computing environment will meet the competing performance requirements? This paper presents the technical challenges, engineering solutions, and results of the NOCC computer-human interface design.

Ellman, Alvin

Method for resource control in parallel environments using program organization and run-time support

A system and method for dynamic scheduling and allocation of resources to parallel applications during the course of their execution. By establishing well-defined interactions between an executing job and the parallel system, the system and method support dynamic reconfiguration of processor partitions, dynamic distribution and redistribution of data, communication among cooperating applications, and various other monitoring actions. The interactions occur only at specific points in the execution of the program where the aforementioned operations can be performed efficiently.

Ekanadham, Kattamuri

Method for resource control in parallel environments using program organization and run-time support

A system and method for dynamic scheduling and allocation of resources to parallel applications during the course of their execution. By establishing well-defined interactions between an executing job and the parallel system, the system and method support dynamic reconfiguration of processor partitions, dynamic distribution and redistribution of data, communication among cooperating applications, and various other monitoring actions. The interactions occur only at specific points in the execution of the program where the aforementioned operations can be performed efficiently.

Ekanadham, Kattamuri

Putting EVM to the Test

IN MANY INSTANCES THERE IS NO FOREWARNING; SCHEDULES slip, costs soar, and the project manager is faced with the near impossible task of explaining why each impact occurred. With contractors performing the majority of the work, the management job can become even more obscure. The simple lack of proximity to the contractor can limit effective communication. Add to that a mixture of cultural differences and a desire for the contractor to portray the most optimistic view of their performance, and you create an even more difficult task for the project manager. This was the scenario when the Habitat Holding Rack (HHR) manager at Marshall Space Flight Center (MSFC), Stacy Counts, was introduced to the overall concept of Earned Value Management (EVM). Faced with increased costs (which eventually resulted in decreased scope of the project), continued schedule slides, and several technical anomalies, she was looking for a way to gain a better handle on the project performance. As a component of the Space Station Biological Research Program (SSBRP), the HHR project is an integral piece of the Program content. The HHR is the first rack hardware to be delivered for the Program and has therefore been the first rack to move through the trials of test and verification-documenting anomalies and technical difficulties that will benefit the other SSBRP rack projects. For these reasons, the HHR maintained high visibility throughout the manufacturing and assembly process, continuing through test and verification activities. Needless to say, the higher visibility emphasized the need for improved performance on this project. And to improve project performance, Stacy first had to figure out how to measure the cost, schedule and technical objectives effectively.

Kerby, Jerald

Scheduling revisited workstations in integrated-circuit fabrication

The cost of building new semiconductor wafer fabrication factories has grown rapidly, and a state-of-the-art fab may cost 250 million dollars or more. Obtaining an acceptable return on this investment requires high productivity from the fabrication facilities. This paper describes the Photo Dispatcher system which was developed to make machine-loading recommendations on a set of key fab machines. Dispatching policies that generally perform well in job shops (e.g., Shortest Remaining Processing Time) perform poorly for workstations such as photolithography which are visited several times by the same lot of silicon wafers. The Photo Dispatcher evaluates the history of workloads throughout the fab and identifies bottleneck areas. The scheduler then assigns priorities to lots depending on where they are headed after photolithography. These priorities are designed to avoid starving bottleneck workstations and to give preference to lots that are headed to areas where they can be processed with minimal waiting. Other factors considered by the scheduler to establish priorities are the nearness of a lot to the end of its process flow and the time that the lot has already been waiting in queue. Simulations that model the equipment and products in one of Texas Instrument's wafer fabs show the Photo Dispatcher can produce a 10 percent improvement in the time required to fabricate integrated circuits.

Kline, Paul J.

Emergency Preparedness for Catastrophic Events at Small and Medium Sized Airports: Lacking or Not?

The implementation of security methods and processes in general has had a decisive impact on the aviation industry. However, efforts to effectively coordinate varied aspects of security protocols between agencies and general aviation components have not been adequately addressed. Whether or not overall security issues, especially with regard to planning for catastrophic terrorist events, have been neglected at the nation's smaller airports is the main topic of this paper. For perspective, the term general aviation is generally accepted to include all flying except for military and scheduled airline operations. Genera aviation makes up more than 1 percent of the U.S. Gross Domestic Product and supports almost 1.3 mission high-skilled jobs in professional services and manufacturing and hence is an important component of the aviation industry (AOPA, n.d.). In both conceptual and practical terms, this paper argues for the proactive management of security planning and repeated security awareness training from both an individual and an organizational perspective within the general aviation venue. The results of a research project incorporating survey data from general aviation and small commercial airport managers as well as Transportation Security Administration (TSA) employees are reported. Survey findings suggest that miscommunication does take place on different organizational levels and that between TSA employees and airport management interaction can be contentious and cooperation diminished. The importance of organizational training for decreasing conflict and increasing security and preparedness is discussed as a primary implication.

Sweet, Kathleen M.

Task path planning, scheduling and learning for free-ranging robot systems

The development of robotics applications for space operations is often restricted by the limited movement available to guided robots. Free ranging robots can offer greater flexibility than physically guided robots in these applications. Presented here is an object oriented approach to path planning and task scheduling for free-ranging robots that allows the dynamic determination of paths based on the current environment. The system also provides task learning for repetitive jobs. This approach provides a basis for the design of free-ranging robot systems which are adaptable to various environments and tasks.

Wakefield, G. Steve

A distributed scheduling algorithm for heterogeneous real-time systems

Much of the previous work on load balancing and scheduling in distributed environments was concerned with homogeneous systems and homogeneous loads. Several of the results indicated that random policies are as effective as other more complex load allocation policies. The effects of heterogeneity on scheduling algorithms for hard real time systems is examined. A distributed scheduler specifically to handle heterogeneities in both nodes and node traffic is proposed. The performance of the algorithm is measured in terms of the percentage of jobs discarded. While a random task allocation is very sensitive to heterogeneities, the algorithm is shown to be robust to such non-uniformities in system components and load.

Zeineldine, Osman

Fatigue in operational settings: Examples from the aviation environment

The need for 24-h operations creates nonstandard and altered work schedules that can lead to cumulative sleep loss and circadian disruption. These factors can lead to fatigue and sleepiness and affect performance and productivity on the job. The approach, research, and results of the NASA Ames Fatigue Countermeasures Program are described to illustrate one attempt to address these issues in the aviation environment. The scientific and operational relevance of these factors is discussed, and provocative issues for future research are presented.

Rosekind, Mark R.