论文信息 - Achieving scalable cluster system analysis and management with a gossip-based network service

Achieving scalable cluster system analysis and management with a gossip-based network service

Clusters of workstations are increasingly used for applications requiring high levels of both performance and reliability. Certain fundamental services are highly desirable to achieve these twin goals of network-based cluster system analysis and management. Among these services is the ability to detect network and node failures and the capability to efficiently determine computer and network load levels. Furthermore, the ability to allow for the distribution of administrative directives is also integral to the goal of cluster management. This paper presents a scalable approach to providing these vital support capabilities for distributed computing integrated into a cluster management system. Previous approaches to cluster management have suffered from problems of scalability and the inability to properly support heterogeneous systems in a non-proprietary fashion. This cluster management system employs gossip techniques to address the problem of scalability in network-based system management. The results of two case studies show that the cluster management system is scalable and has little adverse impact on the performance of sequential and parallel applications running on the managed system.

Alan D. George | David E. Collins | R. A. Quander

[1] Alan D. George,et al. Parallel and Sequential Job Scheduling in Heterogeneous Clusters: A Simulation Study Using Software in the Loop , 2001, Simul..

[2] Barry W. Johnson. Design & analysis of fault tolerant digital systems , 1988 .

[3] R. Ramaswami,et al. Book Review: Design and Analysis of Fault-Tolerant Digital Systems , 1990 .

[4] Alan D. George,et al. Performance analysis of flat and layered gossip services for failure detection and consensus in scalable heterogeneous clusters , 2001, Proceedings 15th International Parallel and Distributed Processing Symposium. IPDPS 2001.

[5] Tore Anders Aamodt. Design and implementation issues for an SCI cluster configuration system , 1998 .

[6] Uyless D. Black. Network Management Standards: SNMP, CMIP, TMN, MIBs and Object Libraries , 1992 .

[7] David H. Bailey,et al. The Nas Parallel Benchmarks , 1991, Int. J. High Perform. Comput. Appl..

[8] Rajkumar Buyya,et al. PARMON: a portable and scalable monitoring system for clusters , 2000 .

[9] Robbert van Renesse,et al. A Gossip-Style Failure Detection Service , 2009 .