DocumentCode
2350067
Title
Achieving scalable cluster system analysis and management with a gossip-based network service
Author
Collins, D.E. ; George, A.D. ; Quander, R.A.
Author_Institution
FSU Coll. of Eng., Florida A&M Univ., Tallahassee, FL, USA
fYear
2001
fDate
2001
Firstpage
49
Lastpage
58
Abstract
Clusters of workstations are increasingly used for applications requiring high levels of both performance and reliability. Certain fundamental services are highly desirable to achieve these twin goals of network-based cluster system analysis and management. Among these services is the ability to detect network and node failures and the capability to efficiently determine computer and network load levels. Furthermore, the ability to allow for the distribution of administrative directives is also integral to the goal of cluster management. This paper presents a scalable approach to providing these vital support capabilities for distributed computing integrated into a cluster management system. Previous approaches to cluster management have suffered from problems of scalability and the inability to properly support heterogeneous systems in a non-proprietary fashion. This cluster management system employs gossip techniques to address the problem of scalability in network-based system management. The results of two case studies show that the cluster management system is scalable and has little adverse impact on the performance of sequential and parallel applications running on the managed system
Keywords
computer network management; computer network reliability; performance evaluation; workstation clusters; administrative directives; cluster system management; computer load levels; distributed computing; gossip-based network service; network failures; network load levels; network-based management; node failures; performance; reliability; scalable cluster system analysis; workstation clusters; Application software; Collision mitigation; Computer network management; Computer network reliability; Distributed computing; Educational institutions; Heart beat; Reliability engineering; Scalability; Workstations;
fLanguage
English
Publisher
ieee
Conference_Titel
Local Computer Networks, 2001. Proceedings. LCN 2001. 26th Annual IEEE Conference on
Conference_Location
Tampa, FL
ISSN
0742-1303
Print_ISBN
0-7695-1321-2
Type
conf
DOI
10.1109/LCN.2001.990767
Filename
990767
Link To Document