• DocumentCode
    2743031
  • Title

    Adaptive real-time monitoring for large-scale networked systems

  • Author

    Prieto, Alberto Gonzalez ; Stadler, Rolf

  • Author_Institution
    Sch. of Electr. Eng., R. Inst. of Technol., Stockholm, Sweden
  • fYear
    2009
  • fDate
    1-5 June 2009
  • Firstpage
    790
  • Lastpage
    795
  • Abstract
    The focus of this thesis is continuous real-time monitoring, which is essential for the realization of adaptive management systems in large-scale dynamic environments. Real-time monitoring provides the necessary input to the decision-making process of network management. We have developed, implemented, and evaluated a design for real-time continuous monitoring of global metrics with performance objectives, such as monitoring overhead and estimation accuracy. Global metrics describe the state of the system as a whole, in contrast to local metrics, such as device counters or local protocol states, which capture the state of a local entity. Global metrics are computed from local metrics using aggregation functions, such as SUM, AVERAGE and MAX. A key part in the design is a model for the distributed monitoring process that relates performance metrics to parameters that tune the behavior of a monitoring protocol. The model has been instrumental in designing a monitoring protocol that is controllable and achieves given performance objectives. Our design has proved to be effective in meeting performance objectives, efficient, adaptive to changes in the networking conditions, controllable along different performance dimensions, and scalable. We have implemented a prototype on a testbed of commercial routers, which proves the feasibility of the design, and, more generally, the feasibility of effective and efficient real-time monitoring in large network environments.
  • Keywords
    large-scale systems; telecommunication network management; adaptive management system; adaptive real-time monitoring; distributed monitoring process; global metrics; large-scale networked system; network management; Adaptive systems; Condition monitoring; Counting circuits; Decision making; Environmental management; Instruments; Large-scale systems; Measurement; Protocols; Real time systems; Adaptive management; large-scale distributed systems; real-time monitoring;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Integrated Network Management, 2009. IM '09. IFIP/IEEE International Symposium on
  • Conference_Location
    Long Island, NY
  • Print_ISBN
    978-1-4244-3486-2
  • Electronic_ISBN
    978-1-4244-3487-9
  • Type

    conf

  • DOI
    10.1109/INM.2009.5188884
  • Filename
    5188884