DocumentCode :
1467795
Title :
Adaptive control of Markov chains with average cost
Author :
Ren, Zhiyuan ; Krogh, Bruce H.
Author_Institution :
Dept. of Electr. & Comput. Eng., Carnegie Mellon Univ., Pittsburgh, PA, USA
Volume :
46
Issue :
4
fYear :
2001
fDate :
4/1/2001 12:00:00 AM
Firstpage :
613
Lastpage :
617
Abstract :
We present an adaptive control scheme for the control of Markov chains to minimize long-run average cost when the system transition and reward structures are unknown. Q-factors are estimated along a single sample path of the system and control actions are applied based on the latest estimates. We prove that an optimal policy is obtained asymptotically with probability one. More importantly, we prove that optimal system performance is achieved as well, which means that the performance of the system can not be bettered even if the system transition and reward structures are known. An example is given to illustrate our adaptive control scheme
Keywords :
Markov processes; adaptive control; convergence; probability; Markov chains; Q-factors; long-run average cost; optimal policy; Adaptive control; Control systems; Convergence; Cost function; Optimal control; Parameter estimation; Q factor; Stochastic processes; Stochastic systems; System performance;
fLanguage :
English
Journal_Title :
Automatic Control, IEEE Transactions on
Publisher :
ieee
ISSN :
0018-9286
Type :
jour
DOI :
10.1109/9.917662
Filename :
917662
Link To Document :
بازگشت