Show simple item record

contributor authorDuong, Thai
contributor authorNguyen
contributor authorNguyen, Thinh
date accessioned2017-05-09T01:27:03Z
date available2017-05-09T01:27:03Z
date issued2016
identifier issn0022-0434
identifier otherds_138_06_061009.pdf
identifier urihttp://yetl.yabesh.ir/yetl/handle/yetl/160690
description abstractMarkov decision process (MDP) is a wellknown framework for devising the optimal decisionmaking strategies under uncertainty. Typically, the decision maker assumes a stationary environment which is characterized by a timeinvariant transition probability matrix. However, in many realworld scenarios, this assumption is not justified, thus the optimal strategy might not provide the expected performance. In this paper, we study the performance of the classic value iteration algorithm for solving an MDP problem under nonstationary environments. Specifically, the nonstationary environment is modeled as a sequence of timevariant transition probability matrices governed by an adiabatic evolution inspired from quantum mechanics. We characterize the performance of the value iteration algorithm subject to the rate of change of the underlying environment. The performance is measured in terms of the convergence rate to the optimal average reward. We show two examples of queuing systems that make use of our analysis framework.
publisherThe American Society of Mechanical Engineers (ASME)
titleAdiabatic Markov Decision Process: Convergence of Value Iteration Algorithm
typeJournal Paper
journal volume138
journal issue6
journal titleJournal of Dynamic Systems, Measurement, and Control
identifier doi10.1115/1.4032875
journal fristpage61009
journal lastpage61009
identifier eissn1528-9028
treeJournal of Dynamic Systems, Measurement, and Control:;2016:;volume( 138 ):;issue: 006
contenttypeFulltext


Files in this item

Thumbnail

This item appears in the following Collection(s)

Show simple item record