| contributor author | Xu, Lingyi | |
| contributor author | Gajic, Zoran | |
| date accessioned | 2026-08-23T08:29:11Z | |
| date available | 2026-08-23T08:29:11Z | |
| date copyright | 2026/01/01 | |
| date issued | 2026 | |
| identifier issn | 0022-0434 | |
| identifier other | ds-24-1282.pdf | |
| identifier uri | http://yetl.yabesh.ir/yetl1/handle/yetl/4316618 | |
| description abstract | Abstract. The paper demonstrates via simulation that the well-known value iteration algorithm of reinforcement learning for the discrete-time linear-quadratic optimal control problem converges very slowly—at most linearly. Despite its slow convergence, the value iteration algorithm still converges even when the initial feedback gain is several orders of magnitude away from the optimal one, and even when the initial feedback gain is not stabilizing, as demonstrated by an example. It is known that the convergence rate of the corresponding policy iteration algorithm is quadratic, assuming the initial feedback gain is stabilizing. We show that the convergence speed of the value iteration algorithm can also be made quadratic by applying ideas from the doubling algorithm used for solving the algebraic Riccati equation. We precisely state a condition required for convergence of the value iteration algorithm, which turns out to be milder than the corresponding condition for the policy iteration algorithm. In addition, we show that the newly proposed value iteration algorithm requires less computational effort than the policy iteration algorithm. With these improvements and observations, we revitalize the value iteration algorithm and demonstrate its superiority over the policy iteration algorithm. | |
| publisher | The American Society of Mechanical Engineers (ASME) | |
| title | Improving Reinforcement Learning Value Iterations in Discrete-Time Linear-Quadratic Optimal Control | |
| type | Journal Paper | |
| journal volume | 148 | |
| journal issue | 1 | |
| journal title | Journal of Dynamic Systems, Measurement, and Control | |
| identifier doi | 10.1115/1.4069531 | |
| journal fristpage | 19 | |
| journal lastpage | 22 | |
| page | 4 | |
| tree | Journal of Dynamic Systems, Measurement, and Control:;2026:;volume( 148 ):;issue:001 | |
| contenttype | Fulltext | |