YaBeSH Engineering and Technology Library

    • Journals
    • PaperQuest
    • YSE Standards
    • YaBeSH
    • Login
    View Item 
    •   YE&T Library
    • ASME
    • Journal of Dynamic Systems, Measurement, and Control
    • View Item
    •   YE&T Library
    • ASME
    • Journal of Dynamic Systems, Measurement, and Control
    • View Item
    • All Fields
    • Source Title
    • Year
    • Publisher
    • Title
    • Subject
    • Author
    • DOI
    • ISBN
    Advanced Search
    JavaScript is disabled for your browser. Some features of this site may not work without it.

    Archive

    Improving Reinforcement Learning Value Iterations in Discrete-Time Linear-Quadratic Optimal Control

    Source: Journal of Dynamic Systems, Measurement, and Control:;2026:;volume( 148 ):;issue:001::page 19
    Author:
    Xu, Lingyi
    ,
    Gajic, Zoran
    DOI: 10.1115/1.4069531
    Publisher: The American Society of Mechanical Engineers (ASME)
    Abstract: Abstract. The paper demonstrates via simulation that the well-known value iteration algorithm of reinforcement learning for the discrete-time linear-quadratic optimal control problem converges very slowly—at most linearly. Despite its slow convergence, the value iteration algorithm still converges even when the initial feedback gain is several orders of magnitude away from the optimal one, and even when the initial feedback gain is not stabilizing, as demonstrated by an example. It is known that the convergence rate of the corresponding policy iteration algorithm is quadratic, assuming the initial feedback gain is stabilizing. We show that the convergence speed of the value iteration algorithm can also be made quadratic by applying ideas from the doubling algorithm used for solving the algebraic Riccati equation. We precisely state a condition required for convergence of the value iteration algorithm, which turns out to be milder than the corresponding condition for the policy iteration algorithm. In addition, we show that the newly proposed value iteration algorithm requires less computational effort than the policy iteration algorithm. With these improvements and observations, we revitalize the value iteration algorithm and demonstrate its superiority over the policy iteration algorithm.
    • Download: (1.191Mb)
    • Show Full MetaData Hide Full MetaData
    • Get RIS
    • Item Order
    • Go To Publisher
    • Statistics

      Improving Reinforcement Learning Value Iterations in Discrete-Time Linear-Quadratic Optimal Control

    URI
    https://yetl.yabesh.ir/yetl1/handle/yetl/4316618
    Collections
    • Journal of Dynamic Systems, Measurement, and Control

    Show full item record

    contributor authorXu, Lingyi
    contributor authorGajic, Zoran
    date accessioned2026-08-23T08:29:11Z
    date available2026-08-23T08:29:11Z
    date copyright2026/01/01
    date issued2026
    identifier issn0022-0434
    identifier otherds-24-1282.pdf
    identifier urihttp://yetl.yabesh.ir/yetl1/handle/yetl/4316618
    description abstractAbstract. The paper demonstrates via simulation that the well-known value iteration algorithm of reinforcement learning for the discrete-time linear-quadratic optimal control problem converges very slowly—at most linearly. Despite its slow convergence, the value iteration algorithm still converges even when the initial feedback gain is several orders of magnitude away from the optimal one, and even when the initial feedback gain is not stabilizing, as demonstrated by an example. It is known that the convergence rate of the corresponding policy iteration algorithm is quadratic, assuming the initial feedback gain is stabilizing. We show that the convergence speed of the value iteration algorithm can also be made quadratic by applying ideas from the doubling algorithm used for solving the algebraic Riccati equation. We precisely state a condition required for convergence of the value iteration algorithm, which turns out to be milder than the corresponding condition for the policy iteration algorithm. In addition, we show that the newly proposed value iteration algorithm requires less computational effort than the policy iteration algorithm. With these improvements and observations, we revitalize the value iteration algorithm and demonstrate its superiority over the policy iteration algorithm.
    publisherThe American Society of Mechanical Engineers (ASME)
    titleImproving Reinforcement Learning Value Iterations in Discrete-Time Linear-Quadratic Optimal Control
    typeJournal Paper
    journal volume148
    journal issue1
    journal titleJournal of Dynamic Systems, Measurement, and Control
    identifier doi10.1115/1.4069531
    journal fristpage19
    journal lastpage22
    page4
    treeJournal of Dynamic Systems, Measurement, and Control:;2026:;volume( 148 ):;issue:001
    contenttypeFulltext
    DSpace software copyright © 2002-2015  DuraSpace
    نرم افزار کتابخانه دیجیتال "دی اسپیس" فارسی شده توسط یابش برای کتابخانه های ایرانی | تماس با یابش
    yabeshDSpacePersian
     
    DSpace software copyright © 2002-2015  DuraSpace
    نرم افزار کتابخانه دیجیتال "دی اسپیس" فارسی شده توسط یابش برای کتابخانه های ایرانی | تماس با یابش
    yabeshDSpacePersian