| contributor author | Xu, Lingyi | |
| contributor author | López Muro, Juan | |
| contributor author | Gajić, Zoran | |
| date accessioned | 2026-08-23T07:58:28Z | |
| date available | 2026-08-23T07:58:28Z | |
| date copyright | 2026/01/01 | |
| date issued | 2026 | |
| identifier issn | 2689-6117 | |
| identifier other | aldsc-25-1045.pdf | |
| identifier uri | http://yetl.yabesh.ir/yetl1/handle/yetl/4315884 | |
| description abstract | Abstract. In this article, we present a new policy iteration approach to solve the linear-quadratic (LQ) optimal control problem for infinite horizon discrete-time dynamic systems under the assumption that the system input matrix is known and the state variables are accessible. The adaptive dynamic programming (ADP) methodology has been applied and widely used in many areas on reinforcement learning of corresponding continuous-time problems. However, the built-in constraint has limited its generalization to discrete-time problems as the prior knowledge of the system matrix is required. With our newly proposed method, the system state matrix can be recovered from the iterative reinforcement learning procedure so that the newly developed partial model-free policy iteration approach can also serve as a linear system identification technique. A realistic data-driven example, based on an online tracking and planning navigation mathematical model of a nonholonomic tracked vehicle under a leader–follower formation, is included to demonstrate the efficiency and robustness of the newly proposed methodology. | |
| publisher | The American Society of Mechanical Engineers (ASME) | |
| title | A New Adaptive Dynamic Programming Reinforcement Learning Method for Discrete-Time Linear-Quadratic Optimal Control Problems | |
| type | Journal Paper | |
| journal volume | 6 | |
| journal issue | 1 | |
| journal title | ASME Letters in Dynamic Systems and Control | |
| identifier doi | 10.1115/1.4069345 | |
| journal fristpage | 140 | |
| journal lastpage | 153 | |
| page | 14 | |
| tree | ASME Letters in Dynamic Systems and Control:;2026:;volume( 006 ):;issue:001 | |
| contenttype | Fulltext | |