PDF
Abstract
The main aim of this paper is to extend the Bayesian approach to finding quadratic optimal control to a wider range of stochastic linear systems. These systems involve an unknown parameter in the drift term, which is observed through a noisy linear channel. The author demonstrates the effectiveness of the Bayesian strategy by comparing the cost it incurs with that of an optimal control that possesses complete knowledge of the parameter. The findings reveal that the Bayesian strategy minimizes the worst-case multiplicative regret. Furthermore, the author provides a proof that the corresponding adaptive scheme is optimal when the unknown system parameter belongs to an infinite space. This result further validates the effectiveness of the Bayesian strategy in handling systems with unknown parameters. In summary, this paper contributes to the generalization of the Bayesian strategy for the optimal control in stochastic linear systems with unknown parameters. It also establishes a theoretical basis for its optimality.
Keywords
Bayesian strategy
/
Kalman-Bucy filter
/
Optimal control
/
Adaptive control
/
Multiplicative regret
/
93C40
/
93C41
/
93E11
Cite this article
Download citation ▾
Feifei Miao.
Optimal Control with Learning on the Fly from Finite to Infinite-Dimensional Systems.
Chinese Annals of Mathematics, Series B 1-22 DOI:10.1007/s11401-026-0027-6
| [1] |
Abbasi-Yadkori Y, Szepesvári C. Regret bounds for the adaptive control of linear quadratic systems, Proceedings of the 24th Annual Conference on Learning Theory. J. Math. Learn. Res. Proc. Track., 2011, 19: 1-26
|
| [2] |
Agarwal N, Bullins B, Hazan E, et al. . Online control with adversarial disturbances. 36th International Conference on Machine Learning, 2019, California, IMLS: 154-165
|
| [3] |
Agarwal N, Hazan E, Singh K. Logarithmic regret for online control, 2019, Canada, Curran Associates: 10175-1018432
|
| [4] |
Anderson B, Moore J. Optimal Control: Linear Quadratic Methods, 1989, NJ, Prentice Hall
|
| [5] |
Auer P, Cesa-Bianchi N, Fischer P. Finite-time analysis of the multiarmed bandit problem. Machine learning, 2002, 47: 235-256
|
| [6] |
Bensoussan A. Stochastic Control of Partially Observable Systems, 1992, Cambridge, Cambridge University Press
|
| [7] |
Bertsekas D P. Dynamic Programming and Optimal Control, 2005, Belmont, MA, Athena Scientific Vol. 1
|
| [8] |
Bubeck S, Cesa-Bianchi N. Regret analysis of stochastic and nonstochastic multi-armed bandit problems. Found. Trends in Mach. Learn., 2012, 5(1): 101-112
|
| [9] |
Carruth J, Eggl M F, Fefferman C, et al. . Controlling unknown linear dynamics with bounded multiplicative regret. Rev. Mat. Iberoam., 2022, 38(7): 2185-2216
|
| [10] |
Cassel A, Cohen A, Koren T. Logarithmic regret for learning linear quadratic regulators efficiently. Proceedings of the 37th International Conference on Machine Learning, 2020, New York, JMLR: 1328-1337 119
|
| [11] |
Cesa-Bianchi N, Lugosi G. Prediction, Learning, and Games, 2006, Cambridge, Cambridge University Press
|
| [12] |
Cohen A, Koren T, Mansour Y. Learning linear-quadratic regulators efficiently with only T\documentclass[12pt]{minimal}\usepackage{amsmath}\usepackage{wasysym}\usepackage{amsfonts}\usepackage{amssymb}\usepackage{amsbsy}\usepackage{mathrsfs}\usepackage{upgreek}\setlength{\oddsidemargin}{-69pt}\begin{document}$${\sqrt T}$$\end{document} Regret. International Conference on Machine Learning, 2019: 1300-1309 97
|
| [13] |
Curtain R F. Infinite-dimensional filtering. SIAM J. Control Optim., 1975, 13(1): 89-104
|
| [14] |
Curtain R F, Ichikawa A. The separation principle for Stochastic evolution equations. SIAM J. Control Optim., 1977, 15(3): 367-383
|
| [15] |
Dean S, Mania H, Matni N, et al. . Regret bounds for robust adaptive control of the linear quadratic regulator, 2018, New York, Curran Associates, Inc.: 4192-420131
|
| [16] |
Fefferman C, Pegueroles B G, Rowley C W, Weber M. Optimal control with learning on the fly: A toy problem. Rev. Mat. Iberoam., 2022, 38(1): 175-187
|
| [17] |
Georgiou T T, Lindquist A. The separation principle in stochastic control, redux. IEEE Trans. Automat. Control, 2013, 58(10): 2481-2494
|
| [18] |
Gurevich D, Goswami D, Fefferman C L, Rowley C W. Optimal control with learning on the fly: System with unknown drift. Learning for Dynamics and Control Conference, 2022, California, PMLR: 870-880 168
|
| [19] |
Hazan E. Introduction to Online Convex Optimization, 2022, California, MIT Press
|
| [20] |
Ibrahimi M, Javanmard A, Roy B. Efficient reinforcement learning for high dimensional linear quadratic systems, 201225
|
| [21] |
Ichikawa A. Dynamic programming approach to stochastic evolution equations. SIAM J. Control Optim., 1979, 17(1): 152-174
|
| [22] |
Lai T L, Robbins H. Asymptotically efficient adaptive allocation rules. Adv. Appl. Math., 1985, 6(1): 4-22
|
| [23] |
Lindquist A. On feedback control of linear stochastic systems. SIAM J. Control, 1973, 11(2): 323-343
|
| [24] |
Lü Q, Zhang X. Mathematical Control Theory for Stochastic Partial Differential Equations, 2021, Charm, Springer-Verlag
|
| [25] |
Powell W B. Approximate Dynamic Programming: Solving the Curses of Dimensionality, 2007, Hoboken, NJ, John Wiley & Sons
|
| [26] |
Speyer J L, Chung W H. Stochastic Processes, Estimation, and Control, 2008, PA, Philadelphia, SIAM
|
| [27] |
Yong J, Zhou X Y. Stochastic Controls: Hamiltonian Systems and HJB Equations, 1999, New York, Springer-Verlag
|
RIGHTS & PERMISSIONS
The Editorial Office of CAM and Springer-Verlag Berlin Heidelberg
Just Accepted
This article has successfully passed peer review and final editorial review, and will soon enter typesetting, proofreading and other publishing processes. The currently displayed version is the accepted final manuscript. The officially published version will be updated with format, DOI and citation information upon launch. We recommend that you pay attention to subsequent journal notifications and preferentially cite the officially published version. Thank you for your support and cooperation.