Optimal Control with Learning on the Fly from Finite to Infinite-Dimensional Systems

Feifei Miao

Chinese Annals of Mathematics, Series B ›› : 1 -22.

PDF
Chinese Annals of Mathematics, Series B ›› :1 -22. DOI: 10.1007/s11401-026-0027-6
Article
research-article
Optimal Control with Learning on the Fly from Finite to Infinite-Dimensional Systems
Author information +
History +
PDF

Abstract

The main aim of this paper is to extend the Bayesian approach to finding quadratic optimal control to a wider range of stochastic linear systems. These systems involve an unknown parameter in the drift term, which is observed through a noisy linear channel. The author demonstrates the effectiveness of the Bayesian strategy by comparing the cost it incurs with that of an optimal control that possesses complete knowledge of the parameter. The findings reveal that the Bayesian strategy minimizes the worst-case multiplicative regret. Furthermore, the author provides a proof that the corresponding adaptive scheme is optimal when the unknown system parameter belongs to an infinite space. This result further validates the effectiveness of the Bayesian strategy in handling systems with unknown parameters. In summary, this paper contributes to the generalization of the Bayesian strategy for the optimal control in stochastic linear systems with unknown parameters. It also establishes a theoretical basis for its optimality.

Keywords

Bayesian strategy / Kalman-Bucy filter / Optimal control / Adaptive control / Multiplicative regret / 93C40 / 93C41 / 93E11

Cite this article

Download citation ▾
Feifei Miao. Optimal Control with Learning on the Fly from Finite to Infinite-Dimensional Systems. Chinese Annals of Mathematics, Series B 1-22 DOI:10.1007/s11401-026-0027-6

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Abbasi-Yadkori Y, Szepesvári C. Regret bounds for the adaptive control of linear quadratic systems, Proceedings of the 24th Annual Conference on Learning Theory. J. Math. Learn. Res. Proc. Track., 2011, 19: 1-26

[2]

Agarwal N, Bullins B, Hazan E, et al. . Online control with adversarial disturbances. 36th International Conference on Machine Learning, 2019, California, IMLS: 154-165

[3]

Agarwal N, Hazan E, Singh K. Logarithmic regret for online control, 2019, Canada, Curran Associates: 10175-1018432

[4]

Anderson B, Moore J. Optimal Control: Linear Quadratic Methods, 1989, NJ, Prentice Hall

[5]

Auer P, Cesa-Bianchi N, Fischer P. Finite-time analysis of the multiarmed bandit problem. Machine learning, 2002, 47: 235-256

[6]

Bensoussan A. Stochastic Control of Partially Observable Systems, 1992, Cambridge, Cambridge University Press

[7]

Bertsekas D P. Dynamic Programming and Optimal Control, 2005, Belmont, MA, Athena Scientific Vol. 1

[8]

Bubeck S, Cesa-Bianchi N. Regret analysis of stochastic and nonstochastic multi-armed bandit problems. Found. Trends in Mach. Learn., 2012, 5(1): 101-112

[9]

Carruth J, Eggl M F, Fefferman C, et al. . Controlling unknown linear dynamics with bounded multiplicative regret. Rev. Mat. Iberoam., 2022, 38(7): 2185-2216

[10]

Cassel A, Cohen A, Koren T. Logarithmic regret for learning linear quadratic regulators efficiently. Proceedings of the 37th International Conference on Machine Learning, 2020, New York, JMLR: 1328-1337 119

[11]

Cesa-Bianchi N, Lugosi G. Prediction, Learning, and Games, 2006, Cambridge, Cambridge University Press

[12]

Cohen A, Koren T, Mansour Y. Learning linear-quadratic regulators efficiently with only T\documentclass[12pt]{minimal}\usepackage{amsmath}\usepackage{wasysym}\usepackage{amsfonts}\usepackage{amssymb}\usepackage{amsbsy}\usepackage{mathrsfs}\usepackage{upgreek}\setlength{\oddsidemargin}{-69pt}\begin{document}$${\sqrt T}$$\end{document} Regret. International Conference on Machine Learning, 2019: 1300-1309 97

[13]

Curtain R F. Infinite-dimensional filtering. SIAM J. Control Optim., 1975, 13(1): 89-104

[14]

Curtain R F, Ichikawa A. The separation principle for Stochastic evolution equations. SIAM J. Control Optim., 1977, 15(3): 367-383

[15]

Dean S, Mania H, Matni N, et al. . Regret bounds for robust adaptive control of the linear quadratic regulator, 2018, New York, Curran Associates, Inc.: 4192-420131

[16]

Fefferman C, Pegueroles B G, Rowley C W, Weber M. Optimal control with learning on the fly: A toy problem. Rev. Mat. Iberoam., 2022, 38(1): 175-187

[17]

Georgiou T T, Lindquist A. The separation principle in stochastic control, redux. IEEE Trans. Automat. Control, 2013, 58(10): 2481-2494

[18]

Gurevich D, Goswami D, Fefferman C L, Rowley C W. Optimal control with learning on the fly: System with unknown drift. Learning for Dynamics and Control Conference, 2022, California, PMLR: 870-880 168

[19]

Hazan E. Introduction to Online Convex Optimization, 2022, California, MIT Press

[20]

Ibrahimi M, Javanmard A, Roy B. Efficient reinforcement learning for high dimensional linear quadratic systems, 201225

[21]

Ichikawa A. Dynamic programming approach to stochastic evolution equations. SIAM J. Control Optim., 1979, 17(1): 152-174

[22]

Lai T L, Robbins H. Asymptotically efficient adaptive allocation rules. Adv. Appl. Math., 1985, 6(1): 4-22

[23]

Lindquist A. On feedback control of linear stochastic systems. SIAM J. Control, 1973, 11(2): 323-343

[24]

Q, Zhang X. Mathematical Control Theory for Stochastic Partial Differential Equations, 2021, Charm, Springer-Verlag

[25]

Powell W B. Approximate Dynamic Programming: Solving the Curses of Dimensionality, 2007, Hoboken, NJ, John Wiley & Sons

[26]

Speyer J L, Chung W H. Stochastic Processes, Estimation, and Control, 2008, PA, Philadelphia, SIAM

[27]

Yong J, Zhou X Y. Stochastic Controls: Hamiltonian Systems and HJB Equations, 1999, New York, Springer-Verlag

RIGHTS & PERMISSIONS

The Editorial Office of CAM and Springer-Verlag Berlin Heidelberg

PDF

0

Accesses

0

Citation

Detail

Sections
Recommended

/