Policy Optimization Provably Converges to Nash Equilibria in Zero-Sum Linear Quadratic Games

机译：策略优化可否在零和线性二次游戏中纳入纳什均衡融合

获取原文

页面导航

摘要
著录项
相似文献
相关主题

摘要

We study the global convergence of policy optimization for finding the Nash equilibria (NE) in zero-sum linear quadratic (LQ) games. To this end, we first investigate the landscape of LQ games, viewing it as a nonconvex-nonconcave saddle-point problem in the policy space. Specifically, we show that despite its nonconvexity and nonconcavity, zero-sum LQ games have the property that the stationary point of the objective function with respect to the linear feedback control policies constitutes the NE of the game. Building upon this, we develop three projected nested-gradient methods that are guaranteed to converge to the NE of the game. Moreover, we show that all these algorithms enjoy both globally sublinear and locally linear convergence rates. Simulation results are also provided to illustrate the satisfactory convergence properties of the algorithms. To the best of our knowledge, this work appears to be the first one to investigate the optimization landscape of LQ games, and provably show the convergence of policy optimization methods to the NE. Our work serves as an initial step toward understanding the theoretical aspects of policy-based reinforcement learning algorithms for zero-sum Markov games in general.

机译：我们研究了在零和线性二次（LQ）游戏中找到了NASH均衡（NE）的政策优化的全球融合。为此，我们首先调查LQ游戏的景观，将其视为策略空间中的非谐波 - 非传播马鞍点问题。具体而言，我们表明，尽管它的非凸起和非扫护性，但零和LQ游戏具有目标函数相对于线性反馈控制策略的静止点构成游戏的NE。建立在此处，我们开发了三种预定的嵌套梯度方法，保证将其融合到游戏的NE。此外，我们表明所有这些算法都享有全局载位和局部线性收敛速率。还提供了模拟结果以说明算法的令人满意的收敛性。据我们所知，这项工作似乎是第一个调查LQ游戏的优化景观的作品，并证明了对NE的政策优化方法的融合。我们的工作是了解零汇率Markov游戏的基于策略的强化学习算法的理论方面的最初步骤。

著录项

来源
《Conference on Neural Information Processing Systems》|2020年|p11149-11940|共13页
会议地点
作者
Kaiqing Zhang; Zhuoran Yang; Tamer Basar;
展开▼
作者单位

展开▼
会议组织
原文格式 PDF
正文语种
中图分类计量学;
关键词

相似文献

外文文献
中文文献
专利

1. Distributed convergence to Nash equilibria in two-network zero-sum games [J] . B. Gharesifard, J. Cortes Automatica . 2013,第6期

机译：两网零和博弈中的纳什均衡分布收敛
2. Discontinuous Nash Equilibria in a Two-Stage Linear-Quadratic Dynamic Game With Linear Constraints [J] . Rajani Singh, Agnieszka Wiszniewska-Matyszkiel IEEE Transactions on Automatic Control . 2019,第7期

机译：具有线性约束的两阶段线性二次动态博弈中的不连续Nash平衡
3. Linear Quadratic Mean Field Games:Decentralized O(1/N)-Nash Equilibria [J] . HUANG Minyi, YANG Xuwei 系统科学与复杂性：英文版 . 2021,第005期

机译：线性二次平均野外游戏：分散o（1 / n）-Nash均衡
4. Policy Optimization Provably Converges to Nash Equilibria in Zero-Sum Linear Quadratic Games [C] . Kaiqing Zhang, Zhuoran Yang, Tamer Basar Conference on Neural Information Processing Systems . 2020

机译：策略优化可否在零和线性二次游戏中纳入纳什均衡融合
5. Nash strategies for dynamic noncooperative linear quadratic sequential games [D] . Shen, Dan 2006

机译：动态非合作式线性二次顺序博弈的纳什策略
6. From Nash Equilibria to Chain Recurrent Sets: An Algorithmic Solution Concept for Game Theory [O] . Christos Papadimitriou, Georgios Piliouras 2018

机译：从纳什均衡到链复发集：游戏理论的算法解决方案概念
7. Distributed Convergence to Nash Equilibria in Two-Network Zero-Sum Games [O] . B. Gharesifard, J. Cortes 2013

机译：在双网络零和游戏中分布到纳什均衡的融合
8. Distributed Convergence to Nash Equilibria in Two-Network Zero-Sum Games. [R] . Gharesifard, B., Cortes, J. 2013

机译：双网零和博弈中纳什均衡的分布收敛性。

Policy Optimization Provably Converges to Nash Equilibria in Zero-Sum Linear Quadratic Games

摘要

著录项

相似文献

相关主题

期刊订阅