首页> 外国专利> TRAINING ACTION SELECTION NEURAL NETWORKS

TRAINING ACTION SELECTION NEURAL NETWORKS

机译：培训行动选择神经网络

页面导航

摘要
著录项
相似文献

摘要

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for training a policy neural network. The policy neural network is used to select actions to be performed by an agent that interacts with an environment by receiving an observation characterizing a state of the environment and performing an action from a set of actions in response to the received observation. A trajectory is obtained from a replay memory, and a final update to current values of the policy network parameters is determined for each training observation in the trajectory. The final updates to the current values of the policy network parameters are determined from selected action updates and leave-one-out updates.

机译：方法，系统和设备，包括在计算机存储介质上编码的计算机程序，用于训练策略神经网络。策略神经网络用于通过接收表征环境状态的观察和响应于接收到的观察来执行与环境的观察来选择与环境交互的代理执行的动作。从重放存储器获得轨迹，并且针对轨迹中的每个训练观察确定策略网络参数的当前值的最终更新。对策略网络参数的当前值的最终更新是根据所选操作更新和休留一次更新的。

著录项

公开/公告号US2021110271A1

专利类型
公开/公告日2021-04-15

原文格式PDF
申请/专利权人 DEEPMIND TECHNOLOGIES LIMITED;
展开▼

申请/专利号US201816603307
发明设计人 MARC GENDRON-BELLEMARE;MOHAMMAD GHESHLAGHI AZAR;AUDRUNAS GRUSLYS;REMI MUNOS;
展开▼

申请日2018-06-11
分类号G06N3/08;G06N3/04;
国家 US
入库时间 2022-08-24 18:14:06

相似文献

专利
外文文献
中文文献