首页> 外文会议>Automatic Speech Recognition amp; Understanding, 2009. ASRU 2009 >The exploration/exploitation trade-off in Reinforcement Learning for dialogue management

【24h】

The exploration/exploitation trade-off in Reinforcement Learning for dialogue management

机译：对话学习中强化学习中的探索/开发权衡

获取原文

页面导航

摘要
著录项
相似文献
相关主题

摘要

Conversational systems use deterministic rules that trigger actions such as requests for confirmation or clarification. More recently, Reinforcement Learning and (Partially Observable) Markov Decision Processes have been proposed for this task. In this paper, we investigate action selection strategies for dialogue management, in particular the exploration/exploitation trade-off and its impact on final reward (i.e. the session reward after optimization has ended) and lifetime reward (i.e. the overall reward accumulated over the learner's lifetime). We propose to use interleaved exploitation sessions as a learning methodology to assess the reward obtained from the current policy. The experiments show a statistically significant difference in final reward of exploitation-only sessions between a system that optimizes lifetime reward and one that maximizes the reward of the final policy.

机译：会话系统使用确定性规则来触发诸如确认或澄清请求之类的操作。最近，针对此任务提出了强化学习和（部分可观察到的）马尔可夫决策过程。在本文中，我们研究了对话管理的行动选择策略，特别是探索/开发权衡及其对最终奖励（即优化结束后的会话奖励）和终生奖励（即学习者积累的总奖励）的影响。一生）。我们建议使用交错式开发会话作为一种学习方法，以评估从当前政策中获得的回报。实验显示，在优化终身奖励的系统和最大化最终政策的奖励的系统之间，仅利用会话的最终奖励在统计上有显着差异。

著录项

来源
《Automatic Speech Recognition amp; Understanding, 2009. ASRU 2009 》|2009年|479-484|共6页
会议地点 Merano(IT);Merano(IT)
作者
Varges Sebastian; Riccardi Giuseppe; Quarteroni Silvia; Ivanov Alexei V.;
展开▼
作者单位

Department of Information Engineering and Computer Science, University of Trento, 38050 Povo di Trento, Italy;

展开▼
会议组织
原文格式 PDF
正文语种
中图分类
关键词

相似文献

外文文献
中文文献
专利

1. Exploration and exploitation balance management in fuzzy reinforcement learning [J] . Vali Derhami, Vahid Johari Majd, Majid Nili Ahmadabadi Fuzzy sets and systems . 2010 ,第4期

机译：模糊强化学习中的勘探与开发平衡管理
2. Learning Exploration/Exploitation Strategies for Single Trajectory Reinforcement Learning [J] . Damien Ernst, Francis Maes, Michael Castronovo, JMLR: Workshop and Conference Proceedings . 2012 ,第2012期

机译：单轨强化学习的学习探索/开发策略
3. Learning and innovation: Exploitation and exploration trade-offs [J] . Changsu Kim, Jaeyong Song, Atul Nerkar Journal of Business Research . 2012 ,第8期

机译：学习与创新：开发与探索的权衡
4. The Exploration/Exploitation Trade-off in Reinforcement Learning for Dialogue Management [C] . Sebastian Varges, Giuseppe Riccardi, Silvia Quarteroni, IEEE Workshop on Automatic Speech Recognition Understanding . 2009

机译：对话管理加固学习的探索/剥削权衡
5. Min-Max Inverse Reinforcement Learning for Learning Bi-Modal Dialogue Policies [D] . Patil, Gandharv. 2020

机译：用于学习双模对话策略的最大最大逆钢筋学习
6. The implied exploration-exploitation trade-off in human motor learning [O] . Holly N Phillips, Nikhil A Howai, Guy-Bart V Stan, 2011

机译：人类运动学习中隐含的探索与开发权衡
7. The Exploration/Exploitation Trade-off in Reinforcement Learning for Dialogue Management [O] . Sebastian Varges, Silvia Quarteroni, Alexei V. Ivanov 2010

机译：对话管理中强化学习中的探索/开发权衡

The exploration/exploitation trade-off in Reinforcement Learning for dialogue management

摘要

著录项

相似文献

相关主题

期刊订阅