Reward-Balancing for Statistical Spoken Dialogue Systems using Multi-objective Reinforcement Learning

July 19, 2017 ยท Declared Dead ยท ๐Ÿ› SIGDIAL Conference

๐Ÿ‘ป CAUSE OF DEATH: Ghosted
No code link whatsoever

"No code URL or promise found in abstract"

Evidence collected by the PWNC Scanner

Authors Stefan Ultes, Paweล‚ Budzianowski, Iรฑigo Casanueva, Nikola Mrkลกiฤ‡, Lina Rojas-Barahona, Pei-Hao Su, Tsung-Hsien Wen, Milica Gaลกiฤ‡, Steve Young arXiv ID 1707.06299 Category cs.CL: Computation & Language Cross-listed stat.ML Citations 12 Venue SIGDIAL Conference Last Checked 5 months ago
Abstract
Reinforcement learning is widely used for dialogue policy optimization where the reward function often consists of more than one component, e.g., the dialogue success and the dialogue length. In this work, we propose a structured method for finding a good balance between these components by searching for the optimal reward component weighting. To render this search feasible, we use multi-objective reinforcement learning to significantly reduce the number of training dialogues required. We apply our proposed method to find optimized component weights for six domains and compare them to a default baseline.
Community shame:
Not yet rated
Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

๐Ÿ“œ Similar Papers

In the same crypt โ€” Computation & Language

๐ŸŒ… ๐ŸŒ… Old Age

Attention Is All You Need

Ashish Vaswani, Noam Shazeer, ... (+6 more)

cs.CL ๐Ÿ› NeurIPS ๐Ÿ“š 166.0K cites 9 years ago

Died the same way โ€” ๐Ÿ‘ป Ghosted