Hierarchical Approaches for Reinforcement Learning in Parameterized Action Space

October 23, 2018 ยท Declared Dead ยท ๐Ÿ› AAAI Spring Symposia

๐Ÿ‘ป CAUSE OF DEATH: Ghosted
No code link whatsoever

"No code URL or promise found in abstract"

Evidence collected by the PWNC Scanner

Authors Ermo Wei, Drew Wicke, Sean Luke arXiv ID 1810.09656 Category cs.LG: Machine Learning Cross-listed cs.AI, stat.ML Citations 36 Venue AAAI Spring Symposia Last Checked 5 months ago
Abstract
We explore Deep Reinforcement Learning in a parameterized action space. Specifically, we investigate how to achieve sample-efficient end-to-end training in these tasks. We propose a new compact architecture for the tasks where the parameter policy is conditioned on the output of the discrete action policy. We also propose two new methods based on the state-of-the-art algorithms Trust Region Policy Optimization (TRPO) and Stochastic Value Gradient (SVG) to train such an architecture. We demonstrate that these methods outperform the state of the art method, Parameterized Action DDPG, on test domains.
Community shame:
Not yet rated
Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

๐Ÿ“œ Similar Papers

In the same crypt โ€” Machine Learning

Died the same way โ€” ๐Ÿ‘ป Ghosted