State Advantage Weighting for Offline RL

October 09, 2022 ยท Declared Dead ยท ๐Ÿ› Tiny Papers @ ICLR

๐Ÿ‘ป CAUSE OF DEATH: Ghosted
No code link whatsoever

"No code URL or promise found in abstract"

Evidence collected by the PWNC Scanner

Authors Jiafei Lyu, Aicheng Gong, Le Wan, Zongqing Lu, Xiu Li arXiv ID 2210.04251 Category cs.LG: Machine Learning Cross-listed cs.AI Citations 9 Venue Tiny Papers @ ICLR Last Checked 5 months ago
Abstract
We present state advantage weighting for offline reinforcement learning (RL). In contrast to action advantage $A(s,a)$ that we commonly adopt in QSA learning, we leverage state advantage $A(s,s^\prime)$ and QSS learning for offline RL, hence decoupling the action from values. We expect the agent can get to the high-reward state and the action is determined by how the agent can get to that corresponding state. Experiments on D4RL datasets show that our proposed method can achieve remarkable performance against the common baselines. Furthermore, our method shows good generalization capability when transferring from offline to online.
Community shame:
Not yet rated
Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

๐Ÿ“œ Similar Papers

In the same crypt โ€” Machine Learning

Died the same way โ€” ๐Ÿ‘ป Ghosted