Reward Structure Shapes the Interaction Between Episodic Exploration and Neural Memory in Reinforcement Learning
By Jai Malegaonkar · Paper · cs.LG
In partially observable reinforcement learning, agents face a dual bottleneck: they must explore to encounter rewarding states and retain that experience in memory to optimize their policies. Exploration bonuses and memory architectures are traditionally evaluated in isolation, l