A Framework for Designing Reward Functions: From Objectives to Features to Human-Aligned Reward Functions
By Di Yang Shi · Paper · cs.LG
We present a formal process to enable non-experts to instantiate and iterate on human-aligned reward functions, i.e. reward functions that adhere to a given preference ordering over trajectories. Given a task described in natural language, our process produces a linear reward fun