A unified value function approach enables safe RL without pre-built safety filters or average-case guarantees—the learned policy maintains strict safety constraints while maximizing task performance.
This paper proposes a new Bellman operator that combines task performance and safety constraints into a single value function for reinforcement learning. Unlike existing methods that either require pre-built safety guarantees or only ensure safety on average, this approach learns both objectives jointly and provably maintains safety at all times during execution.