Using a privileged teacher (trained on high-level inputs like maps and bounding boxes) to supervise a camera-based student during closed-loop fine-tuning is 1000× more sample-efficient than direct RL post-training for autonomous driving.
OPTED improves autonomous driving policies by using a privileged teacher trained with reinforcement learning to guide a camera-based student model during closed-loop fine-tuning. This approach avoids expensive direct RL training in simulation while keeping the policy close to human demonstrations, achieving 1.6-9.5× improvements in driving performance.