Dynamic rubrics + closed-form credit redistribution lets agents learn from trajectory-level feedback on long-horizon tasks without verifiers, outperforming both sparse rewards and static rubric approaches.
DRACO improves how AI agents learn from long-horizon tasks by dynamically generating evaluation rubrics during training and redistributing trajectory-level scores back to individual steps.