By explicitly modeling contact forces alongside visual and kinematic information, this system achieves 82% success on precision assembly tasks—a massive jump from 15% for existing approaches—showing that understanding physical interaction is critical for real-world robot manipulation.
Facet-0 is a robotic foundation model that learns to perform precise assembly tasks by predicting how its actions will affect contact forces with objects. It combines vision, language, and force feedback to generate robot movements that maintain sub-millimeter accuracy during assembly, trained on 1,000 hours of real robot data across multiple manufacturing setups.