By processing multiple sparse observations together rather than single images, and training with a motion-range supervision objective, the model better understands articulation without needing large labeled datasets—it generates synthetic training data procedurally.
FAMOS is a feed-forward model that predicts how articulated objects (like doors, drawers) move and which parts are movable from multiple sparse 3D views.