Feed-forward generative models can reconstruct complex hand-object interactions faster and more reliably than optimization-based methods by learning to correct foundation model errors while respecting physical constraints during generation.
This paper presents 4D-HOF, a fast method for reconstructing 3D hand and object positions/orientations from video. Instead of slow per-video optimization, it uses a generative model trained on diverse data to refine rough estimates from vision models.