By splitting scene-aware interaction generation into affordance prediction and motion synthesis stages, you can train on unpaired datasets and avoid the expensive requirement for full human-object-scene annotations while still generating physically plausible interactions.
This paper presents MAMHOI, a method for generating realistic human-object interactions in 3D scenes by factorizing the problem into two stages: first predicting where interactions can feasibly happen using scene understanding, then generating realistic human-object motion conditioned on those affordances.