Motion-based reasoning (tracking how bodies move) is more efficient and generalizable than pose-based reasoning for incident detection, and knowledge distillation can compress this understanding into models small enough for real-world deployment.
This paper tackles classroom safety monitoring using privacy-preserving computer vision. The authors create a hybrid benchmark mixing synthetic and real classroom data, then propose a lightweight motion-reasoning model that captures how incidents differ in movement patterns (speed, direction, acceleration) rather than just body poses.