1 paper
Rahul Jain, Mayank Patel, Asim Unmesh +1
Text- and image-conditioned video generation models have achieved strong visual fidelity and temporal coherence, but they often fail to generate motion governed by kinematic and ge…