1 paper
Chika Maduabuchi, Jindong Wang
Current text-to-video models can make individual frames look convincing while still getting simple interactions wrong: objects move before contact, an intended action is skipped, a…