1 paper
Yuxuan He, Chaiming Huang, Yifan Wu +4
A short video succeeds not simply because of what it shows, but because of how it schedules attention -- yet current multimodal models lack the structural grammar to parse or produ…