1 paper · 1 filter
Alibay Osmanli, Zixu Cheng, Shaogang Gong
Physical video understanding requires more than naming an event correctly. A model can answer a question about pouring, sliding, or collision from textual regularities while still…