1 paper
Weitong Kong, Di Wen, Kunyu Peng +10
Correcting errors in long-video understanding is disproportionately costly: existing multimodal pipelines produce opaque, end-to-end outputs that expose no intermediate state for i…