1 paper · 1 filter
Louis Mahon, Mirella Lapata
The proliferation of creative video content has driven demand for adapting language models to handle video input and enable multimodal understanding. However, end-to-end models str…