1 paper
Mohamed Eltahir, Ali Habibullah, Yazan Alshoibi +3
Extending language models to video introduces two challenges: representation, where existing methods rely on lossy approximations, and long-context, where caption- or agent-based p…