2 papers
cs.CV2026
GridProbe: Posterior-Probing for Adaptive Test-Time Compute in Long-Video VLMs
Mohamed Eltahir, Lama Ayash, Ali Habibullah +2
Long-video understanding in VLMs is bottlenecked by a single monolithic forward pass over thousands of frames at quadratic attention cost. A common mitigation is to first select a…
cs.CV2026
VideoAtlas: Navigating Long-Form Video in Logarithmic Compute
Mohamed Eltahir, Ali Habibullah, Yazan Alshoibi +3
Extending language models to video introduces two challenges: representation, where existing methods rely on lossy approximations, and long-context, where caption- or agent-based p…