3 papers
cs.CV2026
VeRVE: Versatile Retrieval for Videos via Unified Embeddings
Shaunak Halbe, Bhagyashree Puranik, Jayakrishnan Unnikrishnan +3
Modern video retrieval systems are expected to handle diverse tasks ranging from corpus-level retrieval, fine-grained moment localization to flexible multimodal querying. Specializ…
cs.CV2026
One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding
Zheyu Zhang, Ziqi Pang, Shixing Chen +3
Long video understanding is inherently challenging for vision-language models (VLMs) because of the extensive number of frames. With each video frame typically expanding into tens…
cs.CV2025
What Happens Next? Next Scene Prediction with a Unified Video Model
Xinjie Li, Zhimin Chen, Rui Zhao +3
Recent unified models for joint understanding and generation have significantly advanced visual generation capabilities. However, their focus on conventional tasks like text-to-vid…