4 papers · 1 filter
Inference Compute-Optimal Video Vision Language Models
Peiqi Wang, ShengYun Peng, Xuewen Zhang +5
This work investigates the optimal allocation of inference compute across three key scaling factors in video vision language models: language model size, frame count, and the numbe…
CompCap: Improving Multimodal Large Language Models with Composite Captions
Xiaohui Chen, Satya Narayan Shukla, Mahmoud Azab +8
How well can Multimodal Large Language Models (MLLMs) understand composite images? Composite images (CIs) are synthetic visuals created by merging multiple visual elements, such as…
MMViT: Multiscale Multiview Vision Transformers
Yuchen Liu, Natasha Ong, Kaiyan Peng +8
We present Multiscale Multiview Vision Transformers (MMViT), which introduces multiscale feature maps and multiview encodings to transformer models. Our model encodes different vie…
SVT: Supertoken Video Transformer for Efficient Video Understanding
Chenbin Pan, Rui Hou, Hanchao Yu +3
Whether by processing videos with fixed resolution from start to end or incorporating pooling and down-scaling strategies, existing video transformers process the whole video conte…