3 papers
cs.CV2026
Vision Verification Enhanced Fusion of VLMs for Efficient Visual Reasoning
Selim Furkan Tekin, Yichang Xu, Gaowen Liu +3
With the growing number and diversity of Vision-Language Models (VLMs), many works explore language-based ensemble, collaboration, and routing techniques across multiple VLMs to im…
cs.CV2026
VGGT-World: Transforming VGGT into an Autoregressive Geometry World Model
Xiangyu Sun, Shijie Wang, Fengyi Zhang +5
World models that forecast scene evolution by generating future video frames devote the bulk of their capacity to photometric details, yet the resulting predictions often remain ge…
cs.CV2025
Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router
Yubo Huang, Weiqiang Wang, Sirui Zhao +3
Recent years have witnessed remarkable advances in audio-driven talking head generation. However, existing approaches predominantly focus on single-character scenarios. While some…