9 citations · 9 across the 2 of their papers we have counts for
20 papers · 1 filter
Video Prediction Transformers without Recurrence or Convolution
Yujin Tang, Lu Qi, Xiangtai Li +2
Video prediction has witnessed the emergence of RNN-based models led by ConvLSTM, and CNN-based models led by SimVP. Following the significant success of ViT, recent works have int…
Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark
Haobo Yuan, Yueyi Sun, Yanwei Li +7
Recent advances in Multimodal Large Language Models (MLLMs) have significantly improved performance on tasks such as visual grounding and visual question answering. However, the re…
Rethinking Cross-Generator Image Forgery Detection through DINOv3
Zhenglin Huang, Jason Li, Haiquan Wen +7
As generative models become increasingly diverse and powerful, cross-generator detection has emerged as a new challenge. Existing detection methods often memorize artifacts of spec…
Sa2VA: Marrying SAM2 with MLLM for Dense Grounded Understanding of Images and Videos
Haobo Yuan, Xiangtai Li, Tao Zhang +8
This work presents Sa2VA, the first comprehensive, unified model for dense grounded understanding of both images and videos. Unlike existing multi-modal large language models, whic…
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs
Yikang Zhou, Tao Zhang, Shilin Xu +7
Recent advancements in multimodal large language models (MLLM) have shown a strong ability in visual perception, reasoning abilities, and vision-language understanding. However, th…
CoCo4D: Comprehensive and Complex 4D Scene Generation
Junwei Zhou, Xueting Li, Lu Qi +1
Existing 4D synthesis methods primarily focus on object-level generation or dynamic scene synthesis with limited novel views, restricting their ability to generate multi-view consi…