12 papers
PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence Seeking
Dengxian Gong, Yuanzheng Wu, Haobo Yuan +11
This paper explores multi-turn visual reasoning and observes that MLLMs repeatedly fail to localize the target, leading to long, redundant trajectories. We attribute this failure t…
BusterX: MLLM-Powered AI-Generated Video Forgery Detection and Explanation
Haiquan Wen, Yiwei He, Zhenglin Huang +7
As generative video models become increasingly realistic, detecting AI-generated videos requires systems that offer both accuracy and interpretability. However, applying Multimodal…
SaSaSaSa2VA: 2nd Place of the 5th PVUW MeViS-Text Track
Dengxian Gong, Quanzhu Niu, Shihao Chen +6
Referring video object segmentation (RVOS) commonly grounds targets in videos based on static textual cues. MeViS benchmark extends this by incorporating motion-centric expressions…
InstaVSR: Taming Diffusion for Efficient and Temporally Consistent Video Super-Resolution
Jintong Hu, Bin Chen, Zhenyu Hu +3
Video super-resolution (VSR) seeks to reconstruct high-resolution frames from low-resolution inputs. While diffusion-based methods have substantially improved perceptual quality, e…
Conditional Panoramic Image Generation via Masked Autoregressive Modeling
Chaoyang Wang, Xiangtai Li, Lu Qi +4
Recent progress in panoramic image generation has underscored two critical limitations in existing approaches. First, most methods are built upon diffusion models, which are inhere…
The 1st Solution for 7th LSVOS RVOS Track: SaSaSa2VA
Quanzhu Niu, Dengxian Gong, Shihao Chen +6
Referring video object segmentation (RVOS) requires segmenting and tracking objects in videos conditioned on natural-language expressions, demanding fine-grained understanding of b…