5 papers
VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models
Dong Xing, Jiaxin Chen, Hang Yang +3
Large vision-language models (LVLMs) have demonstrated strong performance in open-ended video understanding, yet they remain prone to fluent responses unsupported by video evidence…
Qwen3-VL-Seg: Unlocking Open-World Referring Segmentation with Vision-Language Grounding
Yuan Yao, Qiushi Yang, Humen Zhong +5
Open-world referring segmentation requires grounding unconstrained language expressions to precise pixel-level regions. Existing multimodal large language models (MLLMs) exhibit st…
McSc: Motion-Corrective Preference Alignment for Video Generation with Self-Critic Hierarchical Reasoning
Qiushi Yang, Yingjie Chen, Yuan Yao +3
Text-to-video (T2V) generation has achieved remarkable progress in producing high-quality videos aligned with textual prompts. However, aligning synthesized videos with nuanced hum…
MoSAM: Motion-Guided Segment Anything Model with Spatial-Temporal Memory Selection
Qiushi Yang, Yuan Yao, Miaomiao Cui +1
The recent Segment Anything Model 2 (SAM2) has demonstrated exceptional capabilities in interactive object segmentation for both images and videos. However, as a foundational model…
Towards Fine-grained Interactive Segmentation in Images and Videos
Yuan Yao, Qiushi Yang, Miaomiao Cui +1
The recent Segment Anything Models (SAMs) have emerged as foundational visual models for general interactive segmentation. Despite demonstrating robust generalization abilities, th…