collaborators

12 papers

cs.CV2026

Clearer Sight, Fewer Lies: Oriented Pickup Preference Optimization for Multimodal Hallucination Mitigation

Xin Zou, Haolin Deng, Yibo Yan +5

Multimodal Large Language Models (MLLMs) are prone to hallucination as their generation preferences are insufficiently calibrated to visual evidence, causing them to fall back on l…

cs.CV2026

Consistency as Inductive Bias: Learning Cross-View Invariance for Robust Multimodal Reasoning

Xin Zou, Haolin Deng, Yibo Yan +6

Inductive biases steer learning toward generalizable solutions by encoding task structure. In this work, we identify a crucial missing bias in MLLMs: cross-view consistency, \texti…

cs.CV2026

Learning from Fine-Grained Visual Discrepancies: Mitigating Multimodal Hallucinations via In-Context Visual Contrastive Optimization

Haolin Deng, Xin Zou, Zhiwei Jin +3

Multimodal hallucination remains a persistent challenge for Vision-Language Models (VLMs). Standard textual Direct Preference Optimization (DPO) often fails to mitigate it due to a…

cs.CV2026

MergeTok: Unified Continuous and Discrete Visual Tokenization via Token Merging

Luyuan Zhang, Siyuan Li, Zedong Wang +7

Most visual tokenizers for image generation are bifurcated into two families with complementary limitations: continuous VAEs offer high-fidelity reconstruction but suffer from dens…

cs.CV2026

Click-to-Ask: An AI Live Streaming Assistant with Offline Copywriting and Online Interactive QA

Ruizhi Yu, Keyang Zhong, Peng Liu +5

Live streaming commerce has become a prominent form of broadcasting in the modern era. To facilitate more efficient and convenient product promotions for streamers, we present Clic…

cs.CV2026

Pluggable Pruning with Contiguous Layer Distillation for Diffusion Transformers

Jian Ma, Qirong Peng, Xujie Zhu +3

Diffusion Transformers (DiTs) have shown exceptional performance in image generation, yet their large parameter counts incur high computational costs, impeding deployment in resour…