3 papers
cs.CV2026
VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models
Dong Xing, Jiaxin Chen, Hang Yang +3
Large vision-language models (LVLMs) have demonstrated strong performance in open-ended video understanding, yet they remain prone to fluent responses unsupported by video evidence…
cs.CV2025
Segment Any RGB-Thermal Model with Language-aided Distillation
Dong Xing, Xianxun Zhu, Wei Zhou +3
The recent Segment Anything Model (SAM) demonstrates strong instance segmentation performance across various downstream tasks. However, SAM is trained solely on RGB data, limiting…
cs.CV2025
PTQ4RIS: Post-Training Quantization for Referring Image Segmentation
Xiaoyan Jiang, Hang Yang, Kaiying Zhu +3
Referring Image Segmentation (RIS), aims to segment the object referred by a given sentence in an image by understanding both visual and linguistic information. However, existing R…