2 papers
cs.CV2026
VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models
Dong Xing, Jiaxin Chen, Hang Yang +3
Large vision-language models (LVLMs) have demonstrated strong performance in open-ended video understanding, yet they remain prone to fluent responses unsupported by video evidence…
cs.CV2025
Segment Any RGB-Thermal Model with Language-aided Distillation
Dong Xing, Xianxun Zhu, Wei Zhou +3
The recent Segment Anything Model (SAM) demonstrates strong instance segmentation performance across various downstream tasks. However, SAM is trained solely on RGB data, limiting…