3 papers
cs.CV2026
Look Clearly Before Answering: Mitigating Hallucinations in LVLMs via Saliency-Driven Perceptual Realignment
Pengxu Chen, Yao Zhu, Guangming Zhu +4
Large vision-language models (LVLMs) have demonstrated remarkable capabilities in multimodal understanding. However, they remain prone to hallucinations, generating responses that…
cs.CV2026
Knowledge-guided Disentanglement with Atomic Actions for Action Recognition
Tianci Wu, Siqi Cao, Guangming Zhu +6
Action recognition in complex scenes often involves multiple concurrent fine-grained actions, making it challenging to model internal action structures. Most existing methods rely…
cs.CL2026
Compatibility-Aware Dynamic Fine-Tuning for Large Language Models
Yucheng Zhou, Junwei Sheng, Qianning Wang +1
Supervised Fine-Tuning (SFT) is the predominant paradigm for aligning large language models (LLMs), yet it suffers from optimization instability and limited generalization. Recent…