7 papers
Look Clearly Before Answering: Mitigating Hallucinations in LVLMs via Saliency-Driven Perceptual Realignment
Pengxu Chen, Yao Zhu, Guangming Zhu +4
Large vision-language models (LVLMs) have demonstrated remarkable capabilities in multimodal understanding. However, they remain prone to hallucinations, generating responses that…
Knowledge-guided Disentanglement with Atomic Actions for Action Recognition
Tianci Wu, Siqi Cao, Guangming Zhu +6
Action recognition in complex scenes often involves multiple concurrent fine-grained actions, making it challenging to model internal action structures. Most existing methods rely…
Bridging Vision and Language Concepts through Optimal Transport Semantic Flow
Chenyang Zhang, Anqi Dong, Guangming Zhu +4
Concept Bottleneck Models (CBMs) promise transparent reasoning by predicting through human-interpretable concepts, yet their effectiveness fundamentally depends on how well visual…
Sketch and Text Synergy: Fusing Structural Contours and Descriptive Attributes for Fine-Grained Image Retrieval
Siyuan Wang, Hanchen Gao, Guangming Zhu +5
Fine-grained image retrieval via hand-drawn sketches or textual descriptions remains a critical challenge due to inherent modality gaps. While hand-drawn sketches capture complex s…
SkeletonContext: Skeleton-side Context Prompt Learning for Zero-Shot Skeleton-based Action Recognition
Ning Wang, Tieyue Wu, Naeha Sharif +5
Zero-shot skeleton-based action recognition aims to recognize unseen actions by transferring knowledge from seen categories through semantic descriptions. Most existing methods typ…
Multi-Granularity Mutual Refinement Network for Zero-Shot Learning
Ning Wang, Long Yu, Cong Hua +5
Zero-shot learning (ZSL) aims to recognize unseen classes with zero samples by transferring semantic knowledge from seen classes. Current approaches typically correlate global visu…