collaborators

7 papers

cs.CV2026

Look Clearly Before Answering: Mitigating Hallucinations in LVLMs via Saliency-Driven Perceptual Realignment

Pengxu Chen, Yao Zhu, Guangming Zhu +4

Large vision-language models (LVLMs) have demonstrated remarkable capabilities in multimodal understanding. However, they remain prone to hallucinations, generating responses that…

cs.CV2026

Knowledge-guided Disentanglement with Atomic Actions for Action Recognition

Tianci Wu, Siqi Cao, Guangming Zhu +6

Action recognition in complex scenes often involves multiple concurrent fine-grained actions, making it challenging to model internal action structures. Most existing methods rely…

cs.CV2026

Bridging Vision and Language Concepts through Optimal Transport Semantic Flow

Chenyang Zhang, Anqi Dong, Guangming Zhu +4

Concept Bottleneck Models (CBMs) promise transparent reasoning by predicting through human-interpretable concepts, yet their effectiveness fundamentally depends on how well visual…

cs.CV2026

Sketch and Text Synergy: Fusing Structural Contours and Descriptive Attributes for Fine-Grained Image Retrieval

Siyuan Wang, Hanchen Gao, Guangming Zhu +5

Fine-grained image retrieval via hand-drawn sketches or textual descriptions remains a critical challenge due to inherent modality gaps. While hand-drawn sketches capture complex s…

cs.CV2026

SkeletonContext: Skeleton-side Context Prompt Learning for Zero-Shot Skeleton-based Action Recognition

Ning Wang, Tieyue Wu, Naeha Sharif +5

Zero-shot skeleton-based action recognition aims to recognize unseen actions by transferring knowledge from seen categories through semantic descriptions. Most existing methods typ…

cs.CV2025

Multi-Granularity Mutual Refinement Network for Zero-Shot Learning

Ning Wang, Long Yu, Cong Hua +5

Zero-shot learning (ZSL) aims to recognize unseen classes with zero samples by transferring semantic knowledge from seen classes. Current approaches typically correlate global visu…