collaborators

7 papers

cs.CV2026

Knowledge-guided Disentanglement with Atomic Actions for Action Recognition

Tianci Wu, Siqi Cao, Guangming Zhu +6

The paper introduces a framework that uses large language models to break down action labels into atomic actions and injects this semantic knowledge into video features to improve…

cs.CV2026

Look Clearly Before Answering: Mitigating Hallucinations in LVLMs via Saliency-Driven Perceptual Realignment

Pengxu Chen, Yao Zhu, Guangming Zhu +4

Large vision-language models (LVLMs) have demonstrated remarkable capabilities in multimodal understanding. However, they remain prone to hallucinations, generating responses that…

cs.CV2026

Bridging Vision and Language Concepts through Optimal Transport Semantic Flow

Chenyang Zhang, Anqi Dong, Guangming Zhu +4

Concept Bottleneck Models (CBMs) promise transparent reasoning by predicting through human-interpretable concepts, yet their effectiveness fundamentally depends on how well visual…

cs.CV2026

Sketch and Text Synergy: Fusing Structural Contours and Descriptive Attributes for Fine-Grained Image Retrieval

Siyuan Wang, Hanchen Gao, Guangming Zhu +5

Fine-grained image retrieval via hand-drawn sketches or textual descriptions remains a critical challenge due to inherent modality gaps. While hand-drawn sketches capture complex s…

cs.CV2026

SkeletonContext: Skeleton-side Context Prompt Learning for Zero-Shot Skeleton-based Action Recognition

Ning Wang, Tieyue Wu, Naeha Sharif +5

Zero-shot skeleton-based action recognition aims to recognize unseen actions by transferring knowledge from seen categories through semantic descriptions. Most existing methods typ…

cs.CV2025

Prompt-guided Disentangled Representation for Action Recognition

Tianci Wu, Guangming Zhu, Jiang Lu +4

Action recognition is a fundamental task in video understanding. Existing methods typically extract unified features to process all actions in one video, which makes it challenging…