collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2025

Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe Prior

Yulin Li, Haokun Gui, Ziyang Fan +4

Recent advances in Video Large Language Models (VLLMs) have achieved remarkable video understanding capabilities, yet face critical efficiency bottlenecks due to quadratic computat…

cs.CV2025

CalibCLIP: Contextual Calibration of Dominant Semantics for Text-Driven Image Retrieval

Bin Kang, Bin Chen, Junjie Wang +3

Existing Visual Language Models (VLMs) suffer structural limitations where a few low contribution tokens may excessively capture global semantics, dominating the information aggreg…

cs.CV2025

Generalized Decoupled Learning for Enhancing Open-Vocabulary Dense Perception

Junjie Wang, Keyu Chen, Yulin Li +4

Dense visual perception tasks have been constrained by their reliance on predefined categories, limiting their applicability in real-world scenarios where visual concepts are unbou…

cs.CV2025

DeCLIP: Decoupled Learning for Open-Vocabulary Dense Perception

Junjie Wang, Bin Chen, Yulin Li +3

Dense visual prediction tasks have been constrained by their reliance on predefined categories, limiting their applicability in real-world scenarios where visual concepts are unbou…

cs.CV2024

Multi-path Exploration and Feedback Adjustment for Text-to-Image Person Retrieval

Bin Kang, Bin Chen, Junjie Wang +1

Text-based person retrieval aims to identify the specific persons using textual descriptions as queries. Existing ad vanced methods typically depend on vision-language pre trained…