4 papers · 1 filter
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression
Yuke Zhu, Chi Xie, Shuang Liang +2
Recent advances on Multi-modal Large Language Models have demonstrated that high-resolution image input is crucial for model capabilities, especially for fine-grained tasks. Howeve…
Compositional Learning in Transformer-Based Human-Object Interaction Detection
Zikun Zhuang, Ruihao Qian, Chi Xie +1
Human-object interaction (HOI) detection is an important part of understanding human activities and visual scenes. The long-tailed distribution of labeled instances is a primary ch…
Described Object Detection: Liberating Object Detection with Flexible Expressions
Chi Xie, Zhao Zhang, Yixuan Wu +3
Detecting objects based on language information is a popular task that includes Open-Vocabulary object Detection (OVD) and Referring Expression Comprehension (REC). In this paper,…
Category Query Learning for Human-Object Interaction Classification
Chi Xie, Fangao Zeng, Yue Hu +2
Unlike most previous HOI methods that focus on learning better human-object features, we propose a novel and complementary approach called category query learning. Such queries are…