activity
20202023
most citedShikra: Unleashing Multimodal LLM's Referential Dialogue Magic

71 citations · 142 across the 17 of their papers we have counts for

collaborators

21 papers

cs.CV2023★ 1 cited

Link-Context Learning for Multimodal LLMs

Yan Tai, Weichen Fan, Zhao Zhang +3

The ability to learn from context with novel concepts, and deliver appropriate responses are essential in human conversations. Despite current Multimodal Large Language Models (MLL…

cs.CV2023

Relation-Aware Distribution Representation Network for Person Clustering with Multiple Modalities

Kaijian Liu, Shixiang Tang, Ziyue Li +4

Person clustering with multi-modal clues, including faces, bodies, and voices, is critical for various tasks, such as movie parsing and identity-based movie editing. Related method…

cs.CV2023★ 71 cited

Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Keqin Chen, Zhao Zhang, Weili Zeng +3

In human conversations, individuals can indicate relevant regions within a scene while addressing others. In turn, the other person can then respond by referring to specific region…

cs.CV2023★ 9 cited

Described Object Detection: Liberating Object Detection with Flexible Expressions

Chi Xie, Zhao Zhang, Yixuan Wu +3

Detecting objects based on language information is a popular task that includes Open-Vocabulary object Detection (OVD) and Referring Expression Comprehension (REC). In this paper,…

cs.CV2023★ 8 cited

Patch-Level Contrasting without Patch Correspondence for Accurate and Dense Contrastive Representation Learning

Shaofeng Zhang, Feng Zhu, Rui Zhao +1

We propose ADCLR: A ccurate and D ense Contrastive Representation Learning, a novel self-supervised learning framework for learning accurate and dense vision representation. To ext…

cs.CV2023★ 22 cited

Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Xiaoshi Wu, Yiming Hao, Keqiang Sun +4

Recent text-to-image generative models can generate high-fidelity images from text inputs, but the quality of these generated images cannot be accurately evaluated by existing eval…