71 citations · 142 across the 17 of their papers we have counts for
21 papers
Link-Context Learning for Multimodal LLMs
Yan Tai, Weichen Fan, Zhao Zhang +3
The ability to learn from context with novel concepts, and deliver appropriate responses are essential in human conversations. Despite current Multimodal Large Language Models (MLL…
Relation-Aware Distribution Representation Network for Person Clustering with Multiple Modalities
Kaijian Liu, Shixiang Tang, Ziyue Li +4
Person clustering with multi-modal clues, including faces, bodies, and voices, is critical for various tasks, such as movie parsing and identity-based movie editing. Related method…
Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Keqin Chen, Zhao Zhang, Weili Zeng +3
In human conversations, individuals can indicate relevant regions within a scene while addressing others. In turn, the other person can then respond by referring to specific region…
Described Object Detection: Liberating Object Detection with Flexible Expressions
Chi Xie, Zhao Zhang, Yixuan Wu +3
Detecting objects based on language information is a popular task that includes Open-Vocabulary object Detection (OVD) and Referring Expression Comprehension (REC). In this paper,…
Patch-Level Contrasting without Patch Correspondence for Accurate and Dense Contrastive Representation Learning
Shaofeng Zhang, Feng Zhu, Rui Zhao +1
We propose ADCLR: A ccurate and D ense Contrastive Representation Learning, a novel self-supervised learning framework for learning accurate and dense vision representation. To ext…
Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
Xiaoshi Wu, Yiming Hao, Keqiang Sun +4
Recent text-to-image generative models can generate high-fidelity images from text inputs, but the quality of these generated images cannot be accurately evaluated by existing eval…