1 citations · 1 across the 1 of their papers we have counts for
3 papers
cs.CV2023★ 1 cited
Link-Context Learning for Multimodal LLMs
Yan Tai, Weichen Fan, Zhao Zhang +3
The ability to learn from context with novel concepts, and deliver appropriate responses are essential in human conversations. Despite current Multimodal Large Language Models (MLL…
cs.CV2023
Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Keqin Chen, Zhao Zhang, Weili Zeng +3
In human conversations, individuals can indicate relevant regions within a scene while addressing others. In turn, the other person can then respond by referring to specific region…
cs.CV2023
Described Object Detection: Liberating Object Detection with Flexible Expressions
Chi Xie, Zhao Zhang, Yixuan Wu +3
Detecting objects based on language information is a popular task that includes Open-Vocabulary object Detection (OVD) and Referring Expression Comprehension (REC). In this paper,…