16 citations · 30 across the 5 of their papers we have counts for
14 papers
Adaptive Keyframe Sampling for Long Video Understanding
Xi Tang, Jihao Qiu, Lingxi Xie +3
Multimodal large language models (MLLMs) have enabled open-world visual understanding by injecting visual input as extra tokens into large language models (LLMs) as contexts. Howev…
YOLOv12: Attention-Centric Real-Time Object Detectors
Yunjie Tian, Qixiang Ye, David Doermann
Enhancing the network architecture of the YOLO framework has been crucial for a long time, but has focused on CNN-based improvements despite the proven superiority of attention mec…
Personalized Large Vision-Language Models
Chau Pham, Hoang Phan, David Doermann +1
The personalization model has gained significant attention in image generation yet remains underexplored for large vision-language models (LVLMs). Beyond generic ones, with persona…
ClawMachine: Learning to Fetch Visual Tokens for Referential Comprehension
Tianren Ma, Lingxi Xie, Yunjie Tian +2
Aligning vision and language concepts at a finer level remains an essential topic of multimodal large language models (MLLMs), particularly for tasks such as referring and groundin…
Artemis: Towards Referential Understanding in Complex Videos
Jihao Qiu, Yuan Zhang, Xi Tang +6
Videos carry rich visual information including object description, action, interaction, etc., but the existing multimodal large language models (MLLMs) fell short in referential un…
Building Vision Models upon Heat Conduction
Zhaozhi Wang, Yue Liu, Yunjie Tian +3
Visual representation models leveraging attention mechanisms are challenged by significant computational overhead, particularly when pursuing large receptive fields. In this study,…