activity
20202025
most citedVMamba: Visual State Space Model

382 citations · 431 across the 15 of their papers we have counts for

collaborators
Showing cs.CVShow all

16 papers · 1 filter

cs.CV2025

Adaptive Keyframe Sampling for Long Video Understanding

Xi Tang, Jihao Qiu, Lingxi Xie +3

Multimodal large language models (MLLMs) have enabled open-world visual understanding by injecting visual input as extra tokens into large language models (LLMs) as contexts. Howev…

cs.CV2025

YOLOv12: Attention-Centric Real-Time Object Detectors

Yunjie Tian, Qixiang Ye, David Doermann

Enhancing the network architecture of the YOLO framework has been crucial for a long time, but has focused on CNN-based improvements despite the proven superiority of attention mec…

cs.CV2024★ 1 cited

Personalized Large Vision-Language Models

Chau Pham, Hoang Phan, David Doermann +1

The personalization model has gained significant attention in image generation yet remains underexplored for large vision-language models (LVLMs). Beyond generic ones, with persona…

cs.CV2024

ClawMachine: Learning to Fetch Visual Tokens for Referential Comprehension

Tianren Ma, Lingxi Xie, Yunjie Tian +2

Aligning vision and language concepts at a finer level remains an essential topic of multimodal large language models (MLLMs), particularly for tasks such as referring and groundin…

cs.CV2024

Artemis: Towards Referential Understanding in Complex Videos

Jihao Qiu, Yuan Zhang, Xi Tang +6

Videos carry rich visual information including object description, action, interaction, etc., but the existing multimodal large language models (MLLMs) fell short in referential un…

cs.CV2024★ 1 cited

Building Vision Models upon Heat Conduction

Zhaozhi Wang, Yue Liu, Yunjie Tian +3

Visual representation models leveraging attention mechanisms are challenged by significant computational overhead, particularly when pursuing large receptive fields. In this study,…