activity
20202025
most citedGraFormer: Graph Convolution Transformer for 3D Pose Estimation

16 citations · 30 across the 5 of their papers we have counts for

collaborators

14 papers

cs.CV2025

Adaptive Keyframe Sampling for Long Video Understanding

Xi Tang, Jihao Qiu, Lingxi Xie +3

Multimodal large language models (MLLMs) have enabled open-world visual understanding by injecting visual input as extra tokens into large language models (LLMs) as contexts. Howev…

cs.CV2025

YOLOv12: Attention-Centric Real-Time Object Detectors

Yunjie Tian, Qixiang Ye, David Doermann

Enhancing the network architecture of the YOLO framework has been crucial for a long time, but has focused on CNN-based improvements despite the proven superiority of attention mec…

cs.CV20241 cited

Personalized Large Vision-Language Models

Chau Pham, Hoang Phan, David Doermann +1

The personalization model has gained significant attention in image generation yet remains underexplored for large vision-language models (LVLMs). Beyond generic ones, with persona…

cs.CV2024

ClawMachine: Learning to Fetch Visual Tokens for Referential Comprehension

Tianren Ma, Lingxi Xie, Yunjie Tian +2

Aligning vision and language concepts at a finer level remains an essential topic of multimodal large language models (MLLMs), particularly for tasks such as referring and groundin…

cs.CV2024

Artemis: Towards Referential Understanding in Complex Videos

Jihao Qiu, Yuan Zhang, Xi Tang +6

Videos carry rich visual information including object description, action, interaction, etc., but the existing multimodal large language models (MLLMs) fell short in referential un…

cs.CV2024

Building Vision Models upon Heat Conduction

Zhaozhi Wang, Yue Liu, Yunjie Tian +3

Visual representation models leveraging attention mechanisms are challenged by significant computational overhead, particularly when pursuing large receptive fields. In this study,…