most citedSimVG: A Simple Framework for Visual Grounding with Decoupled Multi-modal Fusion

6 citations · 8 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CV2025★ 2 cited

Improving Generalized Visual Grounding with Instance-aware Joint Learning

Ming Dai, Wenxuan Cheng, Jiang-Jiang Liu +4

Generalized visual grounding tasks, including Generalized Referring Expression Comprehension (GREC) and Segmentation (GRES), extend the classical visual grounding paradigm by accom…

cs.CV2025

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination

Ming Dai, Wenxuan Cheng, Jiedong Zhuang +4

Recent advances in visual grounding have largely shifted away from traditional proposal-based two-stage frameworks due to their inefficiency and high computational complexity, favo…

cs.LG2025

DASViT: Differentiable Architecture Search for Vision Transformer

Pengjin Wu, Ferrante Neri, Zhenhua Feng

Designing effective neural networks is a cornerstone of deep learning, and Neural Architecture Search (NAS) has emerged as a powerful tool for automating this process. Among the ex…

cs.CV2025

Vision Remember: Recovering Visual Information in Efficient LVLM with Vision Feature Resampling

Ze Feng, Jiang-jiang Liu, Sen Yang +5

The computational expense of redundant vision tokens in Large Vision-Language Models (LVLMs) has led many existing methods to compress them via a vision projector. However, this co…

cs.CV2024★ 6 cited

SimVG: A Simple Framework for Visual Grounding with Decoupled Multi-modal Fusion

Ming Dai, Lingfeng Yang, Yihao Xu +2

Visual grounding is a common vision task that involves grounding descriptive sentences to the corresponding regions of an image. Most existing methods use independent image-text en…