collaborators

6 papers

cs.CV2025

ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding

Hao Lu, Jiahao Wang, Yaolun Zhang +5

Video multimodal large language models (Video-MLLMs) have achieved remarkable progress in video understanding. However, they remain vulnerable to hallucination-producing content in…

cs.CV2025

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation

Tongtian Yue, Longteng Guo, Yepeng Tang +4

Despite the impressive advancements of Large Vision-Language Models (LVLMs), existing approaches suffer from a fundamental bottleneck: inefficient visual-language integration. Curr…

cs.LG2025

Prefix Grouper: Efficient GRPO Training through Shared-Prefix Forward

Zikang Liu, Tongtian Yue, Yepeng Tang +5

Group Relative Policy Optimization (GRPO) enhances policy learning by computing gradients from relative comparisons among candidate outputs that share a common input prefix. Despit…

cs.CV2025

Towards Unified Referring Expression Segmentation Across Omni-Level Visual Target Granularities

Jing Liu, Wenxuan Wang, Yisi Zhang +5

Referring expression segmentation (RES) aims at segmenting the entities' masks that match the descriptive language expression. While traditional RES methods primarily address objec…

cs.CV2025

Image Difference Grounding with Natural Language

Wenxuan Wang, Zijia Zhao, Yisi Zhang +4

Visual grounding (VG) typically focuses on locating regions of interest within an image using natural language, and most existing VG methods are limited to single-image interpretat…

cs.AI2025

VRoPE: Rotary Position Embedding for Video Large Language Models

Zikang Liu, Longteng Guo, Yepeng Tang +6

Rotary Position Embedding (RoPE) has shown strong performance in text-based Large Language Models (LLMs), but extending it to video remains a challenge due to the intricate spatiot…