activity
20242026
most citedDiffusion Feedback Helps CLIP See Better

1 citations · 1 across the 9 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

TimeThink: Reasoning with Time for Video LLMs

Handong Li, Longteng Guo, Zikang Liu +8

Video reasoning requires models to identify and verify temporally localized evidence within long video sequences. Recent Video Large Language Models (Video-LLMs) have shown promisi…

cs.CV2026

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding

Handong Li, Zikang Liu, Longteng Guo +10

Processing long-form videos with Video Large Language Models (Video-LLMs) is computationally prohibitive. Current efficiency methods often compromise fine-grained perception throug…

cs.CV2025

ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding

Hao Lu, Jiahao Wang, Yaolun Zhang +5

Video multimodal large language models (Video-MLLMs) have achieved remarkable progress in video understanding. However, they remain vulnerable to hallucination-producing content in…

cs.CV2025

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation

Tongtian Yue, Longteng Guo, Yepeng Tang +4

Despite the impressive advancements of Large Vision-Language Models (LVLMs), existing approaches suffer from a fundamental bottleneck: inefficient visual-language integration. Curr…

cs.CV2025

Towards Unified Referring Expression Segmentation Across Omni-Level Visual Target Granularities

Jing Liu, Wenxuan Wang, Yisi Zhang +5

Referring expression segmentation (RES) aims at segmenting the entities' masks that match the descriptive language expression. While traditional RES methods primarily address objec…

cs.CV2025

Image Difference Grounding with Natural Language

Wenxuan Wang, Zijia Zhao, Yisi Zhang +4

Visual grounding (VG) typically focuses on locating regions of interest within an image using natural language, and most existing VG methods are limited to single-image interpretat…