11 citations · 13 across the 33 of their papers we have counts for
13 papers · 1 filter
Token-Budget Distillation: Transferring Full-Token Semantics to Compressed Video Vision-Language Models
Xiaoyang Guo, Guoping Luo, Jusheng Zhang +2
Adapting video vision-language models (VLMs) is computationally expensive because video inputs produce a large number of visual tokens, making both fine-tuning and inference costly…
The Fourth Challenge on Image Super-Resolution (4) at NTIRE 2026: Benchmark Results and Method Overview
Zheng Chen, Kai Liu, Jingkai Wang +150
This paper presents the NTIRE 2026 image super-resolution (4) challenge, one of the associated competitions of the NTIRE 2026 Workshop at CVPR 2026. The challenge aims to r…
Process-of-Thought Reasoning for Videos
Jusheng Zhang, Kaitong Cai, Jian Wang +3
Video understanding requires not only recognizing visual content but also performing temporally grounded, multi-step reasoning over long and noisy observations. We propose Process-…
ResAgent: Entropy-based Prior Point Discovery and Visual Reasoning for Referring Expression Segmentation
Yihao Wang, Jusheng Zhang, Ziyi Tang +2
Referring Expression Segmentation (RES) is a core vision-language segmentation task that enables pixel-level understanding of targets via free-form linguistic expressions, supporti…
3D-Agent:Tri-Modal Multi-Agent Collaboration for Scalable 3D Object Annotation
Jusheng Zhang, Yijia Fan, Zimo Wen +2
Driven by applications in autonomous driving robotics and augmented reality 3D object annotation presents challenges beyond 2D annotation including spatial complexity occlusion and…
FlashVLM: Text-Guided Visual Token Selection for Large Multimodal Models
Kaitong Cai, Jusheng Zhang, Jing Yang +4
Large vision-language models (VLMs) typically process hundreds or thousands of visual tokens per image or video frame, incurring quadratic attention cost and substantial redundancy…