activity
20232026
most citedDepthSSC: Monocular 3D Semantic Scene Completion via Depth-Spatial Alignment and Voxel Adaptation

11 citations · 13 across the 33 of their papers we have counts for

collaborators
Showing cs.CVShow all

13 papers · 1 filter

cs.CV2026

Token-Budget Distillation: Transferring Full-Token Semantics to Compressed Video Vision-Language Models

Xiaoyang Guo, Guoping Luo, Jusheng Zhang +2

Adapting video vision-language models (VLMs) is computationally expensive because video inputs produce a large number of visual tokens, making both fine-tuning and inference costly…

cs.CV2026

The Fourth Challenge on Image Super-Resolution (4) at NTIRE 2026: Benchmark Results and Method Overview

Zheng Chen, Kai Liu, Jingkai Wang +150

This paper presents the NTIRE 2026 image super-resolution (4) challenge, one of the associated competitions of the NTIRE 2026 Workshop at CVPR 2026. The challenge aims to r…

cs.CV2026

Process-of-Thought Reasoning for Videos

Jusheng Zhang, Kaitong Cai, Jian Wang +3

Video understanding requires not only recognizing visual content but also performing temporally grounded, multi-step reasoning over long and noisy observations. We propose Process-…

cs.CV2026

ResAgent: Entropy-based Prior Point Discovery and Visual Reasoning for Referring Expression Segmentation

Yihao Wang, Jusheng Zhang, Ziyi Tang +2

Referring Expression Segmentation (RES) is a core vision-language segmentation task that enables pixel-level understanding of targets via free-form linguistic expressions, supporti…

cs.CV2026

3D-Agent:Tri-Modal Multi-Agent Collaboration for Scalable 3D Object Annotation

Jusheng Zhang, Yijia Fan, Zimo Wen +2

Driven by applications in autonomous driving robotics and augmented reality 3D object annotation presents challenges beyond 2D annotation including spatial complexity occlusion and…

cs.CV2025

FlashVLM: Text-Guided Visual Token Selection for Large Multimodal Models

Kaitong Cai, Jusheng Zhang, Jing Yang +4

Large vision-language models (VLMs) typically process hundreds or thousands of visual tokens per image or video frame, incurring quadratic attention cost and substantial redundancy…