Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
V-CORE: Temporally Consistent Video Understanding for Video-LLM
Zhengjian Kang, Qi Chen, Rui Liu +4
Recent Video Large Language Models (Video-LLMs) have shown strong multimodal reasoning capabilities, yet remain challenged by video understanding tasks that require consistent temp…
cs.CV2025
Dual-R-DETR: Resolving Query Competition with Pairwise Routing in Transformer Decoders
Ye Zhang, Qi Chen, Wenyou Huang +2
Detection Transformers (DETR) formulate object detection as a set prediction problem and enable end-to-end training without post-processing. However, object queries in DETR interac…
cs.CV2025
LP-DETR: Layer-wise Progressive Relations for Object Detection
Zhengjian Kang, Ye Zhang, Xiaoyu Deng +2
This paper presents LP-DETR (Layer-wise Progressive DETR), a novel approach that enhances DETR-based object detection through multi-scale relation modeling. Our method introduces l…