Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation
Ruozhen He, Meng Wei, Ziyan Yang +1
Multi-shot video generation extends single-shot generation to coherent visual narratives, yet maintaining consistent characters, objects, and locations across shots remains a chall…
cs.CV2023
In Defense of Clip-based Video Relation Detection
Meng Wei, Long Chen, Wei Ji +2
Video Visual Relation Detection (VidVRD) aims to detect visual relationship triplets in videos using spatial bounding boxes and temporal boundaries. Existing VidVRD methods can be…
cs.CV2021
Vision Transformer with Progressive Sampling
Xiaoyu Yue, Shuyang Sun, Zhanghui Kuang +4
Transformers with powerful global relation modeling abilities have been introduced to fundamental computer vision tasks recently. As a typical example, the Vision Transformer (ViT)…