collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing

Chenxuan Miao, Yutong Feng, Yi Lu +6

Instruction-based video editing (IVE) is an emerging field with broad applications, yet evaluating editing models remains challenging. Existing benchmarks suffer from two major lim…

cs.CV2026

PanoWorld: Towards Spatial Supersensing in 360 Panorama World

Changpeng Wang, Xin Lin, Junhan Liu +5

Multimodal large laboratory models (MLLMs) still struggle with spatial understanding under the dominant perspective-image paradigm, which inherits the narrow field of view of human…

cs.CV2026

DenseScout: Algorithm-System Co-design for Budgeted Tiny Object Selection on Edge Platforms

Zhouzhi Xiong, Zimo Zeng, Yi Chen +3

Deploying high-resolution tiny-object perception on edge platforms requires not only accurate localization, but also selecting a small set of informative patches under compute, tra…

cs.CV2025

From Illusion to Intention: Visual Rationale Learning for Vision-Language Reasoning

Changpeng Wang, Haozhe Wang, Xi Chen +6

Recent advances in vision-language reasoning underscore the importance of thinking with images, where models actively ground their reasoning in visual evidence. Yet, prevailing fra…

cs.CV2025

ROSE: Remove Objects with Side Effects in Videos

Chenxuan Miao, Yutong Feng, Jianshu Zeng +7

Video object removal has achieved advanced performance due to the recent success of video generative models. However, when addressing the side effects of objects, e.g., their shado…

cs.CV2025

TextVidBench: A Benchmark for Long Video Scene Text Understanding

Yangyang Zhong, Ji Qi, Yuan Yao +5

Despite recent progress on the short-video Text-Visual Question Answering (ViteVQA) task - largely driven by benchmarks such as M4-ViteVQA - existing datasets still suffer from lim…