7 papers
OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing
Chenxuan Miao, Yutong Feng, Yi Lu +6
Instruction-based video editing (IVE) is an emerging field with broad applications, yet evaluating editing models remains challenging. Existing benchmarks suffer from two major lim…
PanoWorld: Towards Spatial Supersensing in 360 Panorama World
Changpeng Wang, Xin Lin, Junhan Liu +5
Multimodal large laboratory models (MLLMs) still struggle with spatial understanding under the dominant perspective-image paradigm, which inherits the narrow field of view of human…
DenseScout: Algorithm-System Co-design for Budgeted Tiny Object Selection on Edge Platforms
Xiong Zhouzhi, Zimo Zeng, Yi Chen +3
Deploying tiny object perception on edge platforms is challenging because practical systems must satisfy both strict compute budgets and end-to-end latency constraints. A common st…
Parallelism and Generation Order in Masked Diffusion Language Models: Limits Today, Potential Tomorrow
Yangyang Zhong, Yanmei Gu, Zhengqing Zang +14
Masked Diffusion Language Models (MDLMs) promise parallel token generation and arbitrary-order decoding, yet it remains unclear to what extent current models truly realize these ca…
From Illusion to Intention: Visual Rationale Learning for Vision-Language Reasoning
Changpeng Wang, Haozhe Wang, Xi Chen +6
Recent advances in vision-language reasoning underscore the importance of thinking with images, where models actively ground their reasoning in visual evidence. Yet, prevailing fra…
ROSE: Remove Objects with Side Effects in Videos
Chenxuan Miao, Yutong Feng, Jianshu Zeng +7
Video object removal has achieved advanced performance due to the recent success of video generative models. However, when addressing the side effects of objects, e.g., their shado…