collaborators

8 papers

cs.CV2026

SCOPE-Router: Cost-Aware Open-Set VLM Routing for Execution-Oriented Tasks

Tao Yu, Yifei Qu, Zhiqing Cui +14

Model routing aims to select the most suitable model from a candidate pool for each query, balancing quality and cost. Existing VLM routing research is limited to traditional VQA e…

cs.CV2026

HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization

Tao Cheng, Shi-Zhe Chen, Hao Zhang +3

Chain-of-Thought (CoT) reasoning significantly elevates the complex problem-solving capabilities of multimodal large language models (MLLMs). However, adapting CoT to vision typica…

cs.CV2026

VDE Bench: Evaluating The Capability of Image Editing Models to Modify Visual Documents

Hongzhu Yi, Yujia Yang, Yuanxiang Wang +18

In recent years, image editing models have made significant progress, enabling users to manipulate visual content in a flexible and interactive manner through natural language inst…

cs.CV2026

HY-Himmel Technical Report: Hierarchical Interleaved Multi-stream Motion Encoding for Long Video Understanding

Haopeng Jin, Hongzhu Yi, Wenlong Zhao +6

Long-video understanding with multimodal language models suffers from three compounding bottlenecks: heavy decode cost to obtain dense RGB frames, quadratic token growth with frame…

cs.CV2026

Omni IIE Bench: Benchmarking the Practical Capabilities of Image Editing Models

Yujia Yang, Yuanxiang Wang, Zhenyu Guan +11

While Instruction-based Image Editing (IIE) has achieved significant progress, existing benchmarks pursue task breadth via mixed evaluations. This paradigm obscures a critical fail…

cs.CV2026

Beyond Closed-Pool Video Retrieval: A Benchmark and Agent Framework for Real-World Video Search and Moment Localization

Tao Yu, Yujia Yang, Haopeng Jin +17

Traditional video retrieval benchmarks focus on matching precise descriptions to closed video pools, failing to reflect real-world searches characterized by fuzzy, multi-dimensiona…