Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
What Semantics Survive the Connector? Diagnosing VLM-to-DiT Alignment in Video Editing
Hangyu Lin, Chao Wen, Chengming Xu +4
Flow matching based video generative models have been increasingly relying on prepended Vision-Language Models (VLMs) to handle complex, instruction-based video editing. The prevai…
cs.CV2026
StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering
Ming Xie, Zizheng Huang, Xudong Tan +6
While streaming omni-video understanding demands continuous perception and proactive, real-time interaction, this crucial area remains largely under-explored. Current omni-modal me…
cs.CV2025
VTPerception-R1: Enhancing Multimodal Reasoning via Explicit Visual and Textual Perceptual Grounding
Yizhuo Ding, Mingkang Chen, Zhibang Feng +4
Multimodal large language models (MLLMs) often struggle to ground reasoning in perceptual evidence. We present a systematic study of perception strategies-explicit, implicit, visua…