3 papers
cs.CV2026
StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering
Ming Xie, Zizheng Huang, Xudong Tan +6
While streaming omni-video understanding demands continuous perception and proactive, real-time interaction, this crucial area remains largely under-explored. Current omni-modal me…
cs.CV2025
Enhancing Video Inpainting with Aligned Frame Interval Guidance
Ming Xie, Junqiu Yu, Qiaole Dong +2
Recent image-to-video (I2V) based video inpainting methods have made significant strides by leveraging single-image priors and modeling temporal consistency across masked frames. N…
cs.CV2025
AnyRefill: A Unified, Data-Efficient Framework for Left-Prompt-Guided Vision Tasks
Ming Xie, Chenjie Cao, Yunuo Cai +3
In this paper, we present a novel Left-Prompt-Guided (LPG) paradigm to address a diverse range of reference-based vision tasks. Inspired by the human creative process, we reformula…