collaborators

5 papers

cs.CV2026

Dynamic Resolution Routing for Efficient Egocentric Grounding

Huixin Sun, Wangbo Zhao, Fanyue Wei +3

Egocentric visual grounding requires high-resolution inputs to localize small objects. However, scaling Multimodal Large Language Models to this domain is constrained by the excess…

cs.CV2026

Decouple and Cache: KV Cache Construction for Streaming Video Understanding

Zhanzhong Pang, Dibyadip Chatterjee, Fadime Sener +1

Streaming video understanding requires processing unbounded video streams with limited memory and computation, posing two key challenges. First, continuously constructing new and e…

cs.CV2026

Don't Pause! Every prediction matters in a streaming video

Dibyadip Chatterjee, Zhanzhong Pang, Fadime Sener +2

Streaming video models should respond the moment an event unfolds, not after the moment has passed. Yet existing online VideoQA benchmarks remain largely retrospective. They pause…

cs.CV2025

The Devil is in Attention Sharing: Improving Complex Non-rigid Image Editing Faithfulness via Attention Synergy

Zhuo Chen, Fanyue Wei, Runze Xu +4

Training-free image editing with large diffusion models has become practical, yet faithfully performing complex non-rigid edits (e.g., pose or shape changes) remains highly challen…

cs.CV2025

The Devil is in the Spurious Correlations: Boosting Moment Retrieval with Dynamic Learning

Xinyang Zhou, Fanyue Wei, Lixin Duan +2

Given a textual query along with a corresponding video, the objective of moment retrieval aims to localize the moments relevant to the query within the video. While commendable res…