3 papers
cs.CV2025
Zoomer: Adaptive Image Focus Optimization for Black-box MLLM
Jiaxu Qian, Chendong Wang, Yifan Yang +16
Multimodal large language models (MLLMs) such as GPT-4o, Gemini Pro, and Claude 3.5 have enabled unified reasoning over text and visual inputs, yet they often hallucinate in real w…
cs.CV2025
VoLUT: Efficient Volumetric streaming enhanced by LUT-based super-resolution
Chendong Wang, Anlan Zhang, Yifan Yang +5
3D volumetric video provides immersive experience and is gaining traction in digital media. Despite its rising popularity, the streaming of volumetric video content poses significa…
cs.CV2025
Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning
Chendong Wang, Donglin Bai, Yifan Yang +11
We present \emph{Video-in-the-Loop} (ViTL), a two-stage long-video QA framework that preserves a fixed token budget by first \emph{localizing} question-relevant interval(s) with a…