4 papers
MiVE: Multiscale Vision-language features for reference-guided video Editing
Tong Wang, Meng Zou, Chengjing Wu +4
Reference-guided video editing takes a source video, a text instruction, and a reference image as inputs, requiring the model to faithfully apply the instructed edits while preserv…
UHR-Micro: Diagnosing and Mitigating the Resolution Illusion in Earth Observation VLMs
Shuo Ni, Tong Wang, Jing Zhang +4
Vision-Language Models (VLMs) increasingly operate on ultra-high-resolution (UHR) Earth observation imagery, yet they remain vulnerable to a severe scale mismatch between large-sca…
See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection
Zhiheng Wu, Tong Wang, Shuning Wang +2
Recent advances in Vision-Language Models (VLMs) have benefited from Reinforcement Learning (RL) for enhanced reasoning. However, existing methods still face critical limitations,…
MIRAGE: A Micro-Interaction Relational Architecture for Grounded Exploration in Multi-Figure Artworks
Jui-Cheng Chiu, Yu-Chao Wang, Shengyang Luo +4
Appreciating multi-figure paintings requires understanding how characters relate through subtle cues like gaze alignment, gesture, and spatial arrangement. We present MIRAGE, an ev…