2 papers
cs.CV2026
GeoWeaver: Grounding Visual Tokens with Geometric Evidence before Scene Reasoning
Deshui Miao, Xingsen Huang, Yameng Gu +3
Spatio-temporal reasoning in vision-language models requires visual representations that preserve physical geometry rather than merely semantic appearance. Recent multimodal models…
cs.CV2026
OAMVOS:2nd Report for 5th PVUW MOSE Track
Deshui Miao, Xingsen Huang, Yameng Gu +3
SAM-based dense trackers provide strong short-term mask propagation but remain fragile under long occlusion, fast motion, viewpoint change, and distractors. The problem is especial…