2 papers
cs.CV2026
MeSD: Multi-Evidence Self-Distillation for VideoLLM
Weijie Zhu, Han Fang, Hanyu Fu +13
While reinforcement learning with verifiable rewards provides reliable outcome supervision for VideoLLMs, sequence-level rewards offer limited token-level guidance. On-policy self-…
cs.CV2026
GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding
Hao Li, Han Fang, Zixin Pan +8
Although multimodal large language models (MLLMs) have achieved remarkable progress, understanding 3D spatial relationships from 2D images remains a critical challenge. Existing me…