2 papers
cs.CV2026
V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning
Haoxiang Sun, Zhihang Yi, Langxuan Deng +6
Fine-grained visual reasoning requires multimodal large language models (MLLMs) to identify task-relevant visual evidence and ground their reasoning in local image regions. Existin…
cs.CL2026
Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models
Peiqi Jia, Haonan Jia, Ziqi Miao +3
With the widespread deployment of Multimodal Large Language Models (MLLMs) in social interaction, understanding and controlling their behavior under complex personality conditions…