Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Med-OPD: Improving Medical Vision-Language Models via Evidence-Aware On-Policy Distillation
Yunhang Qian, Jiaquan Yu, Jiawei Liu +3
Medical Vision-Language Models (Med-VLMs) require reliable reasoning from fine-grained visual evidence, yet existing models can produce plausible clinical answers by relying on lan…
cs.CV2026
MagicWorld: Towards Long-Horizon Stability for Interactive Video World Exploration
Guangyuan Li, Bo Li, Jinwei Chen +3
Recent interactive video world model methods generate scene evolution conditioned on user instructions. Although they achieve impressive results, two key limitations remain. First,…
cs.CV2025
CameraMaster: Unified Camera Semantic-Parameter Control for Photography Retouching
Qirui Yang, Yang Yang, Ying Zeng +5
Text-guided diffusion models have greatly advanced image editing and generation. However, achieving physically consistent image retouching with precise parameter control (e.g., exp…