Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Task-Focused Memorization for Multimodal Agents
Tao Zou, Yichen He, Tian Qiu +2
Long-term memory is essential for multimodal agents to build coherent experience, accumulate world knowledge, and achieve continual learning. However, constructing effective memory…
cs.CV2025
VisPlay: Self-Evolving Vision-Language Models from Images
Yicheng He, Chengsong Huang, Zongxia Li +2
Reinforcement learning (RL) provides a principled framework for improving Vision-Language Models (VLMs) on complex reasoning tasks. However, existing RL approaches often rely on hu…
cs.CV2023
Dataset Bias Mitigation in Multiple-Choice Visual Question Answering and Beyond
Zhecan Wang, Long Chen, Haoxuan You +6
Vision-language (VL) understanding tasks evaluate models' comprehension of complex visual scenes through multiple-choice questions. However, we have identified two dataset biases t…