Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
PAVXploreRL: Physical-Action-Visual World Model Reinforcement Learning with Action Exploration
Han Wang, Zijun Wang, Shuoshuo Xue +5
Action-conditioned world models are a key component of embodied AI, serving as scalable policy evaluators that reduce reliance on expensive real-world rollouts. To accurately captu…
cs.CV2026
Self-Captioning Multimodal Interaction Tuning: Amplifying Exploitable Redundancies for Robust Vision Language Models
Yuriel Ryan, Hei Man Ip, Adriel Kuek +2
Current vision language models face hallucination and robustness issues against ambiguous or corrupted modalities. We hypothesize that these issues can be addressed by exploiting t…
cs.CV2024
All in an Aggregated Image for In-Image Learning
Lei Wang, Wanyu Xu, Zhiqiang Hu +5
This paper introduces a new in-context learning (ICL) mechanism called In-Image Learning (IL) that combines demonstration examples, visual cues, and chain-of-thought reasoning…