3 papers
cs.RO2026
HALO: A Unified Vision-Language-Action Model for Embodied Multimodal Chain-of-Thought Reasoning
Quanxin Shou, Fangqi Zhu, Shawn Chen +9
Vision-Language-Action (VLA) models have shown strong performance in robotic manipulation, but often struggle in long-horizon or out-of-distribution scenarios due to the lack of ex…
cs.CV2026
Generating Storytelling Images with Rich Chains-of-Reasoning
Xiujie Song, Qi Jia, Shota Watanabe +4
A single image can convey a compelling story through logically connected visual clues, forming Chains-of-Reasoning (CoRs). We define these semantically rich images as Storytelling…
cs.CV2025
Is Your Image a Good Storyteller?
Xiujie Song, Xiaoyi Pang, Haifeng Tang +2
Quantifying image complexity at the entity level is straightforward, but the assessment of semantic complexity has been largely overlooked. In fact, there are differences in semant…