From the 1 of 5 linked papers with an AI index.
5 papers
WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models
Bohai Gu, Yueyang Yuan, Taiyi Wu +9
Interactive video world models are essential for long-horizon planning and exploration, yet they suffer from compounding errors. Post-training methods such as reinforcement learnin…
RedFlow: Redirect Failure into Action-Level Corrections for Flow-matching VLA Policy
Zhengyang Yan, Junhao Li, Fangqi Zhu +6
RedFlow is an offline reinforcement learning framework that turns failure experiences into action-level corrective supervision for flow-matching vision‑language‑action policies, im…
HALO: A Unified Vision-Language-Action Model for Embodied Multimodal Chain-of-Thought Reasoning
Quanxin Shou, Fangqi Zhu, Shawn Chen +9
Vision-Language-Action (VLA) models have shown strong performance in robotic manipulation, but often struggle in long-horizon or out-of-distribution scenarios due to the lack of ex…
Generating Storytelling Images with Rich Chains-of-Reasoning
Xiujie Song, Qi Jia, Shota Watanabe +4
A single image can convey a compelling story through logically connected visual clues, forming Chains-of-Reasoning (CoRs). We define these semantically rich images as Storytelling…
Is Your Image a Good Storyteller?
Xiujie Song, Xiaoyi Pang, Haifeng Tang +2
Quantifying image complexity at the entity level is straightforward, but the assessment of semantic complexity has been largely overlooked. In fact, there are differences in semant…