From the 1 of 15 linked papers with an AI index.
15 papers
DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
DreamX Team, Rui Chen, Xiangxiang Chu +7
We present \textbf{DreamX-Phi 1.0}, an action-conditioned video world model for robotic manipulation that, given an observed frame, a language instruction, and a prescribed action…
Peak-End-Net: A Peak-End Rule Inspired Framework for Generalizable Video Aesthetic Assessment
Geng Li, Haiwen Li, Rui Chen +3
The paper introduces Peak-End-Net, a lightweight framework that uses the psychological peak‑end rule to assess video aesthetics by combining frame‑wise aesthetic priors from a pret…
M2Note: Continual Evolution of Vision Language Models via Mistake Notebook Learning
Haiwen Li, Jing Tang, Rui Chen +2
Vision Language Models (VLMs) have demonstrated remarkable capabilities in multimodal reasoning tasks, yet they still suffer from recurring failures, such as skipping key visual ch…
DreamX-World 1.0: A General-Purpose Interactive World Model
DreamX Team, Yancheng Bai, Rui Chen +20
DreamX-World 1.0 is a general-purpose interactive text/image-to-video world model for controllable long-horizon generation. It supports camera navigation, revisits to previously ob…
What if Agents Could Imagine? Reinforcing Open-Vocabulary HOI Comprehension through Generation
Zhenlong Yuan, Yue Wang, Dapeng Zhang +9
Multimodal Large Language Models have shown promising capabilities in bridging visual and textual reasoning, yet their reasoning capabilities in Open-Vocabulary Human-Object Intera…
Video-STAR: Reinforcing Open-Vocabulary Action Recognition with Tools
Zhenlong Yuan, Xiangyan Qu, Chengxuan Qian +8
Multimodal large language models (MLLMs) have demonstrated remarkable potential in bridging visual and textual reasoning, yet their reliance on text-centric priors often limits the…