works on

From the 1 of 15 linked papers with an AI index.

collaborators

15 papers

cs.CV2026

DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

DreamX Team, Rui Chen, Xiangxiang Chu +7

We present \textbf{DreamX-Phi 1.0}, an action-conditioned video world model for robotic manipulation that, given an observed frame, a language instruction, and a prescribed action…

cs.CV2026

Peak-End-Net: A Peak-End Rule Inspired Framework for Generalizable Video Aesthetic Assessment

Geng Li, Haiwen Li, Rui Chen +3

The paper introduces Peak-End-Net, a lightweight framework that uses the psychological peak‑end rule to assess video aesthetics by combining frame‑wise aesthetic priors from a pret…

cs.MA2026

M2Note: Continual Evolution of Vision Language Models via Mistake Notebook Learning

Haiwen Li, Jing Tang, Rui Chen +2

Vision Language Models (VLMs) have demonstrated remarkable capabilities in multimodal reasoning tasks, yet they still suffer from recurring failures, such as skipping key visual ch…

cs.CV2026

DreamX-World 1.0: A General-Purpose Interactive World Model

DreamX Team, Yancheng Bai, Rui Chen +20

DreamX-World 1.0 is a general-purpose interactive text/image-to-video world model for controllable long-horizon generation. It supports camera navigation, revisits to previously ob…

cs.CV2026

What if Agents Could Imagine? Reinforcing Open-Vocabulary HOI Comprehension through Generation

Zhenlong Yuan, Yue Wang, Dapeng Zhang +9

Multimodal Large Language Models have shown promising capabilities in bridging visual and textual reasoning, yet their reasoning capabilities in Open-Vocabulary Human-Object Intera…

cs.CV2026

Video-STAR: Reinforcing Open-Vocabulary Action Recognition with Tools

Zhenlong Yuan, Xiangyan Qu, Chengxuan Qian +8

Multimodal large language models (MLLMs) have demonstrated remarkable potential in bridging visual and textual reasoning, yet their reliance on text-centric priors often limits the…