From the 1 of 11 linked papers with an AI index.
11 papers
DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
DreamX Team, Rui Chen, Xiangxiang Chu +7
We present \textbf{DreamX-Phi 1.0}, an action-conditioned video world model for robotic manipulation that, given an observed frame, a language instruction, and a prescribed action…
Evaluation-Verification Reward for Consistent Multi-Reference Image Editing
Yingmao Miao, Pengfei Zhang, Xiaochen Lv +5
While recent image editing models have made rapid progress, multi-reference editing remains challenging, particularly in maintaining visual consistency across references and ensuri…
Peak-End-Net: A Peak-End Rule Inspired Framework for Generalizable Video Aesthetic Assessment
Geng Li, Haiwen Li, Rui Chen +3
The paper introduces Peak-End-Net, a lightweight framework that uses the psychological peak‑end rule to assess video aesthetics by combining frame‑wise aesthetic priors from a pret…
M2Note: Continual Evolution of Vision Language Models via Mistake Notebook Learning
Haiwen Li, Jing Tang, Rui Chen +2
Vision Language Models (VLMs) have demonstrated remarkable capabilities in multimodal reasoning tasks, yet they still suffer from recurring failures, such as skipping key visual ch…
Edit in 2D, Verify in 3D: Reinforcement Learning for Multi-view Consistent Scene Editing
Jiyuan Wang, Chunyu Lin, Lei Sun +8
Leveraging the priors of 2D diffusion models for 3D editing has emerged as a promising paradigm. However, multi-view consistency remains challenging in edited results, and the extr…
DreamX-World 1.0: A General-Purpose Interactive World Model
DreamX Team, Yancheng Bai, Rui Chen +20
DreamX-World 1.0 is a general-purpose interactive text/image-to-video world model for controllable long-horizon generation. It supports camera navigation, revisits to previously ob…