works on

From the 1 of 28 linked papers with an AI index.

collaborators

28 papers

cs.CV2026

Evaluation-Verification Reward for Consistent Multi-Reference Image Editing

Yingmao Miao, Pengfei Zhang, Xiaochen Lv +5

While recent image editing models have made rapid progress, multi-reference editing remains challenging, particularly in maintaining visual consistency across references and ensuri…

cs.CV2026

Peak-End-Net: A Peak-End Rule Inspired Framework for Generalizable Video Aesthetic Assessment

Geng Li, Haiwen Li, Rui Chen +3

The paper introduces Peak-End-Net, a lightweight framework that uses the psychological peak‑end rule to assess video aesthetics by combining frame‑wise aesthetic priors from a pret…

cs.MA2026

M2Note: Continual Evolution of Vision Language Models via Mistake Notebook Learning

Haiwen Li, Jing Tang, Rui Chen +2

Vision Language Models (VLMs) have demonstrated remarkable capabilities in multimodal reasoning tasks, yet they still suffer from recurring failures, such as skipping key visual ch…

cs.CV2026

Edit in 2D, Verify in 3D: Reinforcement Learning for Multi-view Consistent Scene Editing

Jiyuan Wang, Chunyu Lin, Lei Sun +8

Leveraging the priors of 2D diffusion models for 3D editing has emerged as a promising paradigm. However, multi-view consistency remains challenging in edited results, and the extr…

cs.CV2026

DreamX-World 1.0: A General-Purpose Interactive World Model

DreamX Team, Yancheng Bai, Rui Chen +20

DreamX-World 1.0 is a general-purpose interactive text/image-to-video world model for controllable long-horizon generation. It supports camera navigation, revisits to previously ob…

cs.CV2026

What if Agents Could Imagine? Reinforcing Open-Vocabulary HOI Comprehension through Generation

Zhenlong Yuan, Yue Wang, Dapeng Zhang +9

Multimodal Large Language Models have shown promising capabilities in bridging visual and textual reasoning, yet their reasoning capabilities in Open-Vocabulary Human-Object Intera…