From the 1 of 28 linked papers with an AI index.
28 papers
Evaluation-Verification Reward for Consistent Multi-Reference Image Editing
Yingmao Miao, Pengfei Zhang, Xiaochen Lv +5
While recent image editing models have made rapid progress, multi-reference editing remains challenging, particularly in maintaining visual consistency across references and ensuri…
Peak-End-Net: A Peak-End Rule Inspired Framework for Generalizable Video Aesthetic Assessment
Geng Li, Haiwen Li, Rui Chen +3
The paper introduces Peak-End-Net, a lightweight framework that uses the psychological peak‑end rule to assess video aesthetics by combining frame‑wise aesthetic priors from a pret…
M2Note: Continual Evolution of Vision Language Models via Mistake Notebook Learning
Haiwen Li, Jing Tang, Rui Chen +2
Vision Language Models (VLMs) have demonstrated remarkable capabilities in multimodal reasoning tasks, yet they still suffer from recurring failures, such as skipping key visual ch…
Edit in 2D, Verify in 3D: Reinforcement Learning for Multi-view Consistent Scene Editing
Jiyuan Wang, Chunyu Lin, Lei Sun +8
Leveraging the priors of 2D diffusion models for 3D editing has emerged as a promising paradigm. However, multi-view consistency remains challenging in edited results, and the extr…
DreamX-World 1.0: A General-Purpose Interactive World Model
DreamX Team, Yancheng Bai, Rui Chen +20
DreamX-World 1.0 is a general-purpose interactive text/image-to-video world model for controllable long-horizon generation. It supports camera navigation, revisits to previously ob…
What if Agents Could Imagine? Reinforcing Open-Vocabulary HOI Comprehension through Generation
Zhenlong Yuan, Yue Wang, Dapeng Zhang +9
Multimodal Large Language Models have shown promising capabilities in bridging visual and textual reasoning, yet their reasoning capabilities in Open-Vocabulary Human-Object Intera…