Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Edit-R2: Context-Aware Reinforcement Learning for Multi-Turn Image Editing
Yuxiao Ye, Haoran He, Fangyuan Kong +4
Text-guided image editing has advanced rapidly with diffusion models and unified multimodal foundation models. However, most existing methods remain confined to single-turn setting…
cs.AI2025
Simulating the Visual World with Artificial Intelligence: A Roadmap
Jingtong Yue, Ziqi Huang, Zhaoxi Chen +3
The landscape of video generation is shifting, from a focus on generating visually appealing clips to building virtual environments that support interaction and maintain physical p…
cs.AI2025
RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark
Yang Shi, Yuhao Dong, Yue Ding +22
The integration of visual understanding and generation into unified multimodal models represents a significant stride toward general-purpose AI. However, a fundamental question rem…