adaptive inference 1agentic visual reasoning 1diffusion models 1multimodal language models 1multimodal large language models 1multi-reference video editing 1reference tokens 1reinforcement learning 1structured instructions 1tool use adaptiveness 1
From the 2 of 41 linked papers with an AI index.
Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Edit-R2: Context-Aware Reinforcement Learning for Multi-Turn Image Editing
Yuxiao Ye, Haoran He, Fangyuan Kong +4
Text-guided image editing has advanced rapidly with diffusion models and unified multimodal foundation models. However, most existing methods remain confined to single-turn setting…
cs.AI2026
RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark
Yang Shi, Yuhao Dong, Yue Ding +22
The integration of visual understanding and generation into unified multimodal models represents a significant stride toward general-purpose AI. However, a fundamental question rem…
cs.AI2026
Simulating the Visual World with Artificial Intelligence: A Roadmap
Jingtong Yue, Ziqi Huang, Zhaoxi Chen +3
The landscape of video generation is shifting, from a focus on generating visually appealing clips to building virtual environments that support interaction and maintain physical p…