9 papers
Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers
Sicheng Mo, Yuheng Li, Ziyang Leng +2
Multi-agent interactive world models should not only generate consistent observations, but also maintain world states that persist across agents and evolve across views. Existing a…
UniTemp: Unlocking Video Generation in Any Temporal Order via Bidirectional Distillation
Lin Zhang, Sicheng Mo, Zefan Cai +6
Autoregressive video diffusion models have emerged as a promising approach for long video generation, achieving strong performance in streaming settings. However, existing methods…
Personal AI Agent for Camera Roll VQA
Thao Nguyen, Krishna Kumar Singh, Donghyun Kim +2
We study the personal camera roll visual question answering setting. In this setting, a conversational AI assistant can access a user's personal camera roll and retrieve relevant p…
MAOAM: Unified Object and Material Selection with Vision-Language Models
Jaden Park, Valentin Deschaintre, Jason Kuen +5
Selection is a core operation in interactive image editing. To be practical, a user should be able to specify and disambiguate the desired selection region through either text or c…
From Plans to Pixels: Learning to Plan and Orchestrate for Open-Ended Image Editing
Anirudh Sundara Rajan, Krishna Kumar Singh, Yong Jae Lee
Modern image editing models produce realistic results but struggle with abstract, multi step instructions (e.g., ``make this advertisement more vegetarian-friendly''). Prior agent…
Relational Visual Similarity
Thao Nguyen, Sicheng Mo, Krishna Kumar Singh +6
Humans do not just see attribute similarity -- we also see relational similarity. An apple is like a peach because both are reddish fruit, but the Earth is also like a peach: its c…