works on

From the 1 of 5 linked papers with an AI index.

collaborators

5 papers

cs.RO2026

Motubrain: An Advanced World Action Model for Robot Control

Motubrain Team, Chendong Xiang, Fan Bao +17

Motubrain is a unified world action model that jointly learns video and robot actions using a UniDiffuser and Mixture-of-Transformers architecture, enabling policy learning, world…

cs.CV2026

LatentUM: Unleashing the Potential of Interleaved Cross-Modal Reasoning via a Latent-Space Unified Model

Jiachun Jin, Zetong Zhou, Xiao Yang +4

Unified models (UMs) hold promise for their ability to understand and generate content across heterogeneous modalities. Compared to merely generating visual content, the use of UMs…

cs.CV2025

Scaling Group Inference for Diverse and High-Quality Generation

Gaurav Parmar, Or Patashnik, Daniil Ostashev +4

Generative models typically sample outputs independently, and recent inference-time guidance and scaling algorithms focus on improving the quality of individual samples. However, i…

cs.CV2025

Multi-subject Open-set Personalization in Video Generation

Tsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace +7

Video personalization methods allow us to synthesize videos with specific concepts such as people, pets, and places. However, existing methods often focus on limited domains, requi…

cs.CV2025

Object-level Visual Prompts for Compositional Image Generation

Gaurav Parmar, Or Patashnik, Kuan-Chieh Wang +5

We introduce a method for composing object-level visual prompts within a text-to-image diffusion model. Our approach addresses the task of generating semantically coherent composit…