From the 1 of 5 linked papers with an AI index.
5 papers
Motubrain: An Advanced World Action Model for Robot Control
Motubrain Team, Chendong Xiang, Fan Bao +17
Motubrain is a unified world action model that jointly learns video and robot actions using a UniDiffuser and Mixture-of-Transformers architecture, enabling policy learning, world…
LatentUM: Unleashing the Potential of Interleaved Cross-Modal Reasoning via a Latent-Space Unified Model
Jiachun Jin, Zetong Zhou, Xiao Yang +4
Unified models (UMs) hold promise for their ability to understand and generate content across heterogeneous modalities. Compared to merely generating visual content, the use of UMs…
Scaling Group Inference for Diverse and High-Quality Generation
Gaurav Parmar, Or Patashnik, Daniil Ostashev +4
Generative models typically sample outputs independently, and recent inference-time guidance and scaling algorithms focus on improving the quality of individual samples. However, i…
Multi-subject Open-set Personalization in Video Generation
Tsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace +7
Video personalization methods allow us to synthesize videos with specific concepts such as people, pets, and places. However, existing methods often focus on limited domains, requi…
Object-level Visual Prompts for Compositional Image Generation
Gaurav Parmar, Or Patashnik, Kuan-Chieh Wang +5
We introduce a method for composing object-level visual prompts within a text-to-image diffusion model. Our approach addresses the task of generating semantically coherent composit…