13 papers
MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
Mingqiao Ye, Zhaochong An, Zhitong Gao +11
Any-to-any models predict any modality from any combination of others within a single network, a formulation used in multimodal vision and vision-language models, and increasingly…
VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward
Zhaochong An, Orest Kupyn, Théo Uscidda +5
Large-scale video diffusion models achieve impressive visual quality, yet often fail to preserve geometric consistency. Prior approaches improve consistency either by augmenting th…
RAIGen: Rare Attribute Identification in Text-to-Image Generative Models
Silpa Vadakkeeveetil Sreelatha, Dan Wang, Serge Belongie +2
Text-to-image diffusion models achieve impressive generation quality but inherit and amplify training-data biases, skewing coverage of semantic attributes. Prior work addresses thi…
Track2View: 4D-Consistent Camera-Controlled Video Generation via Paired 3D Point Tracks
Feng Qiao, Zhaochong An, Zhexiao Xiong +2
Re-rendering an existing video from a novel camera viewpoint requires the output to follow the prescribed camera trajectory while preserving the appearance and dynamics of the orig…
The Latent Color Subspace: Emergent Order in High-Dimensional Chaos
Mateusz Pach, Jessica Bader, Quentin Bouniot +2
Text-to-image generation models have advanced rapidly, yet achieving fine-grained control over generated images remains difficult, largely due to limited understanding of how seman…
Beyond Binary Success: A Diagnostic Meta-Evaluation Framework for Fine-Grained Manipulation
He-Yang Xu, Pengyuan Zhang, Zongyuan Ge +5
Fine-grained manipulation marks a regime where global scene context no longer suffices, and success hinges on the tight coupling of local attribute grounding, high-fidelity spatial…