11 papers
3D Scene-Adaptive Trajectory-Controllable Human Image Animation with Camera Movement
Deyin Liu, Jicheng Xu, Lin Yuanbo Wu +4
Human image animation, which aims to generate a video of a reference subject following a provided action sequence, has received increasing research interest. With the development o…
ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation
Anindya Mondal, Sauradip Nag, Anjan Dutta
ABACUS is a unified vision-language model that handles object counting, crowd counting, referring-expression counting, and count-faithful image generation without any benchmark-spe…
RAIGen: Rare Attribute Identification in Text-to-Image Generative Models
Silpa Vadakkeeveetil Sreelatha, Dan Wang, Serge Belongie +2
Text-to-image diffusion models achieve impressive generation quality but inherit and amplify training-data biases, skewing coverage of semantic attributes. Prior work addresses thi…
Beyond Consistency: Preserving Temporal Structure in Zero-Shot Video Editing
Deyin Liu, Yisheng Ding, Zhe Jin +3
Existing zero-shot video editing methods rely on pre-trained diffusion models, successfully achieving spatial control and basic temporal consistency but fundamentally fail to prese…
TaleDiffusion: Multi-Character Story Generation with Dialogue Rendering
Ayan Banerjee, Josep Llados, Umapada Pal +1
Text-to-story visualization is challenging due to the need for consistent interaction among multiple characters across frames. Existing methods struggle with character consistency,…
CountLoop: Training-Free High-Instance Image Generation via Iterative Agent Guidance
Anindya Mondal, Ayan Banerjee, Sauradip Nag +3
Diffusion models excel at photorealistic synthesis but struggle with precise object counts, especially in high-density settings. We introduce COUNTLOOP, a training-free framework t…