4 papers
CAPTAIN: Semantic Feature Injection for Memorization Mitigation in Text-to-Image Diffusion Models
Tong Zhang, Carlos Hinojosa, Bernard Ghanem
Diffusion models can unintentionally reproduce training examples, raising privacy and copyright concerns as these systems are increasingly deployed at scale. Existing inference-tim…
MoCA-Video: Motion-Aware Concept Alignment for Consistent Video Editing
Tong Zhang, Juan C Leon Alcazar, Victor Escorcia +1
We present MoCA-Video, a training-free framework for semantic mixing in videos. Operating in the latent space of a frozen video diffusion model, MoCA-Video utilizes class-agnostic…
MatchDiffusion: Training-free Generation of Match-cuts
Alejandro Pardo, Fabio Pizzati, Tong Zhang +4
Match-cuts are powerful cinematic tools that create seamless transitions between scenes, delivering strong visual and metaphorical connections. However, crafting match-cuts is a ch…
Beyond Pixels: Exploring Human-Readable SVG Generation for Simple Images with Vision Language Models
Tong Zhang, Haoyang Liu, Peiyan Zhang +2
In the field of computer graphics, the use of vector graphics, particularly Scalable Vector Graphics (SVG), represents a notable development from traditional pixel-based imagery. S…