28 papers
RankT2I: A Submodular Framework for Discovering Interpretable and Diverse Semantics in Text-to-Image Models
Ritika Allada, Pinar Yanardag
Recent advances in text-to-image (T2I) models have revolutionized the field of image generation and editing. However, identifying semantics that a T2I model can successfully edit i…
Learn Once, Edit Anywhere: Visual Direction Transfer for Diffusion Models
Yusuf Dalva, Hidir Yesiltepe, Pinar Yanardag
The rapid advancement of diffusion models has enabled the generation of high-fidelity images from textual prompts, yet achieving precise, disentangled control over specific attribu…
From Zero to Hero: Training-Free Custom Concept Spawning in World Models
Kiymet Akdemir, Pinar Yanardag
Autoregressive world models have emerged as a powerful paradigm for interactive video generation, allowing users to navigate dynamically generated environments through actions. The…
DTG-Restore: Training-Free Diffusion Refinement for Generative Video Super-Resolution
Hidir Yesiltepe, Koutilya PNVR, Gaurav Pathak +4
Recent progress in video diffusion models has enabled remarkable generative fidelity, yet leveraging these priors for restoration remains limited by the strong coupling between con…
VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion
Hidir Yesiltepe, Jiazhen Hu, Tuna Han Salih Meral +4
Long-rollout causal video diffusion has converged on a fixed-size sliding-window KV cache, with recent progress innovating within this layout by changing which tokens occupy the wi…
AdaState: Self-Evolving Anchors for Streaming Video Generation
Yusuf Dalva, Pinar Yanardag
Autoregressive video diffusion models generate streaming video by producing frames sequentially, conditioning each chunk on previously generated content. These models are structura…