5 papers
S2DiT: Sandwich Diffusion Transformer for Mobile Streaming Video Generation
Lin Zhao, Yushu Wu, Aleksei Lebedev +11
Diffusion Transformers (DiTs) have recently improved video generation quality. However, their heavy computational cost makes real-time or on-device generation infeasible. In this w…
Test-Time Computing for Referring Multimodal Large Language Models
Mingrui Wu, Hao Chen, Jiayi Ji +5
We propose ControlMLLM++, a novel test-time adaptation framework that injects learnable visual prompts into frozen multimodal large language models (MLLMs) to enable fine-grained r…
Prompt Optimization Via Diffusion Language Models
Shiyu Wang, Haolin Chen, Liangwei Yang +8
We propose a diffusion-based framework for prompt optimization that leverages Diffusion Language Models (DLMs) to iteratively refine system prompts through masked denoising. By con…
AlcheMinT: Fine-grained Temporal Control for Multi-Reference Consistent Video Generation
Sharath Girish, Viacheslav Ivanov, Tsai-Shien Chen +3
Recent advances in subject-driven video generation with large diffusion models have enabled personalized content synthesis conditioned on user-provided subjects. However, existing…
Zero-Shot Dynamic Concept Personalization with Grid-Based LoRA
Rameen Abdal, Or Patashnik, Ekaterina Deyneka +5
Recent advances in text-to-video generation have enabled high-quality synthesis from text and image prompts. While the personalization of dynamic concepts, which capture subject-sp…