10 papers
MOSAIK: Multi-Patch Content-Aware Spatial Allocation of Image Tokens for Efficient Generation
Mohammadreza Hami, Mohammadreza Samadi, Chao Gao +1
Pixel-space diffusion models avoid the reconstruction ceiling of latent diffusion models by generating directly in image space. However, their substantially higher token count make…
DBLAST: Dependent Block Drafting for Stochastic Speculative Decoding
Amirmohammad Karimi, Chao Gao, Negar Hassanpour
Speculative decoding accelerates large language models' inference by using a lightweight drafter to propose multiple future tokens and a target model to verify them. While recent b…
RePack then Refine: Efficient Diffusion Transformer with Vision Foundation Model
Guanfang Dong, Luke Schultz, Negar Hassanpour +1
Semantic-rich features from Vision Foundation Models (VFMs) have been leveraged to enhance Latent Diffusion Models (LDMs). However, raw VFM features are typically high-dimensional…
GLASS: Global-Local Aggregation for Inference-time Sparsification of LLMs
Amirmohsen Sattarifard, Sepehr Lavasani, Kunlin Zhang +5
Inference-time sparsification is a promising path to deploy large language models (LLMs) on resource-constrained devices, yet existing training-free methods typically estimate feed…
Griffin: Generative Reference and Layout Guided Image Composition
Aryan Mikaeili, Amirhossein Alimohammadi, Negar Hassanpour +2
Text-to-image models have achieved a level of realism that enables the generation of highly convincing images. However, text-based control can be a limiting factor when more explic…
Cora: Correspondence-aware image editing using few step diffusion
Amirhossein Alimohammadi, Aryan Mikaeili, Sauradip Nag +3
Image editing is an important task in computer graphics, vision, and VFX, with recent diffusion-based methods achieving fast and high-quality results. However, edits requiring sign…