7 papers
MOSAIK: Multi-Patch Content-Aware Spatial Allocation of Image Tokens for Efficient Generation
Mohammadreza Hami, Mohammadreza Samadi, Chao Gao +1
Pixel-space diffusion models avoid the reconstruction ceiling of latent diffusion models by generating directly in image space. However, their substantially higher token count make…
DBLAST: Dependent Block Drafting for Stochastic Speculative Decoding
Amirmohammad Karimi, Chao Gao, Negar Hassanpour
Speculative decoding accelerates large language models' inference by using a lightweight drafter to propose multiple future tokens and a target model to verify them. While recent b…
GLASS: Global-Local Aggregation for Inference-time Sparsification of LLMs
Amirmohsen Sattarifard, Sepehr Lavasani, Kunlin Zhang +5
Inference-time sparsification is a promising path to deploy large language models (LLMs) on resource-constrained devices, yet existing training-free methods typically estimate feed…
Principled Fast and Meta Knowledge Learners for Continual Reinforcement Learning
Ke Sun, Hongming Zhang, Jun Jin +4
Inspired by the human learning and memory system, particularly the interplay between the hippocampus and cerebral cortex, this study proposes a dual-learner framework comprising a…
RAISE: Requirement-Adaptive Evolutionary Refinement for Training-Free Text-to-Image Alignment
Liyao Jiang, Ruichen Chen, Chao Gao +1
Recent text-to-image (T2I) diffusion models achieve remarkable realism, yet faithful prompt-image alignment remains challenging, particularly for complex prompts with multiple obje…
Grounding Degradations in Natural Language for All-In-One Video Restoration
Muhammad Kamran Janjua, Amirhosein Ghasemabadi, Kunlin Zhang +3
In this work, we propose an all-in-one video restoration framework that grounds degradation-aware semantic context of video frames in natural language via foundation models, offeri…