From the 1 of 12 linked papers with an AI index.
12 papers
Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation
Weiming Zhuang, Jiabo Huang, Jingtao Li +4
The paper introduces Argus-Unified, a compact multimodal model that combines image understanding and generation by leveraging pretrained vision-language models and hybrid visual to…
On the Limits of Token Reduction for Efficient Unified Vision Language Training
Siyi Chen, Weiming Zhuang, Jingtao Li +1
Unified vision-language models (VLMs) integrate visual understanding and visual generation within a single autoregressive backbone, but their joint training is computationally expe…
VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations
Maitreya Patel, Jingtao Li, Weiming Zhuang +2
We introduce an efficient, resolution-agnostic autoregressive (AR) image synthesis approach that generalizes to arbitrary resolutions and aspect ratios, narrowing the gap to diffus…
Replay-Free Continual Low-Rank Adaptation with Dynamic Memory
Huancheng Chen, Jingtao Li, Weiming Zhuang +2
We revisit continual learning~(CL), which enables pre-trained vision transformers (ViTs) to sequentially fine-tune on new downstream tasks over time. However, as the scale of these…
Training-Free Layout-to-Image Generation with Marginal Attention Constraints
Huancheng Chen, Jingtao Li, Weiming Zhuang +2
Recently, many text-to-image diffusion models have excelled at generating high-resolution images from text but struggle with precise control over spatial composition and object cou…
Empirical Recipes for Efficient and Compact Vision-Language Models
Jiabo Huang, Zhizhong Li, Sina Sajadmanesh +2
Deploying vision-language models (VLMs) in resource-constrained settings demands low latency and high throughput, yet existing compact VLMs often fall short of the inference speedu…