efficient training 1image generation 1image understanding 1unified multimodal model 1vision-language models 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.CV2026
Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation
Weiming Zhuang, Jiabo Huang, Jingtao Li +4
The paper introduces Argus-Unified, a compact multimodal model that combines image understanding and generation by leveraging pretrained vision-language models and hybrid visual to…
cs.LG2026
StelLA: Subspace Learning in Low-rank Adaptation using Stiefel Manifold
Zhizhong Li, Sina Sajadmanesh, Jingtao Li +1
Low-rank adaptation (LoRA) has been widely adopted as a parameter-efficient technique for fine-tuning large-scale pre-trained models. However, it still lags behind full fine-tuning…
cs.CV2025
ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model
Weitai Kang, Weiming Zhuang, Zhizhong Li +2
Fine-grained multimodal capability in Multimodal Large Language Models (MLLMs) has emerged as a critical research direction, particularly for tackling the visual grounding (VG) pro…