From the 1 of 8 linked papers with an AI index.
8 papers
Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation
Weiming Zhuang, Jiabo Huang, Jingtao Li +4
The paper introduces Argus-Unified, a compact multimodal model that combines image understanding and generation by leveraging pretrained vision-language models and hybrid visual to…
Empirical Recipes for Efficient and Compact Vision-Language Models
Jiabo Huang, Zhizhong Li, Sina Sajadmanesh +2
Deploying vision-language models (VLMs) in resource-constrained settings demands low latency and high throughput, yet existing compact VLMs often fall short of the inference speedu…
UniCompress: Token Compression for Unified Vision-Language Understanding and Generation
Ziyao Wang, Chen Chen, Jingtao Li +4
Unified models aim to support both understanding and generation by encoding images into discrete tokens and processing them alongside text within a single autoregressive framework.…
Neuro-Symbolic Spatial Reasoning in Segmentation
Jiayi Lin, Jiabo Huang, Shaogang Gong
Open-Vocabulary Semantic Segmentation (OVSS) assigns pixel-level labels from an open set of categories, requiring generalization to unseen and unlabelled objects. Using vision-lang…
Seeing Further on the Shoulders of Giants: Knowledge Inheritance for Vision Foundation Models
Jiabo Huang, Chen Chen, Lingjuan Lyu
Vision foundation models (VFMs) are predominantly developed using data-centric methods. These methods require training on vast amounts of data usually with high-quality labels, whi…
UNIFORM: Unifying Knowledge from Large-scale and Diverse Pre-trained Models
Yimu Wang, Weiming Zhuang, Chen Chen +3
In the era of deep learning, the increasing number of pre-trained models available online presents a wealth of knowledge. These models, developed with diverse architectures and tra…