most citedMoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models

4 citations · 10 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CV20242 cited

Token-level Correlation-guided Compression for Efficient Multimodal Document Understanding

Renshan Zhang, Yibo Lyu, Rui Shao +3

Cropping high-resolution document images into multiple sub-images is the most widely used approach for current Multimodal Large Language Models (MLLMs) to do document understanding…

cs.CV20244 cited

MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models

Leyang Shen, Gongwei Chen, Rui Shao +2

Multimodal large language models (MLLMs) have demonstrated impressive capabilities across various vision-language tasks. However, a generalist MLLM typically underperforms compared…

cs.IR2024

Prompt-based Multi-interest Learning Method for Sequential Recommendation

Xue Dong, Xuemeng Song, Tongliang Liu +1

Multi-interest learning method for sequential recommendation aims to predict the next item according to user multi-faceted interests given the user historical interactions. Existin…

cs.CV20232 cited

Knowledge-Aware Prompt Tuning for Generalizable Vision-Language Models

Baoshuo Kan, Teng Wang, Wenpeng Lu +3

Pre-trained vision-language models, e.g., CLIP, working with manually designed prompts have demonstrated great capacity of transfer learning. Recently, learnable prompts achieve st…

cs.CV2023

Contrast-augmented Diffusion Model with Fine-grained Sequence Alignment for Markup-to-Image Generation

Guojin Zhong, Jin Yuan, Pan Wang +3

The recently rising markup-to-image generation poses greater challenges as compared to natural image generation, due to its low tolerance for errors as well as the complex sequence…

cs.CV20232 cited

Set-level Guidance Attack: Boosting Adversarial Transferability of Vision-Language Pre-training Models

Dong Lu, Zhiqiang Wang, Teng Wang +3

Vision-language pre-training (VLP) models have shown vulnerability to adversarial examples in multimodal tasks. Furthermore, malicious adversaries can be deliberately transferred t…