Showing cs.CVShow all
3 papers · 1 filter
cs.CV2024
Awaker2.5-VL: Stably Scaling MLLMs with Parameter-Efficient Mixture of Experts
Jinqiang Long, Yanqi Dai, Guoxing Yang +4
As the research of Multimodal Large Language Models (MLLMs) becomes popular, an advancing MLLM model is typically required to handle various textual and visual tasks (e.g., VQA, De…
cs.CV2023
VDT: General-purpose Video Diffusion Transformers via Mask Modeling
Haoyu Lu, Guoxing Yang, Nanyi Fei +4
This work introduces Video Diffusion Transformer (VDT), which pioneers the use of transformers in diffusion-based video generation. It features transformer blocks with modularized…
cs.CV2021
WenLan: Bridging Vision and Language by Large-Scale Multi-Modal Pre-Training
Yuqi Huo, Manli Zhang, Guangzhen Liu +32
Multi-modal pre-training models have been intensively explored to bridge vision and language in recent years. However, most of them explicitly model the cross-modal interaction bet…