activity
20212026
most citedmPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video

50 citations · 87 across the 7 of their papers we have counts for

collaborators

5 papers

cs.AI2026

Generative Retrieval via Diffusion Transformer with Metric-Ordered Sequence Training and Hybrid-Policy Preference Optimization

Chenghao Liu, Yu Zhang, Zhongtao Jiang +7

Embedding-based retrieval ranks items by their similarity to a query in a shared vector space and usually aims to return the highest-scoring items. In many production settings this…

cs.CV20235 cited

RAMM: Retrieval-augmented Biomedical Visual Question Answering with Multi-modal Pre-training

Zheng Yuan, Qiao Jin, Chuanqi Tan +4

Vision-and-language multi-modal pretraining and fine-tuning have shown great success in visual question answering (VQA). Compared to general domain VQA, the performance of biomedic…

q-bio.MN20233 cited

Molecular Geometry-aware Transformer for accurate 3D Atomic System modeling

Zheng Yuan, Yaoyun Zhang, Chuanqi Tan +3

Molecular dynamic simulations are important in computational physics, chemistry, material, and biology. Machine learning-based methods have shown strong abilities in predicting mol…

cs.CV202350 cited

mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video

Haiyang Xu, Qinghao Ye, Ming Yan +12

Recent years have witnessed a big convergence of language, vision, and multi-modal pretraining. In this work, we present mPLUG-2, a new unified paradigm with modularized design for…

cs.CL2021

From Dense to Sparse: Contrastive Pruning for Better Pre-trained Language Model Compression

Runxin Xu, Fuli Luo, Chengyu Wang +4

Pre-trained Language Models (PLMs) have achieved great success in various Natural Language Processing (NLP) tasks under the pre-training and fine-tuning paradigm. With large quanti…