most citedMIT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

14 citations · 16 across the 5 of their papers we have counts for

collaborators

5 papers

cs.AI2024

LaDiC: Are Diffusion Models Really Inferior to Autoregressive Counterparts for Image-to-Text Generation?

Yuchi Wang, Shuhuai Ren, Rundong Gao +5

Diffusion models have exhibited remarkable capabilities in text-to-image generation. However, their performance in image-to-text generation, specifically image captioning, has lagg…

cs.CV20242 cited

Towards Multimodal Video Paragraph Captioning Models Robust to Missing Modality

Sishuo Chen, Lei Li, Shuhuai Ren +5

Video paragraph captioning (VPC) involves generating detailed narratives for long videos, utilizing supportive modalities such as speech and event boundaries. However, the existing…

cs.CL2024

PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain

Liang Chen, Yichi Zhang, Shuhuai Ren +7

We present PCA-Bench, a multimodal decision-making benchmark for evaluating the integrated capabilities of Multimodal Large Language Models (MLLMs). Departing from previous benchma…

cs.CV2023

TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding

Shuhuai Ren, Sishuo Chen, Shicheng Li +2

Large-scale video-language pre-training has made remarkable strides in advancing video-language understanding tasks. However, the heavy computational burden of video encoding remai…

cs.CV202314 cited

MIT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Lei Li, Yuwei Yin, Shicheng Li +9

Instruction tuning has significantly advanced large language models (LLMs) such as ChatGPT, enabling them to align with human instructions across diverse tasks. However, progress i…