activity
20192023
most citedTowards End-to-End Embodied Decision Making via Multi-modal Large Language Model: Explorations with GPT4-Vision and Beyond

4 citations · 5 across the 4 of their papers we have counts for

collaborators

6 papers

cs.AI20234 cited

Towards End-to-End Embodied Decision Making via Multi-modal Large Language Model: Explorations with GPT4-Vision and Beyond

Liang Chen, Yichi Zhang, Shuhuai Ren +6

In this study, we explore the potential of Multimodal Large Language Models (MLLMs) in improving embodied decision-making processes for agents. While Large Language Models (LLMs) h…

cs.CL2021

Dynamic Knowledge Distillation for Pre-trained Language Models

Lei Li, Yankai Lin, Shuhuai Ren +3

Knowledge distillation~(KD) has been proved effective for compressing large-scale pre-trained language models. However, existing methods conduct KD statically, e.g., the student mo…

cs.CL2021

Text AutoAugment: Learning Compositional Augmentation Policy for Text Classification

Shuhuai Ren, Jinchao Zhang, Lei Li +2

Data augmentation aims to enrich training samples for alleviating the overfitting issue in low-resource or class-imbalanced situations. Traditional methods first devise task-specif…

cs.CL20211 cited

Learning Relation Alignment for Calibrated Cross-modal Retrieval

Shuhuai Ren, Junyang Lin, Guangxiang Zhao +5

Despite the achievements of large-scale multimodal pre-training approaches, cross-modal retrieval, e.g., image-text retrieval, remains a challenging task. To bridge the semantic ga…

cs.CL2020

CascadeBERT: Accelerating Inference of Pre-trained Language Models via Calibrated Complete Models Cascade

Lei Li, Yankai Lin, Deli Chen +4

Dynamic early exiting aims to accelerate the inference of pre-trained language models (PLMs) by emitting predictions in internal layers without passing through the entire model. In…

cs.CV2019

DCA: Diversified Co-Attention towards Informative Live Video Commenting

Zhihan Zhang, Zhiyi Yin, Shuhuai Ren +2

We focus on the task of Automatic Live Video Commenting (ALVC), which aims to generate real-time video comments with both video frames and other viewers' comments as inputs. A majo…