collaborators

7 papers

cs.CL2024

SocialBench: Sociality Evaluation of Role-Playing Conversational Agents

Hongzhan Chen, Hehong Chen, Ming Yan +8

Large language models (LLMs) have advanced the development of various AI conversational agents, including role-playing conversational agents that mimic diverse characters and human…

cs.CL2024

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Qinghao Ye, Haiyang Xu, Guohai Xu +15

Large language models (LLMs) have demonstrated impressive zero-shot abilities on a variety of open-ended tasks, while recent research has also explored the use of LLMs for multi-mo…

cs.LG2024

TiMix: Text-aware Image Mixing for Effective Vision-Language Pre-training

Chaoya Jiang, Wei ye, Haiyang Xu +4

Self-supervised Multi-modal Contrastive Learning (SMCL) remarkably advances modern Vision-Language Pre-training (VLP) models by aligning visual and linguistic modalities. Due to no…

cs.CV2024

Hallucination Augmented Contrastive Learning for Multimodal Large Language Model

Chaoya Jiang, Haiyang Xu, Mengfan Dong +7

Multi-modal large language models (MLLMs) have been shown to efficiently integrate natural language with visual information to handle multi-modal tasks. However, MLLMs still face a…

cs.MM2024

COPA: Efficient Vision-Language Pre-training Through Collaborative Object- and Patch-Text Alignment

Chaoya Jiang, Haiyang Xu, Wei Ye +7

Vision-Language Pre-training (VLP) methods based on object detection enjoy the rich knowledge of fine-grained object-text alignment but at the cost of computationally expensive inf…

cs.CL2024

AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation

Junyang Wang, Yuhang Wang, Guohai Xu +8

Despite making significant progress in multi-modal tasks, current Multi-modal Large Language Models (MLLMs) encounter the significant challenge of hallucinations, which may lead to…