13 citations · 18 across the 2 of their papers we have counts for
4 papers
UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Jiabo Ye, Anwen Hu, Haiyang Xu +11
Text is ubiquitous in our visual world, conveying crucial information, such as in documents, websites, and everyday photographs. In this work, we propose UReader, a first explorati…
COPA: Efficient Vision-Language Pre-training Through Collaborative Object- and Patch-Text Alignment
Chaoya Jiang, Haiyang Xu, Wei Ye +7
Vision-Language Pre-training (VLP) methods based on object detection enjoy the rich knowledge of fine-grained object-text alignment but at the cost of computationally expensive inf…
CValues: Measuring the Values of Chinese Large Language Models from Safety to Responsibility
Guohai Xu, Jiayi Liu, Ming Yan +11
With the rapid evolution of large language models (LLMs), there is a growing concern that they may pose risks or have negative social impacts. Therefore, evaluation of human values…
mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Jiabo Ye, Anwen Hu, Haiyang Xu +10
Document understanding refers to automatically extract, analyze and comprehend information from various types of digital documents, such as a web page. Existing Multi-model Large L…