activity
20212024
most citedmPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video

50 citations · 178 across the 34 of their papers we have counts for

collaborators

34 papers

cs.CL2024

A Simple yet Effective Training-free Prompt-free Approach to Chinese Spelling Correction Based on Large Language Models

Houquan Zhou, Zhenghua Li, Bo Zhang +5

This work proposes a simple training-free prompt-free approach to leverage large language models (LLMs) for the Chinese spelling correction (CSC) task, which is totally different f…

hep-ph2024

Phenomenological study of heavy neutral gauge boson in the left-right symmetric model at future muon collider

Zongyang Lu, Jianing Qin, Honglei Li +5

The exotic neutral gauge boson is a powerful candidate for the new physics beyond the standard model. As a promising model, the left-right symmetric model has been proposed to expl…

cs.CV2024

Platypus: A Generalized Specialist Model for Reading Text in Various Forms

Peng Wang, Zhaohai Li, Jun Tang +4

Reading text from images (either natural scenes or documents) has been a long-standing research topic for decades, due to the high technical challenge and wide application range. P…

cs.CV20246 cited

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Jiabo Ye, Haiyang Xu, Haowei Liu +6

Multi-modal Large Language Models (MLLMs) have demonstrated remarkable capabilities in executing instructions for a variety of single-image tasks. Despite this progress, significan…

cs.AI20241 cited

ReAct Meets ActRe: When Language Agents Enjoy Training Data Autonomy

Zonghan Yang, Peng Li, Ming Yan +3

Language agents have demonstrated autonomous decision-making abilities by reasoning with foundation models. Recently, efforts have been made to train language agents for performanc…

cs.CV20241 cited

OmniParser: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition

Jianqiang Wan, Sibo Song, Wenwen Yu +6

Recently, visually-situated text parsing (VsTP) has experienced notable advancements, driven by the increasing demand for automated document understanding and the emergence of Gene…