collaborators

6 papers

cs.AI2025

A Comprehensive Survey and Guide to Multimodal Large Language Models in Vision-Language Tasks

Chia Xin Liang, Pu Tian, Caitlyn Heqi Yin +7

This survey and application guide to multimodal large language models(MLLMs) explores the rapidly developing field of MLLMs, examining their architectures, applications, and impact…

cs.CV2025

MMLongCite: A Benchmark for Evaluating Fidelity of Long-Context Vision-Language Models

Keyan Zhou, Zecheng Tang, Lingfeng Ming +8

The rapid advancement of large vision language models (LVLMs) has led to a significant expansion of their context windows. However, an extended context window does not guarantee th…

cs.CL2025

Do Thinking Tokens Help or Trap? Towards More Efficient Large Reasoning Model

Bowen Ding, Yuhan Chen, Futing Wang +2

Large Reasoning Models (LRMs) excel at solving complex problems but face an overthinking dilemma. When handling simple tasks, they often produce verbose responses overloaded with t…

cs.CV2025

Ocean-OCR: Towards General OCR Application via a Vision-Language Model

Song Chen, Xinyu Guo, Yadong Li +10

Multimodal large language models (MLLMs) have shown impressive capabilities across various domains, excelling in processing and understanding information from multiple modalities.…

cs.CL2025

Baichuan-Omni-1.5 Technical Report

Yadong Li, Jun Liu, Tao Zhang +89

We introduce Baichuan-Omni-1.5, an omni-modal model that not only has omni-modal understanding capabilities but also provides end-to-end audio generation capabilities. To achieve f…

cs.CL2024

Marco-LLM: Bridging Languages via Massive Multilingual Training for Cross-Lingual Enhancement

Lingfeng Ming, Bo Zeng, Chenyang Lyu +17

Large Language Models (LLMs) have achieved remarkable progress in recent years; however, their excellent performance is still largely limited to major world languages, primarily En…