6 papers
A Comprehensive Survey and Guide to Multimodal Large Language Models in Vision-Language Tasks
Chia Xin Liang, Pu Tian, Caitlyn Heqi Yin +7
This survey and application guide to multimodal large language models(MLLMs) explores the rapidly developing field of MLLMs, examining their architectures, applications, and impact…
MMLongCite: A Benchmark for Evaluating Fidelity of Long-Context Vision-Language Models
Keyan Zhou, Zecheng Tang, Lingfeng Ming +8
The rapid advancement of large vision language models (LVLMs) has led to a significant expansion of their context windows. However, an extended context window does not guarantee th…
Do Thinking Tokens Help or Trap? Towards More Efficient Large Reasoning Model
Bowen Ding, Yuhan Chen, Futing Wang +2
Large Reasoning Models (LRMs) excel at solving complex problems but face an overthinking dilemma. When handling simple tasks, they often produce verbose responses overloaded with t…
Ocean-OCR: Towards General OCR Application via a Vision-Language Model
Song Chen, Xinyu Guo, Yadong Li +10
Multimodal large language models (MLLMs) have shown impressive capabilities across various domains, excelling in processing and understanding information from multiple modalities.…
Baichuan-Omni-1.5 Technical Report
Yadong Li, Jun Liu, Tao Zhang +89
We introduce Baichuan-Omni-1.5, an omni-modal model that not only has omni-modal understanding capabilities but also provides end-to-end audio generation capabilities. To achieve f…
Marco-LLM: Bridging Languages via Massive Multilingual Training for Cross-Lingual Enhancement
Lingfeng Ming, Bo Zeng, Chenyang Lyu +17
Large Language Models (LLMs) have achieved remarkable progress in recent years; however, their excellent performance is still largely limited to major world languages, primarily En…