5 papers
DeepEyesV2: Toward Agentic Multimodal Model
Jack Hong, Chenxiao Zhao, ChengLin Zhu +3
Agentic multimodal models should not only comprehend text and images, but also actively invoke external tools, such as code execution environments and web search, and integrate the…
EduFlow: Advancing MLLMs' Problem-Solving Proficiency through Multi-Stage, Multi-Perspective Critique
Chenglin Zhu, Tao Zhang, Chong Li +3
Multimodal large language models (MLLMs) still perform poorly on scientific tasks, particularly those requiring multi-step and interpretable reasoning. Their limitations include in…
Efficient Medical VIE via Reinforcement Learning
Lijun Liu, Ruiyang Li, Zhaocheng Liu +5
Visual Information Extraction (VIE) converts unstructured document images into structured formats like JSON, critical for medical applications such as report analysis and online co…
K12Vista: Exploring the Boundaries of MLLMs in K-12 Education
Chong Li, Chenglin Zhu, Tao Zhang +3
Multimodal large language models have demonstrated remarkable reasoning capabilities in various visual tasks. However, their abilities in K12 scenarios are still systematically und…
Baichuan-Omni-1.5 Technical Report
Yadong Li, Jun Liu, Tao Zhang +89
We introduce Baichuan-Omni-1.5, an omni-modal model that not only has omni-modal understanding capabilities but also provides end-to-end audio generation capabilities. To achieve f…