8 papers
EduFlow: Advancing MLLMs' Problem-Solving Proficiency through Multi-Stage, Multi-Perspective Critique
Chenglin Zhu, Tao Zhang, Chong Li +3
Multimodal large language models (MLLMs) still perform poorly on scientific tasks, particularly those requiring multi-step and interpretable reasoning. Their limitations include in…
K12Vista: Exploring the Boundaries of MLLMs in K-12 Education
Chong Li, Chenglin Zhu, Tao Zhang +3
Multimodal large language models have demonstrated remarkable reasoning capabilities in various visual tasks. However, their abilities in K12 scenarios are still systematically und…
S2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models
Yuanbo Fang, Haoze Sun, Jun Liu +5
End-to-end speech large language models ((LLMs)) extend the capabilities of text-based models to directly process and generate audio tokens. However, this often leads to a decline…
CFBench: A Comprehensive Constraints-Following Benchmark for LLMs
Tao Zhang, Chenglin Zhu, Yanjun Shen +10
The adeptness of Large Language Models (LLMs) in comprehending and following natural language instructions is critical for their deployment in sophisticated real-world applications…
Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction
Tianpeng Li, Jun Liu, Tao Zhang +11
We introduce Baichuan-Audio, an end-to-end audio large language model that seamlessly integrates audio understanding and generation. It features a text-guided aligned speech genera…
Ocean-OCR: Towards General OCR Application via a Vision-Language Model
Song Chen, Xinyu Guo, Yadong Li +10
Multimodal large language models (MLLMs) have shown impressive capabilities across various domains, excelling in processing and understanding information from multiple modalities.…