papers

Publications (12)

cs.LG2024

Baichuan Alignment Technical Report

Mingan Lin, Fan Yang, Yanjun Shen +21

We introduce Baichuan Alignment, a detailed analysis of the alignment techniques employed in the Baichuan series of models. This represents the industry's first comprehensive accou…

cs.CL2025

CFBench: A Comprehensive Constraints-Following Benchmark for LLMs

Tao Zhang, Chenglin Zhu, Yanjun Shen +10

The adeptness of Large Language Models (LLMs) in comprehending and following natural language instructions is critical for their deployment in sophisticated real-world applications…

cs.CV2025

Ocean-OCR: Towards General OCR Application via a Vision-Language Model

Song Chen, Xinyu Guo, Yadong Li +10

Multimodal large language models (MLLMs) have shown impressive capabilities across various domains, excelling in processing and understanding information from multiple modalities.…

cs.CV2024

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge

Yaqi Zhao, Yuanyang Yin, Lin Li +7

Does seeing always mean knowing? Large Vision-Language Models (LVLMs) integrate separately pre-trained vision and language components, often using CLIP-ViT as vision backbone. Howe…

cs.CL2025

Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Tianpeng Li, Jun Liu, Tao Zhang +11

We introduce Baichuan-Audio, an end-to-end audio large language model that seamlessly integrates audio understanding and generation. It features a text-guided aligned speech genera…

cs.AI2024

Baichuan-Omni Technical Report

Yadong Li, Haoze Sun, Mingan Lin +23

The salient multimodal capabilities and interactive experience of GPT-4o highlight its critical role in practical applications, yet it lacks a high-performing open-source counterpa…

cs.AI2025

EduFlow: Advancing MLLMs' Problem-Solving Proficiency through Multi-Stage, Multi-Perspective Critique

Chenglin Zhu, Tao Zhang, Chong Li +3

Multimodal large language models (MLLMs) still perform poorly on scientific tasks, particularly those requiring multi-step and interpretable reasoning. Their limitations include in…

cs.CL2024

PAS: Data-Efficient Plug-and-Play Prompt Augmentation System

Miao Zheng, Hao Liang, Fan Yang +16

In recent years, the rise of Large Language Models (LLMs) has spurred a growing demand for plug-and-play AI systems. Among the various AI techniques, prompt engineering stands out…

cs.CL2025

Baichuan-Omni-1.5 Technical Report

Yadong Li, Jun Liu, Tao Zhang +89

We introduce Baichuan-Omni-1.5, an omni-modal model that not only has omni-modal understanding capabilities but also provides end-to-end audio generation capabilities. To achieve f…

cs.CL2024

BaichuanSEED: Sharing the Potential of ExtensivE Data Collection and Deduplication by Introducing a Competitive Large Language Model Baseline

Guosheng Dong, Da Pan, Yiding Sun +17

The general capabilities of Large Language Models (LLM) highly rely on the composition and selection on extensive pretraining datasets, treated as commercial secrets by several ins…

cs.CV2026

MathScape: Benchmarking Multimodal Large Language Models in Real-World Mathematical Contexts

Hao Liang, Linzhuang Sun, Minxuan Zhou +7

With the rapid progress of Multimodal LLMs, evaluating their mathematical reasoning capabilities has become an increasingly important research direction. In particular, visual-text…

cs.AI2025

K12Vista: Exploring the Boundaries of MLLMs in K-12 Education

Chong Li, Chenglin Zhu, Tao Zhang +3

Multimodal large language models have demonstrated remarkable reasoning capabilities in various visual tasks. However, their abilities in K12 scenarios are still systematically und…