5 papers
MedConsultBench: A Full-Cycle, Fine-Grained, Process-Aware Benchmark for Medical Consultation Agents
Chuhan Qiao, Jianghua Huang, Daxing Zhao +5
Current evaluations of medical consultation agents often prioritize outcome-oriented tasks, frequently overlooking the end-to-end process integrity and clinical safety essential fo…
CFBench: A Comprehensive Constraints-Following Benchmark for LLMs
Tao Zhang, Chenglin Zhu, Yanjun Shen +10
The adeptness of Large Language Models (LLMs) in comprehending and following natural language instructions is critical for their deployment in sophisticated real-world applications…
Baichuan 2: Open Large-scale Language Models
Aiyuan Yang, Bin Xiao, Bingning Wang +52
Large language models (LLMs) have demonstrated remarkable performance on a variety of natural language tasks based on just a few examples of natural language instructions, reducing…
Baichuan-Omni Technical Report
Yadong Li, Haoze Sun, Mingan Lin +23
The salient multimodal capabilities and interactive experience of GPT-4o highlight its critical role in practical applications, yet it lacks a high-performing open-source counterpa…
Baichuan Alignment Technical Report
Mingan Lin, Fan Yang, Yanjun Shen +21
We introduce Baichuan Alignment, a detailed analysis of the alignment techniques employed in the Baichuan series of models. This represents the industry's first comprehensive accou…