activity
20242026
collaborators

6 papers

cs.AI2026

MedConsultBench: A Full-Cycle, Fine-Grained, Process-Aware Benchmark for Medical Consultation Agents

Chuhan Qiao, Jianghua Huang, Daxing Zhao +5

Current evaluations of medical consultation agents often prioritize outcome-oriented tasks, frequently overlooking the end-to-end process integrity and clinical safety essential fo…

cs.CL2025

CFBench: A Comprehensive Constraints-Following Benchmark for LLMs

Tao Zhang, Chenglin Zhu, Yanjun Shen +10

The adeptness of Large Language Models (LLMs) in comprehending and following natural language instructions is critical for their deployment in sophisticated real-world applications…

cs.CL2025

Baichuan 2: Open Large-scale Language Models

Aiyuan Yang, Bin Xiao, Bingning Wang +52

Large language models (LLMs) have demonstrated remarkable performance on a variety of natural language tasks based on just a few examples of natural language instructions, reducing…

cs.AI2024

Baichuan-Omni Technical Report

Yadong Li, Haoze Sun, Mingan Lin +23

The salient multimodal capabilities and interactive experience of GPT-4o highlight its critical role in practical applications, yet it lacks a high-performing open-source counterpa…

cs.LG2024

Baichuan Alignment Technical Report

Mingan Lin, Fan Yang, Yanjun Shen +21

We introduce Baichuan Alignment, a detailed analysis of the alignment techniques employed in the Baichuan series of models. This represents the industry's first comprehensive accou…

cs.CL2024

SysBench: Can Large Language Models Follow System Messages?

Yanzhao Qin, Tao Zhang, Yanjun Shen +8

Large Language Models (LLMs) have become instrumental across various applications, with the customization of these models to specific scenarios becoming increasingly critical. Syst…