activity
20242026
collaborators

7 papers

cs.AI2026

H+ Embedding: Harmonizing Global and Token-Level Retrieval with Context-Dependent Phrases

Shusen Zhang, Junyi Hu, Ye Feng +6

Terminology-intensive retrieval, especially in medical settings, depends on preserving multi-word entities, abbreviations, numerical constraints, and compositional concepts. Howeve…

cs.CL2026

Med-R: Enhancing Medical Retrieval-Augmented Reasoning of LLMs via Progressive Reinforcement Learning

Keer Lu, Zheng Liang, Youquan Li +8

In medical scenarios, effectively retrieving external knowledge and leveraging it for rigorous logical reasoning is of significant importance. Despite their potential, existing wor…

cs.CL2025

Med-R: Crafting Trustworthy LLM Physicians via Retrieval and Reasoning of Evidence-Based Medicine

Keer Lu, Zheng Liang, Da Pan +6

Large Language Models (LLMs) have exhibited remarkable capabilities in clinical scenarios. Despite their potential, existing works face challenges when applying LLMs to medical set…

cs.LG2025

Baichuan-M2: Scaling Medical Capability with Large Verifier System

M2 Team, Chengfeng Dou, Chong Liu +31

As large language models (LLMs) advance in conversational and reasoning capabilities, their practical application in healthcare has become a critical research focus. However, there…

cs.CL2025

VersaTune: An Efficient Data Composition Framework for Training Multi-Capability LLMs

Keer Lu, Keshi Zhao, Zhuoran Zhang +8

As demonstrated by the proprietary Large Language Models (LLMs) such as GPT and Claude series, LLMs have the potential to achieve remarkable proficiency across a wide range of doma…

cs.CL2025

Baichuan-Omni-1.5 Technical Report

Yadong Li, Jun Liu, Tao Zhang +89

We introduce Baichuan-Omni-1.5, an omni-modal model that not only has omni-modal understanding capabilities but also provides end-to-end audio generation capabilities. To achieve f…