collaborators

5 papers

cs.CL2026

ClinConsensus: A Physician-Calibrated Benchmark for Evaluating Clinical Rubric Coverage in Chinese Medical LLMs

Xiang Zheng, Han Li, Wenjie Luo +15

Open-ended medical LLM evaluation remains weakly grounded in physician-calibrated coverage of clinically relevant response criteria, especially in localized clinical settings. We i…

cs.AI2026

ClawTrace: Cost-Aware Tracing for LLM Agent Skill Distillation

Boqin Yuan, Yue Su, Renchu Song +2

Skill-distillation pipelines learn reusable rules from LLM agent trajectories, but they lack a key signal: how much each step costs. Without per-step cost, a pipeline cannot distin…

cs.CL2026

Capability Conditioned Scaffolding for Professional Human LLM Collaboration

Sen Yang, Yinglei Ma

Large language model personalization typically adapts outputs to user preferences and style but does not account for differences in user evaluation capacity across domains of exper…

cs.AI2025

Multi-LLM Collaborative Search for Complex Problem Solving

Sen Yang, Yafu Li, Wai Lam +1

Large language models (LLMs) often struggle with complex reasoning tasks due to their limitations in addressing the vast reasoning space and inherent ambiguities of natural languag…

cs.CL2025

From Drafts to Answers: Unlocking LLM Potential via Aggregation Fine-Tuning

Yafu Li, Zhilin Wang, Tingchen Fu +3

Scaling data and model size has been proven effective for boosting the performance of large language models. In addition to training-time scaling, recent studies have revealed that…