Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
ClinConsensus: A Physician-Calibrated Benchmark for Evaluating Clinical Rubric Coverage in Chinese Medical LLMs
Xiang Zheng, Han Li, Wenjie Luo +15
Open-ended medical LLM evaluation remains weakly grounded in physician-calibrated coverage of clinically relevant response criteria, especially in localized clinical settings. We i…
cs.CL2026
Capability Conditioned Scaffolding for Professional Human LLM Collaboration
Sen Yang, Yinglei Ma
Large language model personalization typically adapts outputs to user preferences and style but does not account for differences in user evaluation capacity across domains of exper…
cs.CL2025
From Drafts to Answers: Unlocking LLM Potential via Aggregation Fine-Tuning
Yafu Li, Zhilin Wang, Tingchen Fu +3
Scaling data and model size has been proven effective for boosting the performance of large language models. In addition to training-time scaling, recent studies have revealed that…