Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
SciCustom: A Framework for Custom Evaluation of Scientific Capabilities in Large Language Models
Yiyang Gu, Junwei Yang, Junyu Luo +15
Large language models (LLMs) are increasingly applied to scientific research, yet existing evaluations often fail to reflect the fine-grained capabilities required in practice. Mos…
cs.CL2026
CoPA: Benchmarking Personalized Question Answering with Data-Informed Cognitive Factors
Hang Su, Zequn Liu, Chen Hu +3
While LLMs have demonstrated remarkable potential in Question Answering (QA), evaluating personalization remains a critical bottleneck. Existing paradigms predominantly rely on lex…