3 papers
cs.CL2026
From Coarse to Fine: Benchmarking and Reward Modeling for Writing-Centric Generation Tasks
Qingyu Ren, Tianjun Pan, Xingzhou Chen +1
Large language models have achieved remarkable progress in text generation but still struggle with generative writing tasks. In terms of evaluation, existing benchmarks evaluate wr…
cs.CL2026
OPSDL: On-Policy Self-Distillation for Long-Context Language Models
Xinsen Zhang, Zhenkai Ding, Tianjun Pan +4
Extending the effective context length of large language models (LLMs) remains a central challenge for real-world applications. While recent post-training methods have made progres…
cs.AI2026
RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following
Tianjun Pan, Xuan Lin, Wenyan Yang +7
Rubric-based evaluation has become a prevailing paradigm for evaluating instruction following in large language models (LLMs). Despite its widespread use, the reliability of these…