4 papers
TeachBench: A Syllabus-Grounded Framework for Evaluating Teaching Ability in Large Language Models
Zheng Li, Siyao Song, Jingyuan Ma +4
Large language models (LLMs) show promise as teaching assistants, yet their teaching capability remains insufficiently evaluated. Existing benchmarks mainly focus on problem-solvin…
SCoRE: Benchmarking Long-Chain Reasoning in Commonsense Scenarios
Weidong Zhan, Yue Wang, Nan Hu +12
Currently, long-chain reasoning remains a key challenge for large language models (LLMs) because natural texts lack sufficient explicit reasoning data. However, existing benchmarks…
Be a Multitude to Itself: A Prompt Evolution Framework for Red Teaming
Rui Li, Peiyi Wang, Jingyuan Ma +3
Large Language Models (LLMs) have gained increasing attention for their remarkable capacity, alongside concerns about safety arising from their potential to produce harmful content…
Plug-and-Play Training Framework for Preference Optimization
Jingyuan Ma, Rui Li, Zheng Li +2
Recently, preference optimization methods such as DPO have significantly enhanced large language models (LLMs) in wide tasks including dialogue and question-answering. However, cur…