Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Learning an Efficient Multi-Turn Dialogue Evaluator from Multiple LLM Judges
Yuqi Tang, Kehua Feng, Yunfeng Wang +6
Evaluating the conversational abilities of large language models (LLMs) remains a challenging task. Current mainstream approaches primarily rely on the "LLM-as-a-judge" paradigm, w…
cs.CL2024
AutoAct: Automatic Agent Learning from Scratch for QA via Self-Planning
Shuofei Qiao, Ningyu Zhang, Runnan Fang +5
Language agents have achieved considerable performance on various complex question-answering tasks by planning with external tools. Despite the incessant exploration in this field,…