Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
InfoQuest: Evaluating Multi-Turn Dialogue Agents for Open-Ended Conversations with Hidden Context
Bryan L. M. de Oliveira, Luana G. B. Martins, Bruno Brandão +1
Large language models excel at following explicit instructions, but they often struggle with ambiguous or incomplete user requests, defaulting to verbose, generic responses instead…
cs.CL2025
Uncertainty-Aware Step-wise Verification with Generative Reward Models
Zihuiwen Ye, Luckeciano Carvalho Melo, Younesse Kaddar +3
Complex multi-step reasoning tasks, such as solving mathematical problems, remain challenging for large language models (LLMs). While outcome supervision is commonly used, process…