Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
SOPBench: Evaluating Language Agents at Following Standard Operating Procedures and Constraints
Zekun Li, Shinda Huang, Jiangtian Wang +8
As language agents increasingly automate critical tasks, their ability to follow domain-specific standard operating procedures (SOPs), policies, and constraints when taking actions…
cs.CL2025
Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data
Xinyi Wang, Antonis Antoniades, Yanai Elazar +4
The impressive capabilities of large language models (LLMs) have sparked debate over whether these models genuinely generalize to unseen tasks or predominantly rely on memorizing v…
cs.CL2025
DebUnc: Improving Large Language Model Agent Communication With Uncertainty Metrics
Luke Yoffe, Alfonso Amayuelas, William Yang Wang
Multi-agent debates have been introduced to improve the accuracy of Large Language Models (LLMs) by having multiple agents discuss solutions to a problem over several rounds of deb…