4 papers · 1 filter
TOD-ProcBench: Benchmarking Complex Instruction-Following in Task-Oriented Dialogues
Sarik Ghazarian, Abhinav Gullapalli, Swair Shah +4
In real-world task-oriented dialogue (TOD) settings, agents are required to strictly adhere to complex instructions while conducting multi-turn conversations with customers. These…
Semantic Volume: Quantifying and Detecting both External and Internal Uncertainty in LLMs
Xiaomin Li, Zhou Yu, Ziji Zhang +4
Large language models (LLMs) have demonstrated remarkable performance across diverse tasks by encoding vast amounts of factual knowledge. However, they are still prone to hallucina…
When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs
Xiaomin Li, Zhou Yu, Zhiwei Zhang +5
Reasoning-enhanced large language models (RLLMs), whether explicitly trained for reasoning or prompted via chain-of-thought (CoT), have achieved state-of-the-art performance on man…
How and Where to Translate? The Impact of Translation Strategies in Cross-lingual LLM Prompting
Aman Gupta, Yingying Zhuang, Zhou Yu +2
Despite advances in the multilingual capabilities of Large Language Models (LLMs), their performance varies substantially across different languages and tasks. In multilingual retr…