Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Agentic Abstention: Do Agents Know When to Stop Instead of Act?
Han Luo, Bingbing Wen, Lucy Lu Wang
LLM agents are expected to act over multiple turns, using search, browsing interfaces, and terminal tools to complete user goals. Yet not every goal is well specified or achievable…
cs.AI2025
Do Language Models Mirror Human Confidence? Exploring Psychological Insights to Address Overconfidence in LLMs
Chenjun Xu, Bingbing Wen, Bin Han +3
Psychology research has shown that humans are poor at estimating their performance on tasks, tending towards underconfidence on easy tasks and overconfidence on difficult tasks. We…
cs.AI2025
MMMG: a Comprehensive and Reliable Evaluation Suite for Multitask Multimodal Generation
Jihan Yao, Yushi Hu, Yujie Yi +9
Automatically evaluating multimodal generation presents a significant challenge, as automated metrics often struggle to align reliably with human evaluation, especially for complex…