10 papers
Agentic Abstention: Do Agents Know When to Stop Instead of Act?
Han Luo, Bingbing Wen, Lucy Lu Wang
LLM agents are expected to act over multiple turns, using search, browsing interfaces, and terminal tools to complete user goals. Yet not every goal is well specified or achievable…
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation
Katelyn Xiaoying Mei, Yi-Li Hsu, Minjoon Choi +5
Human evaluation plays a critical role in assessing the quality of generated text. However, the reliability and reproducibility of these evaluations depend on transparent and well-…
STOP: Structured On-Policy Pruning of Long-Form Reasoning in Low-Data Regimes
Chenjun Xu, Zhennan Zhou, Zhan Su +3
Long chain-of-thought (Long CoT) reasoning improves performance on multi-step problems, but it also induces overthinking: models often generate low-yield reasoning that increases i…
MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining
Bingbing Wen, Sirajul Salekin, Feiyang Kang +4
Domain reweighting can improve sample efficiency and downstream generalization, but data-mixture optimization for multimodal midtraining remains largely unexplored. Current multimo…
SusBench: An Online Benchmark for Evaluating Dark Pattern Susceptibility of Computer-Use Agents
Longjie Guo, Chenjie Yuan, Mingyuan Zhong +7
As LLM-based computer-use agents (CUAs) begin to autonomously interact with real-world interfaces, understanding their vulnerability to manipulative interface designs becomes incre…
Clarify or Answer: Reinforcement Learning for Agentic VQA with Context Under-specification
Zongwan Cao, Bingbing Wen, Lucy Lu Wang
Real-world visual question answering (VQA) is often context-dependent: an image-question pair may be under-specified, such that the correct answer depends on external information t…