1 citations · 1 across the 6 of their papers we have counts for
9 papers
FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights
Zhen Wang, Fan Bai, Zhongyan Luo +9
Autonomous agents powered by large language models (LLMs) promise to accelerate scientific discovery end-to-end, but rigorously evaluating their capacity for verifiable discovery r…
BAGEN: Are LLM Agents Budget-Aware?
Yuxiang Lin, Zihan Wang, Mengyang Liu +9
While agents are increasingly spending more resources, today agent cost is mostly measured only after execution. A Budget-Aware Agent (BAGEN) should treat budget as an active contr…
Knowing but Not Showing: LLMs Recognize Ambiguity but Rarely Ask Clarifying Questions
Jinyan Su, Claire Cardie
User queries are often underspecified and may admit multiple valid interpretations. Rather than silently making assumptions about the user's intent, a helpful assistant should surf…
Clarification Is Not Enough: Post-Clarification Answering Remains the Bottleneck in Multi-Turn QA
Jinyan Su, Jennifer Healey
Pluralistic alignment requires systems to adapt to diverse user values, communication styles, and contextual assumptions. We believe that a foundational prerequisite for such align…
The Illusion of Specialization: Unveiling the Domain-Invariant "Standing Committee" in Mixture-of-Experts Models
Yan Wang, Yitao Xu, Nanhan Shen +3
Mixture of Experts models are widely assumed to achieve domain specialization through sparse routing. In this work, we question this assumption by introducing COMMITTEEAUDIT, a pos…
CLIPer: Tailoring Diverse User Preference via Classifier-Guided Inference-Time Personalization
Jinyan Su, Jinpeng Zhou, Claire Cardie +1
Personalized LLMs can significantly enhance user experiences by tailoring responses to preferences such as helpfulness, conciseness, and humor. However, fine-tuning models to addre…