most citedFIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights

1 citations · 1 across the 6 of their papers we have counts for

collaborators

9 papers

cs.AI20261 cited

FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights

Zhen Wang, Fan Bai, Zhongyan Luo +9

Autonomous agents powered by large language models (LLMs) promise to accelerate scientific discovery end-to-end, but rigorously evaluating their capacity for verifiable discovery r…

cs.LG2026

BAGEN: Are LLM Agents Budget-Aware?

Yuxiang Lin, Zihan Wang, Mengyang Liu +9

While agents are increasingly spending more resources, today agent cost is mostly measured only after execution. A Budget-Aware Agent (BAGEN) should treat budget as an active contr…

cs.CL2026

Knowing but Not Showing: LLMs Recognize Ambiguity but Rarely Ask Clarifying Questions

Jinyan Su, Claire Cardie

User queries are often underspecified and may admit multiple valid interpretations. Rather than silently making assumptions about the user's intent, a helpful assistant should surf…

cs.CL2026

Clarification Is Not Enough: Post-Clarification Answering Remains the Bottleneck in Multi-Turn QA

Jinyan Su, Jennifer Healey

Pluralistic alignment requires systems to adapt to diverse user values, communication styles, and contextual assumptions. We believe that a foundational prerequisite for such align…

cs.LG2026

The Illusion of Specialization: Unveiling the Domain-Invariant "Standing Committee" in Mixture-of-Experts Models

Yan Wang, Yitao Xu, Nanhan Shen +3

Mixture of Experts models are widely assumed to achieve domain specialization through sparse routing. In this work, we question this assumption by introducing COMMITTEEAUDIT, a pos…

cs.CL2026

CLIPer: Tailoring Diverse User Preference via Classifier-Guided Inference-Time Personalization

Jinyan Su, Jinpeng Zhou, Claire Cardie +1

Personalized LLMs can significantly enhance user experiences by tailoring responses to preferences such as helpfulness, conciseness, and humor. However, fine-tuning models to addre…