4 papers
Deep Research Comparator: A Platform For Fine-grained Human Annotations of Deep Research Agents
Prahaladh Chandrahasan, Jiahe Jin, Zhihan Zhang +9
Effectively evaluating deep research agents that autonomously search the web, analyze information, and generate reports remains a major challenge, particularly when it comes to ass…
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning
Tevin Wang, Chenyan Xiong
Rule-based rewards offer a promising strategy for improving reinforcement learning from human feedback (RLHF), but current approaches often rely on manual rule engineering. We pres…
Understand User Opinions of Large Language Models via LLM-Powered In-the-Moment User Experience Interviews
Mengqiao Liu, Tevin Wang, Cassandra A. Cohen +2
Which large language model (LLM) is better? Every evaluation tells a story, but what do users really think about current LLMs? This paper presents CLUE, an LLM-powered interviewer…
Interpret and Control Dense Retrieval with Sparse Latent Features
Hao Kang, Tevin Wang, Chenyan Xiong
Dense embeddings deliver strong retrieval performance but often lack interpretability and controllability. This paper introduces a novel approach using sparse autoencoders (SAE) to…