collaborators

5 papers

cs.AI2025

Deep Research Comparator: A Platform For Fine-grained Human Annotations of Deep Research Agents

Prahaladh Chandrahasan, Jiahe Jin, Zhihan Zhang +9

Effectively evaluating deep research agents that autonomously search the web, analyze information, and generate reports remains a major challenge, particularly when it comes to ass…

cs.LG2025

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning

Tevin Wang, Chenyan Xiong

Rule-based rewards offer a promising strategy for improving reinforcement learning from human feedback (RLHF), but current approaches often rely on manual rule engineering. We pres…

cs.CL2025

Understand User Opinions of Large Language Models via LLM-Powered In-the-Moment User Experience Interviews

Mengqiao Liu, Tevin Wang, Cassandra A. Cohen +2

Which large language model (LLM) is better? Every evaluation tells a story, but what do users really think about current LLMs? This paper presents CLUE, an LLM-powered interviewer…

cs.CL2024

RAGViz: Diagnose and Visualize Retrieval-Augmented Generation

Tevin Wang, Jingyuan He, Chenyan Xiong

Retrieval-augmented generation (RAG) combines knowledge from domain-specific sources into large language models to ground answer generation. Current RAG systems lack customizable v…

cs.IR2024

Interpret and Control Dense Retrieval with Sparse Latent Features

Hao Kang, Tevin Wang, Chenyan Xiong

Dense embeddings deliver strong retrieval performance but often lack interpretability and controllability. This paper introduces a novel approach using sparse autoencoders (SAE) to…