collaborators

6 papers

cs.CL2026

Deep Research, Shallow Evaluation: A Case Study in Meta-Evaluation for Long-Form QA Benchmarks

Jena D. Hwang, Varsha Kishore, Amanpreet Singh +9

Recent advances have made long-form report-generating systems widely available. This has prompted evaluation frameworks that use LLM-as-judge protocols and claim verification, alon…

cs.HC2026

Understanding Usage and Engagement in AI-Powered Scientific Research Tools: The Asta Interaction Dataset

Dany Haddad, Dan Bareket, Joseph Chee Chang +19

AI-powered scientific research tools are rapidly being integrated into research workflows, yet the field lacks a clear lens into how researchers use these systems in real-world set…

cs.CL2026

SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks

Yilun Zhao, Kaiyan Zhang, Tiansheng Hu +15

We present SciArena, an open and collaborative platform for evaluating foundation models on scientific literature-grounded tasks. Unlike traditional benchmarks for scientific liter…

cs.HC2025

Facets, Taxonomies, and Syntheses: Navigating Structured Representations in LLM-Assisted Literature Review

Raymond Fok, Joseph Chee Chang, Marissa Radensky +4

Comprehensive literature review requires synthesizing vast amounts of research -- a labor intensive and cognitively demanding process. Most prior work focuses either on helping res…

cs.HC2025

Social-RAG: Retrieving from Group Interactions to Socially Ground AI Generation

Ruotong Wang, Xinyi Zhou, Lin Qiu +3

AI agents are increasingly tasked with making proactive suggestions in online spaces where groups collaborate, yet risk being unhelpful or even annoying if they fail to match group…

cs.CL2025

Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference

Mingqi Gao, Yixin Liu, Xinyu Hu +3

Evaluating and ranking the capabilities of different LLMs is crucial for understanding their performance and alignment with human preferences. Due to the high cost and time-consumi…