3 citations · 3 across the 5 of their papers we have counts for
4 papers · 1 filter
IMPersona: Evaluating Individual Level LM Impersonation
Quan Shi, Carlos E. Jimenez, Stephen Dong +4
As language models achieve increasingly human-like capabilities in conversational text generation, a critical question emerges: to what extent can these systems simulate the charac…
Atom of Thoughts for Markov LLM Test-Time Scaling
Fengwei Teng, Quan Shi, Zhaoyang Yu +4
Large Language Models (LLMs) have achieved significant performance gains through test-time scaling methods. However, existing approaches often incur redundant computations due to t…
BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval
Hongjin Su, Howard Yen, Mengzhou Xia +12
Existing retrieval benchmarks primarily consist of information-seeking queries (e.g., aggregated questions from search engines) where keyword or semantic-based retrieval is usually…
Can Language Models Solve Olympiad Programming?
Quan Shi, Michael Tang, Karthik Narasimhan +1
Computing olympiads contain some of the most challenging problems for humans, requiring complex algorithmic reasoning, puzzle solving, in addition to generating efficient code. How…