4 papers
The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations
Emma Yanyang Kong, JJ Tan, Ishan Gupta +8
LLM-as-a-Judge, which leverages a large language model to evaluate natural language generated by another AI application or model, has become a standard, scalable approach for accel…
MicroVerse: An Instrument for Measuring Self-Authored Identity Drift in Long-Horizon Multi-Agent Language-Model Simulations
Sky Ng, Brihi Joshi, Ishan Gupta +47
Long-horizon, multi-agent language model (LM) simulations are widely proposed for studying social behavior, yet instruments to measure whether persona-conditioned agents maintain i…
PersonaEval: Persona-Based User Simulation for Evaluating Interactive Applications
Yifan Simon Liu, Qianfeng Wen, Yilan Fan +40
Real user studies are important for understanding how people interact with systems under test or already deployed. In practice, however, they are often costly, time-consuming, and…
LLM Reasoning for Subjective Tasks: Failure Modes, Mitigation, and Dynamic Reasoning Routing
Juncheng Dong, Ding Tong, Ishan Gupta +1
Recommendation systems thrive on personalization, where ''correctness'' is rarely a binary truth but a matter of subjective human preference. As Large Language Models (LLMs) are de…