2 papers
cs.AI2026
From Prompting to Behavioral Alignment: Personalized LLM Judges for Recommendation Evaluation
Alireza S. Ziabari, Kat Ellis, Colleen Chan +1
Traditional offline recommendation evaluation relies heavily on complex, manually maintained feature pipelines that are difficult to scale. While Large Language Models (LLMs) offer…
cs.AI2026
LLM Reasoning for Subjective Tasks: Failure Modes, Mitigation, and Dynamic Reasoning Routing
Juncheng Dong, Ding Tong, Ishan Gupta +1
Recommendation systems thrive on personalization, where ''correctness'' is rarely a binary truth but a matter of subjective human preference. As Large Language Models (LLMs) are de…