168 citations · 644 across the 36 of their papers we have counts for
Showing 2026Show all
2 papers · 1 filter
cs.HC2026
Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations
Preethi Seshadri, Samuel Cahyawijaya, Ayomide Odumakinde +2
Agentic benchmarks increasingly rely on LLM-simulated users to scalably evaluate agent performance, yet the robustness, validity, and fairness of this approach remain unexamined. T…
cs.LG2026
Entropy-Aligned Decoding of LMs for Better Writing and Reasoning
Kareem Ahmed, Sameer Singh
Language models (LMs) are trained on billions of tokens in an attempt to recover the true language distribution. Still, vanilla random sampling from LMs yields low quality generati…