6 citations · 6 across the 1 of their papers we have counts for
1 paper
Dingli Yu, Simran Kaur, Arushi Gupta +3
With LLMs shifting their role from statistical modeling of language to serving as general-purpose AI agents, how should LLM evaluations change? Arguably, a key ability of an AI age…