1 paper
Sepehr Harfi, Ahmad Salimi, Dongming Shen +1
Most LLM benchmarks score how well a model responds to explicit requests. They leave unmeasured a different conversational ability: noticing and acting on needs the user has implie…