1 paper
Andy Dai, Zexue He, Zhenyu Zhang +2
Current benchmarks for language models primarily evaluate execution on fully specified tasks. However, real user tasks are often ambiguous. Users arrive with incomplete, explorator…