7 citations · 7 across the 1 of their papers we have counts for
1 paper · 1 filter
Shunyu Yao, Noah Shinn, Pedram Razavi +1
Existing benchmarks do not test language agents on their interaction with human users or ability to follow domain-specific rules, both of which are vital for deploying them in real…