1 paper
David Castillo-Bolado, Joseph Davidson, Finlay Gray +1
We introduce a dynamic benchmarking system for conversational agents that evaluates their performance through a single, simulated, and lengthy user↔agent interactio…