1 paper
Eftychia Makri, Nikolaos Nakis, Laura Sisson +4
Here we introduce the Olfactory Perception (OP) benchmark, designed to assess the capability of large language models (LLMs) to reason about smell. The benchmark contains 1,010 que…