1 citations · 1 across the 2 of their papers we have counts for
3 papers
TutorBench: A Benchmark To Assess Tutoring Capabilities Of Large Language Models
Rakshith S Srinivasa, Zora Che, Chen Bo Calvin Zhang +11
As students increasingly adopt large language models (LLMs) as learning aids, it is crucial to build models that are adept at handling the nuances of tutoring: they need to identif…
MultiNRC: A Challenging and Native Multilingual Reasoning Evaluation Benchmark for LLMs
Alexander R. Fabbri, Diego Mares, Jorge Flores +5
Although recent Large Language Models (LLMs) have shown rapid improvement on reasoning benchmarks in English, the evaluation of such LLMs' multilingual reasoning capability across…
MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
Ved Sirdeshmukh, Kaustubh Deshpande, Johannes Mols +7
We present MultiChallenge, a pioneering benchmark evaluating large language models (LLMs) on conducting multi-turn conversations with human users, a crucial yet underexamined capab…