4 papers
Tendem: A Hybrid AI+Human Platform
Konstantin Chernyshev, Ekaterina Artemova, Viacheslav Zhukov +8
Tendem is a hybrid system where AI handles structured, repeatable work and Human Experts step in when the models fail or to verify results. Each result undergoes a comprehensive qu…
JEEM: Vision-Language Understanding in Four Arabic Dialects
Karima Kadaoui, Hanin Atwany, Hamdan Al-Ali +7
We introduce JEEM, a benchmark designed to evaluate Vision-Language Models (VLMs) on visual understanding across four Arabic-speaking countries: Jordan, The Emirates, Egypt, and Mo…
U-MATH: A University-Level Benchmark for Evaluating Mathematical Skills in LLMs
Konstantin Chernyshev, Vitaliy Polshkov, Ekaterina Artemova +4
The current evaluation of mathematical skills in LLMs is limited, as existing benchmarks are either relatively small, primarily focus on elementary and high-school problems, or lac…
Beemo: Benchmark of Expert-edited Machine-generated Outputs
Ekaterina Artemova, Jason Lucas, Saranya Venkatraman +4
The rapid proliferation of large language models (LLMs) has increased the volume of machine-generated texts (MGTs) and blurred text authorship in various domains. However, most exi…