5 papers
Tendem: A Hybrid AI+Human Platform
Konstantin Chernyshev, Ekaterina Artemova, Viacheslav Zhukov +8
Tendem is a hybrid system where AI handles structured, repeatable work and Human Experts step in when the models fail or to verify results. Each result undergoes a comprehensive qu…
U-MATH: A University-Level Benchmark for Evaluating Mathematical Skills in LLMs
Konstantin Chernyshev, Vitaliy Polshkov, Ekaterina Artemova +4
The current evaluation of mathematical skills in LLMs is limited, as existing benchmarks are either relatively small, primarily focus on elementary and high-school problems, or lac…
JEEM: Vision-Language Understanding in Four Arabic Dialects
Karima Kadaoui, Hanin Atwany, Hamdan Al-Ali +7
We introduce JEEM, a benchmark designed to evaluate Vision-Language Models (VLMs) on visual understanding across four Arabic-speaking countries: Jordan, The Emirates, Egypt, and Mo…
Beemo: Benchmark of Expert-edited Machine-generated Outputs
Ekaterina Artemova, Jason Lucas, Saranya Venkatraman +4
The rapid proliferation of large language models (LLMs) has increased the volume of machine-generated texts (MGTs) and blurred text authorship in various domains. However, most exi…
Hands-On Tutorial: Labeling with LLM and Human-in-the-Loop
Ekaterina Artemova, Akim Tsvigun, Dominik Schlechtweg +4
Training and deploying machine learning models relies on a large amount of human-annotated data. As human labeling becomes increasingly expensive and time-consuming, recent researc…