5 papers
Misalignment Bounty: Crowdsourcing AI Agent Misbehavior
Rustem Turtayev, Natalia Fedorova, Oleg Serikov +3
Advanced AI systems sometimes act in ways that differ from human intent. To gather clear, reproducible examples, we ran the Misalignment Bounty: a crowdsourced project that collect…
Preliminary Ranking of WMT25 General Machine Translation Systems
Tom Kocmi, Eleftherios Avramidis, Rachel Bawden +25
We present the preliminary rankings of machine translation (MT) systems submitted to the WMT25 General Machine Translation Shared Task, as determined by automatic evaluation metric…
Voices of Freelance Professional Writers on AI: Limitations, Expectations, and Fears
Anastasiia Ivanova, Natalia Fedorova, Sergei Tilga +1
The rapid development of AI-driven tools, particularly large language models (LLMs), is reshaping professional writing. Still, key aspects of their adoption such as languages suppo…
JEEM: Vision-Language Understanding in Four Arabic Dialects
Karima Kadaoui, Hanin Atwany, Hamdan Al-Ali +7
We introduce JEEM, a benchmark designed to evaluate Vision-Language Models (VLMs) on visual understanding across four Arabic-speaking countries: Jordan, The Emirates, Egypt, and Mo…
Hands-On Tutorial: Labeling with LLM and Human-in-the-Loop
Ekaterina Artemova, Akim Tsvigun, Dominik Schlechtweg +4
Training and deploying machine learning models relies on a large amount of human-annotated data. As human labeling becomes increasingly expensive and time-consuming, recent researc…