8 papers · 1 filter
Enoki: Efficient Multi-Level Hallucination Detection
Elisei Rykov, Timur Ionov, Nikolay Ivanov +5
Ensuring factuality remains a critical challenge for deploying LLMs in high-stakes settings. Existing hallucination detectors usually operate at a single level: claim-level methods…
MemToC: Benchmarking Memory-Tool Conflict Resolution in Large Language Models
Arseniy Varlamov, Rishat Zinnatullin, Elisei Rykov +2
Tool-augmented LLMs must arbitrate between two fallible sources when a tool return conflicts with their parametric memory, yet existing evaluations measure source preference withou…
Learning Selective LLM Autonomy from Copilot Feedback in Enterprise Customer Support Workflows
Nikita Borovkov, Elisei Rykov, Olga Tsymboi +4
We present a deployed system that automates end-to-end customer support workflows inside an enterprise Business Process Management (BPM) platform. The approach is scalable in produ…
Multimodal Evaluation of Russian-language Architectures
Artem Chervyakov, Ulyana Isaeva, Anton Emelyanov +15
Multimodal large language models (MLLMs) are currently at the center of research attention, showing rapid progress in scale and capabilities, yet their intelligence, limitations, a…
When Models Lie, We Learn: Multilingual Span-Level Hallucination Detection with PsiloQA
Elisei Rykov, Kseniia Petrushina, Maksim Savkin +6
Hallucination detection remains a fundamental challenge for the safe and reliable deployment of large language models (LLMs), especially in applications requiring factual accuracy.…
SmurfCat at PAN 2024 TextDetox: Alignment of Multilingual Transformers for Text Detoxification
Elisei Rykov, Konstantin Zaytsev, Ivan Anisimov +1
This paper presents a solution for the Multilingual Text Detoxification task in the PAN-2024 competition of the SmurfCat team. Using data augmentation through machine translation a…