4 papers
Learning Selective LLM Autonomy from Copilot Feedback in Enterprise Customer Support Workflows
Nikita Borovkov, Elisei Rykov, Olga Tsymboi +4
We present a deployed system that automates end-to-end customer support workflows inside an enterprise Business Process Management (BPM) platform. The approach is scalable in produ…
Multimodal Evaluation of Russian-language Architectures
Artem Chervyakov, Ulyana Isaeva, Anton Emelyanov +15
Multimodal large language models (MLLMs) are currently at the center of research attention, showing rapid progress in scale and capabilities, yet their intelligence, limitations, a…
When Models Lie, We Learn: Multilingual Span-Level Hallucination Detection with PsiloQA
Elisei Rykov, Kseniia Petrushina, Maksim Savkin +6
Hallucination detection remains a fundamental challenge for the safe and reliable deployment of large language models (LLMs), especially in applications requiring factual accuracy.…
Don't Fight Hallucinations, Use Them: Estimating Image Realism using NLI over Atomic Facts
Elisei Rykov, Kseniia Petrushina, Kseniia Titova +2
Quantifying the realism of images remains a challenging problem in the field of artificial intelligence. For example, an image of Albert Einstein holding a smartphone violates comm…