4 papers
SIRIN: A Unified Toolkit for Detecting Contextual Hallucinations in Retrieval-Augmented and Memory-Grounded LLM Systems
Julia Belikova, Rauf Parchiev, Mikhail Filimonov +3
SIRIN (Semantic Inconsistency Recognition and Inspection Nexus) is a unified toolkit and interactive web UI for detecting contextual hallucinations (fluent, plausible responses uns…
Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluation
Julia Belikova, Rauf Parchiev, Evgeny Egorov +4
Procedural memory is increasingly used to improve LLM agents on recurring workplace tasks, yet its ability to produce reusable skills remains poorly understood. We introduce AFTER,…
Hallucination Detection in LLMs with Topological Divergence on Attention Graphs
Alexandra Bazarova, Andrei Volodichev, Aleksandr Yugay +10
Hallucination, i.e., generating factually incorrect content, remains a critical challenge for large language models (LLMs). We introduce TOHA, a TOpology-based HAllucination detect…
Data-efficient Meta-models for Evaluation of Context-based Questions and Answers in LLMs
Julia Belikova, Konstantin Polev, Rauf Parchiev +1
Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) systems are increasingly deployed in industry applications, yet their reliability remains hampered by challeng…