3 papers
cs.AI2026
SIRIN: A Unified Toolkit for Detecting Contextual Hallucinations in Retrieval-Augmented and Memory-Grounded LLM Systems
Julia Belikova, Rauf Parchiev, Mikhail Filimonov +3
SIRIN (Semantic Inconsistency Recognition and Inspection Nexus) is a unified toolkit and interactive web UI for detecting contextual hallucinations (fluent, plausible responses uns…
cs.AI2026
Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluation
Julia Belikova, Rauf Parchiev, Evgeny Egorov +4
Procedural memory is increasingly used to improve LLM agents on recurring workplace tasks, yet its ability to produce reusable skills remains poorly understood. We introduce AFTER,…
cs.CL2025
Data-efficient Meta-models for Evaluation of Context-based Questions and Answers in LLMs
Julia Belikova, Konstantin Polev, Rauf Parchiev +1
Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) systems are increasingly deployed in industry applications, yet their reliability remains hampered by challeng…