From the 1 of 69 linked papers with an AI index.
6 papers · 1 filter
EvoMem: Memory-Augmented Evolution for Code Optimization
Viktor Volkov, Valentin Khrulkov, Andrey V. Galichin +8
Successful mutation strategies in evolutionary code search may contain reusable knowledge that is useful beyond a single run, and in some cases may transfer across related tasks an…
Search, Fail, Recover: A Training Framework for Correction-Aware Reasoning
Dmitry Beresnev, Vladimir Makharev, Roman Khalikov +2
Many reasoning tasks are not well described by a single left-to-right chain: a solver may need to pursue a plausible branch, observe delayed failure, and return to the latest prefi…
TheoremBench: Evaluating LLMs on Theorem Proving in Formal Mathematics
QuocViet Pham, Elvir Karimov, Andrey Galichin +1
LLMs have recently achieved strong results on formal proving benchmarks. However, existing evaluations remain heavily concentrated on competition-style problems and often fail to c…
Harnessing non-adversarial robustness in large language models
Qinghua Zhou, Ellina Aleshina, Andrey Lovyagin +6
The work presents an approach for addressing the challenge of robustness in Large Language Models (LLMs) to alterations and potential errors caused by semantically similar but text…
Multi-Agent GraphRAG: A Text-to-Cypher Framework for Labeled Property Graphs
Anton Gusarov, Anastasia Volkova, Valentin Khrulkov +3
While Retrieval-Augmented Generation (RAG) methods commonly draw information from unstructured documents, the emerging paradigm of GraphRAG aims to leverage structured data such as…
How to Evaluate Medical AI
Ilia Kopanichuk, Petr Anokhin, Vladimir Shaposhnikov +5
The integration of artificial intelligence (AI) into medical diagnostic workflows requires robust and consistent evaluation methods to ensure reliability, clinical relevance, and t…