4 papers
INTRYGUE: Induction-Aware Entropy Gating for Reliable RAG Uncertainty Estimation
Alexandra Bazarova, Andrei Volodichev, Daria Kotova +1
While retrieval-augmented generation (RAG) significantly improves the factual reliability of LLMs, it does not eliminate hallucinations, so robust uncertainty quantification (UQ) r…
UNVEILING: What Makes Linguistics Olympiad Puzzles Tricky for LLMs?
Mukund Choudhary, KV Aditya Srivatsa, Gaurja Aeron +6
Large language models (LLMs) have demonstrated potential in reasoning tasks, but their performance on linguistics puzzles remains consistently poor. These puzzles, often derived fr…
What Makes Cryptic Crosswords Challenging for LLMs?
Abdelrahman Sadallah, Daria Kotova, Ekaterina Kochmar
Cryptic crosswords are puzzles that rely on general knowledge and the solver's ability to manipulate language on different levels, dealing with various types of wordplay. Previous…
Are LLMs Good Cryptic Crossword Solvers?
Abdelrahman Sadallah, Daria Kotova, Ekaterina Kochmar
Cryptic crosswords are puzzles that rely not only on general knowledge but also on the solver's ability to manipulate language on different levels and deal with various types of wo…