6 papers
Future Confidence Distillation in Large Language Models
Sahil Kale
Reliable confidence estimation is essential for deploying large language models (LLMs) in confidence-aware systems, where downstream decisions such as retrieval, tool use, and adap…
Lie to Me: Knowledge Graphs for Robust Hallucination Self-Detection in LLMs
Sahil Kale, Antonio Luca Alfeo
Hallucinations, the generation of apparently convincing yet false statements, remain a major barrier to the safe deployment of LLMs. Building on the strong performance of self-dete…
Mirage of Mastery: Memorization Tricks LLMs into Artificially Inflated Self-Knowledge
Sahil Kale
When artificial intelligence mistakes memorization for intelligence, it creates a dangerous mirage of reasoning. Existing studies treat memorization and self-knowledge deficits in…
Look It Up: Analysing Internal Web Search Capabilities of Modern LLMs
Sahil Kale
Modern large language models integrate web search to provide real-time answers, yet it remains unclear whether they are efficiently calibrated to use search when it is actually nee…
TeXpert: A Multi-Level Benchmark for Evaluating LaTeX Code Generation by LLMs
Sahil Kale, Vijaykant Nadadur
LaTeX's precision and flexibility in typesetting have made it the gold standard for the preparation of scientific documentation. Large Language Models (LLMs) present a promising op…
Line of Duty: Evaluating LLM Self-Knowledge via Consistency in Feasibility Boundaries
Sahil Kale, Vijaykant Nadadur
As LLMs grow more powerful, their most profound achievement may be recognising when to say "I don't know". Existing studies on LLM self-knowledge have been largely constrained by h…