2 papers
cs.CL2026
LEDGER: A Long-Context Benchmark of Corporate Annual Reports for Grounded Financial Retrieval and Extraction
Charles Moslonka, Amaury de Vitry, Arthur Garnier +2
Finance reporting is a natural proving ground for large language models, and the very-long-context capabilities of recent models across all sizes make rigorous evaluation in this d…
cs.CL2026
Learned Hallucination Detection in Black-Box LLMs using Token-level Entropy Production Rate
Charles Moslonka, Hicham Randrianarivo, Arthur Garnier +1
Hallucinations in Large Language Model (LLM) outputs for Question Answering (QA) tasks can critically undermine their real-world reliability. This paper introduces a methodology fo…