1 paper
Abha Jha, Akanksha Mahajan, Ashwath Vaithinathan Aravindan +3
Large Language Models (LLMs) often produce hallucinated or unverifiable content, undermining their reliability in factual domains. This work investigates Reinforcement Learning wit…