2 papers
cs.CL2024
VERITAS: A Unified Approach to Reliability Evaluation
Rajkumar Ramamurthy, Meghana Arakkal Rajeev, Oliver Molenschot +2
Large language models (LLMs) often fail to synthesize information from their context to generate an accurate response. This renders them unreliable in knowledge intensive settings…
cs.CL2024
Self-rationalization improves LLM as a fine-grained judge
Prapti Trivedi, Aditya Gulati, Oliver Molenschot +7
LLM-as-a-judge models have been used for evaluating both human and AI generated content, specifically by providing scores and rationales. Rationales, in addition to increasing tran…