2 papers
cs.CL2026
Beyond Majority Vote: Multi-Perspective Adjudication for Medical Hallucination Detection
Joe Cecil, Marjorie Freedman
Understanding the frequency of factual errors in chatbot-generated text and evaluating systems that detect these errors is critical for determining chatbot safety. Yet factual-erro…
cs.CL2023
Remember what you did so you know what to do next
Manuel R. Ciosici, Alex Hedges, Yash Kankanampati +3
We explore using a moderately sized large language model (GPT-J 6B parameters) to create a plan for a simulated robot to achieve 30 classes of goals in ScienceWorld, a text game si…