2 papers
cs.CL2026
ConflictScore: Identifying and Measuring How Language Models Handle Conflicting Evidence
Siyi Liu, Aaron Halfaker, Dan Roth +1
Existing metrics for factuality and faithfulness evaluate whether an answer is supported or contradicted by its grounding documents, but they fail to capture when both supporting a…
cs.IR2025
Multi-Field Adaptive Retrieval
Millicent Li, Tongfei Chen, Benjamin Van Durme +1
Document retrieval for tasks such as search and retrieval-augmented generation typically involves datasets that are unstructured: free-form text without explicit internal structure…