Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Benchmarking Deflection and Hallucination in Large Vision-Language Models
Nicholas Moratelli, Christopher Davis, Leonardo F. R. Ribeiro +2
Large Vision-Language Models (LVLMs) increasingly rely on retrieval to answer knowledge-intensive multimodal questions. Existing benchmarks overlook conflicts between visual and te…
cs.CL2025
GaRAGe: A Benchmark with Grounding Annotations for RAG Evaluation
Ionut-Teodor Sorodoc, Leonardo F. R. Ribeiro, Rexhina Blloshmi +2
We present GaRAGe, a large RAG benchmark with human-curated long-form answers and annotations of each grounding passage, allowing a fine-grained evaluation of whether LLMs can iden…
cs.CL2025
Prompting open-source and commercial language models for grammatical error correction of English learner text
Christopher Davis, Andrew Caines, Ãistein Andersen +6
Thanks to recent advances in generative AI, we are able to prompt large language models (LLMs) to produce texts which are fluent and grammatical. In addition, it has been shown tha…