Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377
To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…
cs.CL2025
ReProCon: Scalable and Resource-Efficient Few-Shot Biomedical Named Entity Recognition
Jeongkyun Yoo, Nela Riddle, Andrew Hoblitzell
Named Entity Recognition (NER) in biomedical domains faces challenges due to data scarcity and imbalanced label distributions, especially with fine-grained entity types. We propose…
cs.CL2024
MetaCheckGPT -- A Multi-task Hallucination Detector Using LLM Uncertainty and Meta-models
Rahul Mehta, Andrew Hoblitzell, Jack O'Keefe +2
Hallucinations in large language models (LLMs) have recently become a significant problem. A recent effort in this direction is a shared task at Semeval 2024 Task 6, SHROOM, a Shar…