1 citations · 1 across the 2 of their papers we have counts for
3 papers
cs.AI2025
Fantastic Bugs and Where to Find Them in AI Benchmarks
Sang Truong, Yuheng Tu, Michael Hardy +8
Benchmarks are pivotal in driving AI progress, and invalid benchmark questions frequently undermine their reliability. Manually identifying and correcting errors among thousands of…
cs.CL2025
Detecting Corpus-Level Knowledge Inconsistencies in Wikipedia with Large Language Models
Sina J. Semnani, Jirayu Burapacheep, Arpandeep Khatua +3
Wikipedia is the largest open knowledge corpus, widely used worldwide and serving as a key resource for training large language models (LLMs) and retrieval-augmented generation (RA…
cs.LG2024★ 1 cited
Your Classifier Can Be Secretly a Likelihood-Based OOD Detector
Jirayu Burapacheep, Yixuan Li
The ability to detect out-of-distribution (OOD) inputs is critical to guarantee the reliability of classification models deployed in an open environment. A fundamental challenge in…