5 papers
Fantastic Bugs and Where to Find Them in AI Benchmarks
Sang Truong, Yuheng Tu, Michael Hardy +8
Benchmarks are pivotal in driving AI progress, and invalid benchmark questions frequently undermine their reliability. Manually identifying and correcting errors among thousands of…
Detecting Corpus-Level Knowledge Inconsistencies in Wikipedia with Large Language Models
Sina J. Semnani, Jirayu Burapacheep, Arpandeep Khatua +3
Wikipedia is the largest open knowledge corpus, widely used worldwide and serving as a key resource for training large language models (LLMs) and retrieval-augmented generation (RA…
Your Classifier Can Be Secretly a Likelihood-Based OOD Detector
Jirayu Burapacheep, Yixuan Li
The ability to detect out-of-distribution (OOD) inputs is critical to guarantee the reliability of classification models deployed in an open environment. A fundamental challenge in…
ColorSwap: A Color and Word Order Dataset for Multimodal Evaluation
Jirayu Burapacheep, Ishan Gaur, Agam Bhatia +1
This paper introduces the ColorSwap dataset, designed to assess and improve the proficiency of multimodal models in matching objects with their colors. The dataset is comprised of…
ARGS: Alignment as Reward-Guided Search
Maxim Khanov, Jirayu Burapacheep, Yixuan Li
Aligning large language models with human objectives is paramount, yet common approaches including RLHF suffer from unstable and resource-intensive training. In response to this ch…