5 papers · 1 filter
Understanding Why Language Models Hallucinate: Testing Reasoning Against Priors
Yangfan Hu, Xuhan Tong, Haoyue Bai +5
Large language models often produce hallucinated answers that violate prompt-level constraints. A key diagnostic question is whether these failures reflect missing knowledge, or wh…
Improving Task Diversity in Label Efficient Supervised Finetuning of LLMs
Abhinav Arabelly, Jagrut Nemade, Robert D Nowak +1
Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse domains, but developing high-performing models for specialized applications often requires sub…
Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench
Reuben Narad, Siddharth Suresh, Jiayi Chen +5
We present HumorBench, a benchmark designed to evaluate large language models' (LLMs) ability to reason about and explain sophisticated humor in cartoon captions. As reasoning mode…
GPT-4o as the Gold Standard: A Scalable and General Purpose Approach to Filter Language Model Pretraining Data
Jifan Zhang, Ziyue Luo, Jia Liu +2
Large language models require vast amounts of high-quality training data, but effective filtering of web-scale datasets remains a significant challenge. This paper demonstrates tha…
An Experimental Design Framework for Label-Efficient Supervised Finetuning of Large Language Models
Gantavya Bhatt, Yifang Chen, Arnav M. Das +9
Supervised finetuning (SFT) on instruction datasets has played a crucial role in achieving the remarkable zero-shot generalization capabilities observed in modern large language mo…