4 papers
Xpertbench: Expert Level Tasks with Rubrics-Based Evaluation
Xue Liu, Xin Ma, Yuxin Ma +36
As Large Language Models (LLMs) exhibit plateauing performance on conventional benchmarks, a pivotal challenge persists: evaluating their proficiency in complex, open-ended tasks c…
Efficient Reasoning via Thought Compression for Language Segmentation
Qing Zhou, Shiyu Zhang, Yuyu Jia +4
Chain-of-thought (CoT) reasoning has significantly improved the performance of large multimodal models in language-guided segmentation, yet its prohibitive computational cost, stem…
Improving Instruct Models for Free: A Study on Partial Adaptation
Ozan İrsoy, Pengxiang Cheng, Jennifer L. Chen +3
Instruct models, obtained from various instruction tuning or post-training steps, are commonly deemed superior and more usable than their base counterpart. While the model gains in…
Evaluating the Retrieval Robustness of Large Language Models
Shuyang Cao, Karthik Radhakrishnan, David Rosenberg +4
Retrieval-augmented generation (RAG) generally enhances large language models' (LLMs) ability to solve knowledge-intensive tasks. But RAG may also lead to performance degradation d…