4 papers · 1 filter
Am I More Pointwise or Pairwise? Revealing Position Bias in Rubric-Based LLM-as-a-Judge
Yuzheng Xu, Tosho Hirasawa, Tadashi Kozuno +1
Large language models are widely employed as evaluators, a paradigm commonly referred to as LLM-as-a-judge. Prior research has predominantly examined point-wise or pair-wise evalua…
MaterialFigBENCH: benchmark dataset with figures for evaluating college-level materials science problem-solving abilities of multimodal large language models
Michiko Yoshitake, Yuta Suzuki, Ryo Igarashi +2
We present MaterialFigBench, a benchmark dataset designed to evaluate the ability of multimodal large language models (LLMs) to solve university-level materials science problems th…
WarrantScore: Modeling Warrants between Claims and Evidence for Substantiation Evaluation in Peer Reviews
Kiyotada Mori, Shohei Tanaka, Tosho Hirasawa +3
The scientific peer-review process is facing a shortage of human resources due to the rapid growth in the number of submitted papers. The use of language models to reduce the human…
Where is the answer? Investigating Positional Bias in Language Model Knowledge Extraction
Kuniaki Saito, Kihyuk Sohn, Chen-Yu Lee +1
Large language models require updates to remain up-to-date or adapt to new domains by fine-tuning them with new documents. One key is memorizing the latest information in a way tha…