11 papers
Am I More Pointwise or Pairwise? Revealing Position Bias in Rubric-Based LLM-as-a-Judge
Yuzheng Xu, Tosho Hirasawa, Tadashi Kozuno +1
Large language models are widely employed as evaluators, a paradigm commonly referred to as LLM-as-a-judge. Prior research has predominantly examined point-wise or pair-wise evalua…
Adaptable Method for Crystal Design across Diverse Constraints and Objectives with Pretrained Property Predictors
Akihiro Fujii, Yoshitaka Ushiku, Koji Shimizu +2
Advanced crystal design can accelerate materials discovery across applications from photovoltaics to spintronics. Practical design must satisfy multiple properties and physical con…
HalDec-Bench: Benchmarking Hallucination Detector in Image Captioning
Kuniaki Saito, Risa Shinoda, Shohei Tanaka +3
Hallucination detection in captions (HalDec) assesses a vision-language model's ability to correctly align image content with text by identifying errors in captions that misreprese…
HalDec-Bench: Benchmarking Hallucination Detector in Image Captioning
Kuniaki Saito, Risa Shinoda, Shohei Tanaka +3
Hallucination detection in captions (HalDec) assesses a vision-language model's ability to correctly align image content with text by identifying errors in captions that misreprese…
MaterialFigBENCH: benchmark dataset with figures for evaluating college-level materials science problem-solving abilities of multimodal large language models
Michiko Yoshitake, Yuta Suzuki, Ryo Igarashi +2
We present MaterialFigBench, a benchmark dataset designed to evaluate the ability of multimodal large language models (LLMs) to solve university-level materials science problems th…
WarrantScore: Modeling Warrants between Claims and Evidence for Substantiation Evaluation in Peer Reviews
Kiyotada Mori, Shohei Tanaka, Tosho Hirasawa +3
The scientific peer-review process is facing a shortage of human resources due to the rapid growth in the number of submitted papers. The use of language models to reduce the human…