4 papers
Self-Preference Bias in LLM-as-a-Judge
Koki Wataoka, Tsubasa Takahashi, Ryokan Ri
Automated evaluation leveraging large language models (LLMs), commonly referred to as LLM evaluators or LLM-as-a-judge, has been widely used in measuring the performance of dialogu…
MergePrint: Merge-Resistant Fingerprints for Robust Black-box Ownership Verification of Large Language Models
Shojiro Yamabe, Futa Waseda, Tsubasa Takahashi +1
Protecting the intellectual property of Large Language Models (LLMs) has become increasingly critical due to the high cost of training. Model merging, which integrates multiple exp…
One-D-Piece: Image Tokenizer Meets Quality-Controllable Compression
Keita Miwa, Kento Sasaki, Hidehisa Arai +2
Current image tokenization methods require a large number of tokens to capture the information contained within images. Although the amount of information varies across images, mos…
ACT-Bench: Towards Action Controllable World Models for Autonomous Driving
Hidehisa Arai, Keishi Ishihara, Tsubasa Takahashi +1
World models have emerged as promising neural simulators for autonomous driving, with the potential to supplement scarce real-world data and enable closed-loop evaluations. However…