34 papers
Attention-Guided Layer Selection for Contrastive Decoding in Large Language Models
Yusuke Sakai, Natthawut Kertkeidkachorn, Kiyoaki Shirai
Contrastive decoding methods such as DoLa improve the factuality of Large Language Models (LLMs) by contrasting the output distributions of mature and premature layers. However, Do…
Noisy-Channel Minimum Bayes Risk Decoding
Yusuke Sakai, Hidetaka Kamigaito, Taro Watanabe
Minimum Bayes Risk (MBR) decoding yields more robust and higher-quality text generation than maximum a posteriori (MAP) decoding by selecting hypotheses that maximize expected util…
Multilinguality of Large Language Models From a Structural Perspective
Haruki Sakajo, Yusuke Sakai, Hidetaka Kamigaito +1
Large language models (LLMs) have excelled in processing multiple languages through pre- and post-training on multilingual data, even though English dominates the training data. Pr…
Enhancing Factuality through Consensus and Consistency in Summarization Using Minimum Bayes Risk Decoding
Riza Setiawan Soetedjo, Yusuke Sakai, Hidetaka Kamigaito +3
Improving the quality of model-generated summaries, especially factuality, the accuracy of a summary with respect to its source content, remains a challenge. While reranking could…
CArtBench: Evaluating Vision-Language Models on Chinese Art Understanding, Interpretation, and Authenticity
Xuefeng Wei, Zhixuan Wang, Xuan Zhou +5
We introduce CARTBENCH, a museum-grounded benchmark for evaluating vision-language models (VLMs) on Chinese artworks beyond short-form recognition and QA. CARTBENCH comprises four…
StructLens: A Structural Lens for Language Models via Maximum Spanning Trees
Haruki Sakajo, Frederikus Hudi, Yusuke Sakai +2
Language exhibits inherent structures, a property that explains both language acquisition and language change. Given this characteristic, we expect language models to manifest thei…