16 papers
Noisy-Channel Minimum Bayes Risk Decoding
Yusuke Sakai, Hidetaka Kamigaito, Taro Watanabe
Minimum Bayes Risk (MBR) decoding yields more robust and higher-quality text generation than maximum a posteriori (MAP) decoding by selecting hypotheses that maximize expected util…
Multilinguality of Large Language Models From a Structural Perspective
Haruki Sakajo, Yusuke Sakai, Hidetaka Kamigaito +1
Large language models (LLMs) have excelled in processing multiple languages through pre- and post-training on multilingual data, even though English dominates the training data. Pr…
Enhancing Factuality through Consensus and Consistency in Summarization Using Minimum Bayes Risk Decoding
Riza Setiawan Soetedjo, Yusuke Sakai, Hidetaka Kamigaito +3
Improving the quality of model-generated summaries, especially factuality, the accuracy of a summary with respect to its source content, remains a challenge. While reranking could…
CArtBench: Evaluating Vision-Language Models on Chinese Art Understanding, Interpretation, and Authenticity
Xuefeng Wei, Zhixuan Wang, Xuan Zhou +5
We introduce CARTBENCH, a museum-grounded benchmark for evaluating vision-language models (VLMs) on Chinese artworks beyond short-form recognition and QA. CARTBENCH comprises four…
Routing by Analogy: kNN-Augmented Expert Assignment for Mixture-of-Experts
Boxuan Lyu, Soichiro Murakami, Hidetaka Kamigaito +1
Mixture-of-Experts (MoE) architectures scale large language models efficiently by employing a parametric ``router'' to dispatch tokens to a sparse subset of experts. Typically, thi…
StructLens: A Structural Lens for Language Models via Maximum Spanning Trees
Haruki Sakajo, Frederikus Hudi, Yusuke Sakai +2
Language exhibits inherent structures, a property that explains both language acquisition and language change. Given this characteristic, we expect language models to manifest thei…