20 papers
YOMI-Bench: A Benchmark for Evaluating Kanji Reading and Phonological Understanding of LLMs for Japanese
Ryota Mibayashi, Hiroya Takamura, Hitomi Yanaka
We propose YOMI-Bench, a benchmark for evaluating kanji reading and phonological understanding of large language models (LLMs) for Japanese. In Japanese, a single kanji character o…
Low-Latency Real-Time Audio Game Commentary System via LLM-Based Parallel Text Generation
Ryota Kawamatsu, Anum Afzal, Yuki Saito +5
We present a low-latency real-time audio game commentary system that generates spoken commentary directly from live gameplay video. In this end-to-end setting, a key bottleneck is…
HOME-KGQA: A Benchmark Dataset for Multimodal Knowledge Graph Question Answering on Household Daily Activities
Shusaku Egami, Aoi Ohta, Tomoki Tsujimura +5
Large Language Models (LLMs) provide flexible natural language processing capabilities, while knowledge graphs (KGs) offer explicit and structured knowledge. Integrating these two…
Multimodal Task Interference: A Benchmark and Analysis of History-Target Mismatch in Multimodal LLMs
Masayuki Kawarada, Tatsuya Ishigaki, Hiroya Takamura
Task interference, the performance degradation caused by task switches within a single conversation, has been studied exclusively in text-only settings despite the growing prevalen…
A Comparative Study of Demonstration Selection for Practical Large Language Models-based Next POI Prediction
Ryo Nishida, Masayuki Kawarada, Tatsuya Ishigaki +2
This paper investigates demonstration selection strategies for predicting a user's next point-of-interest (POI) using large language models (LLMs), aiming to accurately forecast a…
Real-Time Generation of Game Video Commentary with Multimodal LLMs: Pause-Aware Decoding Approaches
Anum Afzal, Yuki Saito, Hiroya Takamura +5
Real-time video commentary generation provides textual descriptions of ongoing events in videos. It supports accessibility and engagement in domains such as sports, esports, and li…