activity
20242026
collaborators

5 papers

cs.CL2026

Revisiting Non-Verbatim Memorization in Large Language Models: The Role of Entity Surface Forms

Yuto Nishida, Naoki Shikoda, Yosuke Kishinami +4

Understanding what kinds of factual knowledge large language models (LLMs) memorize is essential for evaluating their reliability and limitations. Entity-based QA is a common frame…

cs.SE2026

TimeMachine-bench: A Benchmark for Evaluating Model Capabilities in Repository-Level Migration Tasks

Ryo Fujii, Makoto Morishita, Kazuki Yano +1

With the advancement of automated software engineering, research focus is increasingly shifting toward practical tasks reflecting the day-to-day work of software engineers. Among t…

cs.CL2025

JaParaPat: A Large-Scale Japanese-English Parallel Patent Application Corpus

Masaaki Nagata, Katsuki Chousa, Norihito Yasuda

We constructed JaParaPat (Japanese-English Parallel Patent Application Corpus), a bilingual corpus of more than 300 million Japanese-English sentence pairs from patent applications…

cs.CL2025

MQM-Chat: Multidimensional Quality Metrics for Chat Translation

Yunmeng Li, Jun Suzuki, Makoto Morishita +2

The complexities of chats pose significant challenges for machine translation models. Recognizing the need for a precise evaluation metric to address the issues of chat translation…

cs.CL2024

An Investigation of Warning Erroneous Chat Translations in Cross-lingual Communication

Yunmeng Li, Jun Suzuki, Makoto Morishita +2

Machine translation models are still inappropriate for translating chats, despite the popularity of translation software and plug-in applications. The complexity of dialogues poses…