5 papers
The Blindness of Document-Level Translation Evaluation
Ahrii Kim, Vilém Zouhar, Chanjun Park +1
Document-level machine translation (MT) evaluation extends segment-level protocols by presenting full documents to annotators, on the assumption that such presentation elicits docu…
Discourse Dependency: A Continuous Criterion for Translation Difficulty
Ahrii Kim, Chanjun Park, Seong-heum Kim
Recent calls for harder machine translation benchmarks have not clarified what difficulty should mean. We argue that one meaningful and currently unmeasured axis is referential rea…
Tokenization and Morphological Fidelity in Uralic NLP: A Cross-Lingual Evaluation
Nuo Xu, Ahrii Kim
Subword tokenization critically affects Natural Language Processing (NLP) performance, yet its behavior in morphologically rich and low-resource language families remains under-exp…
Do LLMs Truly Benefit from Longer Context in Automatic Post-Editing?
Ahrii Kim, Seong-heum Kim
Automatic post-editing (APE) aims to refine machine translations by correcting residual errors. Although recent large language models (LLMs) demonstrate strong translation capabili…
Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
Sara Papi, Javier Garcia Gilabert, Zachary Hopton +8
As Large Language Models (LLMs) expand beyond text, integrating speech as a native modality has given rise to SpeechLLMs, which directly process spoken language and enable speech-t…