2 papers
cs.IR2026
A Shared IPTC Topic Space for Cross-Source Topic Modelling
Din Iskakov, Sebastian Gonçalves, Marco Idiat +4
Comparing topic attention across different media is hindered by a fundamental modelling problem: topic models fitted separately to each corpus produce corpus-specific topic spaces…
cs.CL2024
Adapting Chat Language Models Using Only Target Unlabeled Language Data
Atsuki Yamaguchi, Terufumi Morishita, Aline Villavicencio +1
Vocabulary expansion (VE) is the de-facto approach to language adaptation of large language models (LLMs) by adding new tokens and continuing pre-training on target data. While thi…