2 papers
cs.CV2026
PLaMo 2.1-VL Technical Report
Tommi Kerola, Yuya Masuda, Takashi Masuko +5
We introduce PLaMo 2.1-VL, a lightweight Vision Language Model (VLM) for autonomous devices, available in 8B and 2B variants and designed for local and edge deployment with Japanes…
eess.AS2024
Enhancing Large Language Model-based Speech Recognition by Contextualization for Rare and Ambiguous Words
Kento Nozawa, Takashi Masuko, Toru Taniguchi
We develop a large language model (LLM) based automatic speech recognition (ASR) system that can be contextualized by providing keywords as prior information in text prompts. We ad…