3 papers
cs.AI2026
ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding
Keuntae Kim, Beomseok Lee, Hyunwoo Kim +1
Vision Language Models (VLMs) achieve strong reasoning with Chain-of-Thought (CoT) prompting but incur high sequential-generation cost, error accumulation, and limited self-correct…
cs.CL2025
NAVER LABS Europe Submission to the Instruction-following Track
Beomseok Lee, Marcely Zanon Boito, Laurent Besacier +1
In this paper we describe NAVER LABS Europe submission to the instruction-following speech processing short track at IWSLT 2025. We participate in the constrained settings, develop…
cs.CL2024
Speech Foundation Models and Crowdsourcing for Efficient, High-Quality Data Collection
Beomseok Lee, Marco Gaido, Ioan Calapodescu +2
While crowdsourcing is an established solution for facilitating and scaling the collection of speech data, the involvement of non-experts necessitates protocols to ensure final dat…