3 papers
eess.AS2026
GigaChat Audio: Time-aware Large Audio Language Model
Aleksandr Kutsakov, Mariia Sadovina, Georgii Gospodinov +4
Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware audio LLM that answers questions with explicit timestamps over up to 1…
eess.AS2026
GigaAM Multilingual: Foundation Model for Underrepresented Languages
Andrei Kuzmenko, Alexandr Maximenko, Aleksandr Kutsakov +5
Despite recent scaling successes, multilingual ASR performance remains highly uneven, with long-tail languages suffering from severe data scarcity. This work addresses the challeng…
eess.AS2025
GigaAM: Efficient Self-Supervised Learner for Speech Recognition
Aleksandr Kutsakov, Alexandr Maximenko, Georgii Gospodinov +2
Self-Supervised Learning (SSL) has demonstrated strong performance in speech processing, particularly in automatic speech recognition. In this paper, we explore an SSL pretraining…