5 papers
Giga-Embeddings: Mixture-of-Experts Encoders for High-Throughput Text Embeddings
Egor Kolodin, Egor Krasnoperov, Evgeniy Kosarev +1
We introduce Giga-Embeddings, a family of text embedding models designed to combine strong retrieval quality with efficient serving. Its largest member is a sparse 10B-parameter Mi…
GigaChat Audio: Time-aware Large Audio Language Model
Aleksandr Kutsakov, Mariia Sadovina, Georgii Gospodinov +4
Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware audio LLM that answers questions with explicit timestamps over up to 1…
GigaAM Multilingual: Foundation Model for Underrepresented Languages
Andrei Kuzmenko, Alexandr Maximenko, Aleksandr Kutsakov +5
Despite recent scaling successes, multilingual ASR performance remains highly uneven, with long-tail languages suffering from severe data scarcity. This work addresses the challeng…
GigaEmbeddings: Efficient Russian Language Embedding Model
Egor Kolodin, Daria Khomich, Nikita Savushkin +2
We introduce GigaEmbeddings, a novel framework for training high-performance Russian-focused text embeddings through hierarchical instruction tuning of the decoder-only LLM designe…
GigaAM: Efficient Self-Supervised Learner for Speech Recognition
Aleksandr Kutsakov, Alexandr Maximenko, Georgii Gospodinov +2
Self-Supervised Learning (SSL) has demonstrated strong performance in speech processing, particularly in automatic speech recognition. In this paper, we explore an SSL pretraining…