activity
20232025
most citedSEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages

2 citations · 5 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CL2025

SparQLe: Speech Queries to Text Translation Through LLMs

Amirbek Djanibekov, Hanan Aldarmaki

With the growing influence of Large Language Models (LLMs), there is increasing interest in integrating speech representations with them to enable more seamless multi-modal process…

cs.SD2025★ 1 cited

Music for All: Representational Bias and Cross-Cultural Adaptability of Music Generation Models

Atharva Mehta, Shivam Chauhan, Amirbek Djanibekov +3

The advent of Music-Language Models has greatly enhanced the automatic music generation capability of AI systems, but they are also limited in their coverage of the musical genres…

cs.CV2024★ 1 cited

All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages

Ashmal Vayani, Dinura Dissanayake, Hasindri Watawana +66

Existing Large Multimodal Models (LMMs) generally focus on only a few regions and languages. As LMMs continue to improve, it is increasingly important to ensure they understand cul…

cs.CL2024★ 1 cited

Dialectal Coverage And Generalization in Arabic Speech Recognition

Amirbek Djanibekov, Hawau Olamide Toyin, Raghad Alshalan +2

Developing robust automatic speech recognition (ASR) systems for Arabic requires effective strategies to manage its diversity. Existing ASR systems mainly cover the modern standard…

cs.CL2024★ 2 cited

SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages

Holy Lovenia, Rahmad Mahendra, Salsabil Maulana Akbar +58

Southeast Asia (SEA) is a region rich in linguistic diversity and cultural variety, with over 1,300 indigenous languages and a population of 671 million people. However, prevailing…

cs.CL2023

ArTST: Arabic Text and Speech Transformer

Hawau Olamide Toyin, Amirbek Djanibekov, Ajinkya Kulkarni +1

We present ArTST, a pre-trained Arabic text and speech transformer for supporting open-source speech technologies for the Arabic language. The model architecture follows the unifie…