4 papers · 1 filter
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio
Abdul Basit Tonmoy, Kazi Fardinul Hoque, Md. Shahrier Islam Arham +1
A single embedding space that covers text, images, video, and audio lets one index serve every query a user can pose. Embedding models built on vision-language backbones now lead t…
Alif: Advancing Urdu Large Language Models via Multilingual Synthetic Data Distillation
Muhammad Ali Shafique, Kanwal Mehreen, Muhammad Arham +3
Developing a high-performing large language models (LLMs) for low-resource languages such as Urdu, present several challenges. These challenges include the scarcity of high-quality…
Improving Multilingual Capabilities with Cultural and Local Knowledge in Large Language Models While Enhancing Native Performance
Ram Mohan Rao Kadiyala, Siddartha Pullakhandam, Siddhant Gupta +6
Large Language Models (LLMs) have shown remarkable capabilities, but their development has primarily focused on English and other high-resource languages, leaving many languages un…
1-800-SHARED-TASKS @ NLU of Devanagari Script Languages: Detection of Language, Hate Speech, and Targets using LLMs
Jebish Purbey, Siddartha Pullakhandam, Kanwal Mehreen +4
This paper presents a detailed system description of our entry for the CHiPSAL 2025 shared task, focusing on language detection, hate speech identification, and target detection in…