5 papers · 1 filter
Neither Here Nor There: Cross-Lingual Representation Dynamics of Code-Mixed Text in Multilingual Encoders
Debajyoti Mazumder, Divyansh Pathak, Prashant Kodali +1
Multilingual encoder-based language models are widely adopted for code-mixed analysis tasks, yet we know surprisingly little about how they represent code-mixed inputs internally -…
Adapting Multilingual Models to Code-Mixed Tasks via Model Merging
Prashant Kodali, Vaishnavi Shivkumar, Swarang Joshi +3
We study model merging as a practical alternative to conventional adaptation strategies for code-mixed NLP. Starting from a multilingual base model, we: (i) perform continued pre-t…
UPDESH: Synthesizing Grounded Instruction Tuning Data for 13 Indic Languages
Pranjal A. Chitale, Varun Gumma, Sanchit Ahuja +4
Developing culturally grounded multilingual AI systems remains challenging, particularly for low-resource languages. While synthetic data offers promise, its effectiveness in multi…
HashSet -- A Dataset For Hashtag Segmentation
Prashant Kodali, Akshala Bhatnagar, Naman Ahuja +2
Hashtag segmentation is the task of breaking a hashtag into its constituent tokens. Hashtags often encode the essence of user-generated posts, along with information like topic and…
Battling Hateful Content in Indic Languages HASOC '21
Aditya Kadam, Anmol Goel, Jivitesh Jain +7
The extensive rise in consumption of online social media (OSMs) by a large number of people poses a critical problem of curbing the spread of hateful content on these platforms. Wi…