activity
20162026
most citedGlot500: Scaling Multilingual Corpora and Language Models to 500 Languages

17 citations · 20 across the 8 of their papers we have counts for

collaborators

12 papers

cs.AI2026

AI Research Preference Models

Thomas Simon Foster, Bassel Al Omari, Tingchen Fu +30

AI research agents (AIRA) can now carry machine learning experiments from proposal through implementation and evaluation. Yet progress on frontier tasks is throttled by the cost of…

cs.CV2024★ 1 cited

Learning Visual Prompts for Guiding the Attention of Vision Transformers

Razieh Rezaei, Masoud Jalili Sabet, Jindong Gu +3

Visual prompting infuses visual information into the input image to adapt models toward specific predictions and tasks. Recently, manually crafted markers such as red circles are s…

cs.CL2023★ 17 cited

Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages

Ayyoob Imani, Peiqin Lin, Amir Hossein Kargaran +8

The NLP community has mainly focused on scaling Large Language Models (LLMs) vertically, i.e., making them better for about 100 languages. We instead scale LLMs horizontally: we cr…

cs.CL2022★ 1 cited

Graph-Based Multilingual Label Propagation for Low-Resource Part-of-Speech Tagging

Ayyoob Imani, Silvia Severini, Masoud Jalili Sabet +2

Part-of-Speech (POS) tagging is an important component of the NLP pipeline, but many low-resource languages lack labeled data for training. An established method for training a POS…

cs.CL2022★ 1 cited

Don't Forget Cheap Training Signals Before Building Unsupervised Bilingual Word Embeddings

Silvia Severini, Viktor Hangya, Masoud Jalili Sabet +2

Bilingual Word Embeddings (BWEs) are one of the cornerstones of cross-lingual transfer of NLP models. They can be built using only monolingual corpora without supervision leading t…

cs.CL2022

CaMEL: Case Marker Extraction without Labels

Leonie Weissweiler, Valentin Hofmann, Masoud Jalili Sabet +1

We introduce CaMEL (Case Marker Extraction without Labels), a novel and challenging task in computational morphology that is especially relevant for low-resource languages. We prop…