2 citations · 3 across the 5 of their papers we have counts for
7 papers
CGELBank: CGEL as a Framework for English Syntax Annotation
Brett Reynolds, Aryaman Arora, Nathan Schneider
We introduce the syntactic formalism of the \textit{Cambridge Grammar of the English Language} (CGEL) to the world of treebanking through the CGELBank project. We discuss some issu…
MASALA: Modelling and Analysing the Semantics of Adpositions in Linguistic Annotation of Hindi
Aryaman Arora, Nitin Venkateswaran, Nathan Schneider
We present a completed, publicly available corpus of annotated semantic relations of adpositions and case markers in Hindi. We used the multilingual SNACS annotation scheme, which…
Estimating the Entropy of Linguistic Distributions
Aryaman Arora, Clara Meister, Ryan Cotterell
Shannon entropy is often a quantity of interest to linguists studying the communicative capacity of human language. However, entropy must typically be estimated from observed data…
Computational historical linguistics and language diversity in South Asia
Aryaman Arora, Adam Farris, Samopriya Basu +1
South Asia is home to a plethora of languages, many of which severely lack access to new language technologies. This linguistic diversity also results in a research environment con…
PASTRIE: A Corpus of Prepositions Annotated with Supersense Tags in Reddit International English
Michael Kranzlein, Emma Manning, Siyao Peng +4
We present the Prepositions Annotated with Supersense Tags in Reddit International English ("PASTRIE") corpus, a new dataset containing manually annotated preposition supersenses o…
Hindi-Urdu Adposition and Case Supersenses v1.0
Aryaman Arora, Nitin Venkateswaran, Nathan Schneider
These are the guidelines for the application of SNACS (Semantic Network of Adposition and Case Supersenses; Schneider et al. 2018) to Modern Standard Hindi of Delhi. SNACS is an in…