activity
20182020
most citedHappyDB: A Corpus of 100,000 Crowdsourced Happy Moments

31 citations · 36 across the 4 of their papers we have counts for

collaborators

8 papers

cs.DB20202 cited

Adaptive Rule Discovery for Labeling Text Data

Sainyam Galhotra, Behzad Golshan, Wang-Chiew Tan

Creating and collecting labeled data is one of the major bottlenecks in machine learning pipelines and the emergence of automated feature generation techniques such as deep learnin…

cs.CL20201 cited

Enhancing Review Comprehension with Domain-Specific Commonsense

Aaron Traylor, Chen Chen, Behzad Golshan +6

Review comprehension has played an increasingly important role in improving the quality of online services and products and commonsense knowledge can further enhance review compreh…

cs.CL2020

SubjQA: A Dataset for Subjectivity and Review Comprehension

Johannes Bjerva, Nikita Bhutani, Behzad Golshan +2

Subjectivity is the expression of internal opinions or beliefs which cannot be objectively observed or verified, and has been shown to be important for sentiment analysis and word-…

cs.CL20192 cited

Essentia: Mining Domain-Specific Paraphrases with Word-Alignment Graphs

Danni Ma, Chen Chen, Behzad Golshan +1

Paraphrases are important linguistic resources for a wide variety of NLP applications. Many techniques for automatic paraphrase mining from general corpora have been proposed. Whil…

cs.CL2019

Emu: Enhancing Multilingual Sentence Embeddings with Semantic Specialization

Wataru Hirota, Yoshihiko Suhara, Behzad Golshan +1

We present Emu, a system that semantically enhances multilingual sentence embeddings. Our framework fine-tunes pre-trained multilingual sentence embeddings using two main component…

cs.SI2018

A Team-Formation Algorithm for Faultline Minimization

Sanaz Bahargam, Behzad Golshan, Theodoros Lappas +1

In recent years, the proliferation of online resumes and the need to evaluate large populations of candidates for on-site and virtual teams have led to a growing interest in automa…