5 papers
Cyberbullying Classifiers are Sensitive to Model-Agnostic Perturbations
Chris Emmery, Ákos Kádár, Grzegorz Chrupała +1
A limited amount of studies investigates the role of model-agnostic adversarial behavior in toxic content classification. As toxicity classifiers predominantly rely on lexical cues…
Turing: an Accurate and Interpretable Multi-Hypothesis Cross-Domain Natural Language Database Interface
Peng Xu, Wenjie Zi, Hamidreza Shahidi +7
A natural language database interface (NLDB) can democratize data-driven insights for non-technical users. However, existing Text-to-SQL semantic parsers cannot achieve high enough…
Subword Pooling Makes a Difference
Judit Ács, Ákos Kádár, András Kornai
Contextual word-representations became a standard in modern natural language processing systems. These models use subword tokenization to handle large vocabularies and unknown word…
Adversarial Stylometry in the Wild: Transferable Lexical Substitution Attacks on Author Profiling
Chris Emmery, Ákos Kádár, Grzegorz Chrupała
Written language contains stylistic cues that can be exploited to automatically infer a variety of potentially sensitive author information. Adversarial stylometry intends to attac…
Bootstrapping Disjoint Datasets for Multilingual Multimodal Representation Learning
Ákos Kádár, Grzegorz Chrupała, Afra Alishahi +1
Recent work has highlighted the advantage of jointly learning grounded sentence representations from multiple languages. However, the data used in these studies has been limited to…