565 citations · 779 across the 29 of their papers we have counts for
35 papers
What do Reward Models Memorize?
Ivo Verhoeven, Pushkar Mishra, Ekaterina Shutova
This paper studies what discriminatively trained reward models (RMs) memorize by measuring counterfactual memorization on two human preference datasets. We show that RMs 1) misallo…
Yesterday's News: Benchmarking Multi-Dimensional Out-of-Distribution Generalization of Misinformation Detection Models
Ivo Verhoeven, Pushkar Mishra, Ekaterina Shutova
This article introduces misinfo-general, a benchmark dataset for evaluating misinformation models' ability to perform out-of-distribution generalization. Misinformation changes rap…
A (More) Realistic Evaluation Setup for Generalisation of Community Models on Malicious Content Detection
Ivo Verhoeven, Pushkar Mishra, Rahel Beloch +2
Community models for malicious content detection, which take into account the context from a social graph alongside the content itself, have shown remarkable performance on benchma…
What's the Meaning of Superhuman Performance in Today's NLU?
Simone Tedeschi, Johan Bos, Thierry Declerck +9
In the last five years, there has been a significant focus in Natural Language Processing (NLP) on developing larger Pretrained Language Models (PLMs) and introducing benchmarks su…
How do languages influence each other? Studying cross-lingual data sharing during LM fine-tuning
Rochelle Choenni, Dan Garrette, Ekaterina Shutova
Multilingual large language models (MLLMs) are jointly trained on data from many different languages such that representation of individual languages can benefit from other languag…
CK-Transformer: Commonsense Knowledge Enhanced Transformers for Referring Expression Comprehension
Zhi Zhang, Helen Yannakoudakis, Xiantong Zhen +1
The task of multimodal referring expression comprehension (REC), aiming at localizing an image region described by a natural language expression, has recently received increasing a…