12 citations · 22 across the 11 of their papers we have counts for
10 papers · 1 filter
EvalCards: A Framework for Standardized Evaluation Reporting
Ruchira Dhar, Danae Sanchez Villegas, Antonia Karamolegkou +11
Evaluation has long been a central concern in NLP, and transparent reporting practices are more critical than ever in today's landscape of rapidly released open-access models. Draw…
Evaluating Webcam-based Gaze Data as an Alternative for Human Rationale Annotations
Stephanie Brandl, Oliver Eberle, Tiago Ribeiro +2
Rationales in the form of manually annotated input spans usually serve as ground truth when evaluating explainability methods in NLP. They are, however, time-consuming and often bi…
WebQAmGaze: A Multilingual Webcam Eye-Tracking-While-Reading Dataset
Tiago Ribeiro, Stephanie Brandl, Anders Søgaard +1
We present WebQAmGaze, a multilingual low-cost eye-tracking-while-reading dataset, designed as the first webcam-based eye-tracking corpus of reading to support the development of e…
Domain-Specific Word Embeddings with Structure Prediction
Stephanie Brandl, David Lassner, Anne Baillot +1
Complementary to finding good general word embeddings, an important question for representation learning is to find dynamic word embeddings, e.g., across time or domain. Current me…
Every word counts: A multilingual analysis of individual human alignment with model attention
Stephanie Brandl, Nora Hollenstein
Human fixation patterns have been shown to correlate strongly with Transformer-based attention. Those correlation analyses are usually carried out without taking into account indiv…
Evaluating Deep Taylor Decomposition for Reliability Assessment in the Wild
Stephanie Brandl, Daniel Hershcovich, Anders Søgaard
We argue that we need to evaluate model interpretability methods 'in the wild', i.e., in situations where professionals make critical decisions, and models can potentially assist t…