3 citations · 3 across the 4 of their papers we have counts for
4 papers
Scaling Attention Head Analysis via Gradient-Based Attribution in Context-Aware Machine Translation
Paweł Mąka, Yusuf Can Semerci, Jan Scholtes +1
In this paper, we introduce a gradient-based head attribution strategy where the Token-level Max-Margin loss is backpropagated to the attention maps. This framework enables a large…
Analyzing the Attention Heads for Pronoun Disambiguation in Context-aware Machine Translation Models
Paweł Mąka, Yusuf Can Semerci, Jan Scholtes +1
In this paper, we investigate the role of attention heads in Context-aware Machine Translation models for pronoun disambiguation in the English-to-German and English-to-French lang…
Sequence Shortening for Context-Aware Machine Translation
Paweł Mąka, Yusuf Can Semerci, Jan Scholtes +1
Context-aware Machine Translation aims to improve translations of sentences by incorporating surrounding sentences as context. Towards this task, two main architectures have been a…
TENT: Tensorized Encoder Transformer for Temperature Forecasting
Onur Bilgin, Paweł Mąka, Thomas Vergutz +1
Reliable weather forecasting is of great importance in science, business, and society. The best performing data-driven models for weather prediction tasks rely on recurrent or conv…