activity
20152023
most citedQUOTUS: The Structure of Political Media Coverage as Revealed by Quoting Patterns

15 citations · 22 across the 4 of their papers we have counts for

collaborators
Showing cs.CLShow all

9 papers · 1 filter

cs.CL2023

Joint Dropout: Improving Generalizability in Low-Resource Neural Machine Translation through Phrase Pair Variables

Ali Araabi, Vlad Niculae, Christof Monz

Despite the tremendous success of Neural Machine Translation (NMT), its performance on low-resource language pairs still remains subpar, partly due to the limited ability to handle…

cs.CL2023

Viewing Knowledge Transfer in Multilingual Machine Translation Through a Representational Lens

David Stap, Vlad Niculae, Christof Monz

We argue that translation quality alone is not a sufficient metric for measuring knowledge transfer in multilingual neural machine translation. To support this claim, we introduce…

cs.CL2020

Understanding the Mechanics of SPIGOT: Surrogate Gradients for Latent Structure Learning

Tsvetomila Mihaylova, Vlad Niculae, André F. T. Martins

Latent structure models are a powerful tool for modeling language data: they can mitigate the error propagation and annotation bottleneck in pipeline systems, while simultaneously…

cs.CL2019

Adaptively Sparse Transformers

Gonçalo M. Correia, Vlad Niculae, André F. T. Martins

Attention mechanisms have become ubiquitous in NLP. Recent architectures, notably the Transformer, learn powerful context-aware word representations through layered, multi-headed a…

cs.CL20197 cited

Sparse Sequence-to-Sequence Models

Ben Peters, Vlad Niculae, André F. T. Martins

Sequence-to-sequence models are a powerful workhorse of NLP. Most variants employ a softmax transformation in both their attention mechanism and output layer, leading to dense alig…

cs.CL2018

Towards Dynamic Computation Graphs via Sparse Latent Structure

Vlad Niculae, André F. T. Martins, Claire Cardie

Deep NLP models benefit from underlying structures in the data---e.g., parse trees---typically extracted using off-the-shelf parsers. Recent attempts to jointly learn the latent st…