15 citations · 22 across the 3 of their papers we have counts for
14 papers
Understanding the Mechanics of SPIGOT: Surrogate Gradients for Latent Structure Learning
Tsvetomila Mihaylova, Vlad Niculae, André F. T. Martins
Latent structure models are a powerful tool for modeling language data: they can mitigate the error propagation and annotation bottleneck in pipeline systems, while simultaneously…
Efficient Marginalization of Discrete and Structured Latent Variables via Sparsity
Gonçalo M. Correia, Vlad Niculae, Wilker Aziz +1
Training neural network models with discrete (categorical or structured) latent variables can be computationally challenging, due to the need for marginalization over large or comb…
Sparse and Continuous Attention Mechanisms
André F. T. Martins, António Farinhas, Marcos Treviso +3
Exponential families are widely used in machine learning; they include many distributions in continuous and discrete domains (e.g., Gaussian, Dirichlet, Poisson, and categorical di…
LP-SparseMAP: Differentiable Relaxed Optimization for Sparse Structured Prediction
Vlad Niculae, André F. T. Martins
Structured prediction requires manipulating a large number of combinatorial structures, e.g., dependency trees or alignments, either as latent or output variables. Recently, the Sp…
Adaptively Sparse Transformers
Gonçalo M. Correia, Vlad Niculae, André F. T. Martins
Attention mechanisms have become ubiquitous in NLP. Recent architectures, notably the Transformer, learn powerful context-aware word representations through layered, multi-headed a…
Sparse Sequence-to-Sequence Models
Ben Peters, Vlad Niculae, André F. T. Martins
Sequence-to-sequence models are a powerful workhorse of NLP. Most variants employ a softmax transformation in both their attention mechanism and output layer, leading to dense alig…