4 citations · 12 across the 7 of their papers we have counts for
3 papers · 1 filter
DORB: Dynamically Optimizing Multiple Rewards with Bandits
Ramakanth Pasunuru, Han Guo, Mohit Bansal
Policy gradients-based reinforcement learning has proven to be a promising approach for directly optimizing non-differentiable evaluation metrics for language generation tasks. How…
Evaluating Interactive Summarization: an Expansion-Based Framework
Ori Shapira, Ramakanth Pasunuru, Hadar Ronen +3
Allowing users to interact with multi-document summarizers is a promising direction towards improving and customizing summary results. Different ideas for interactive summarization…
Multi-Source Domain Adaptation for Text Classification via DistanceNet-Bandits
Han Guo, Ramakanth Pasunuru, Mohit Bansal
Domain adaptation performance of a learning algorithm on a target domain is a function of its source domain error and a divergence measure between the data distribution of these tw…