Publications (10)
Card-660: Cambridge Rare Word Dataset - a Reliable Benchmark for Infrequent Word Representation Models
Mohammad Taher Pilehvar, Dimitri Kartsaklis, Victor Prokhorov +1
Rare word representation has recently enjoyed a surge of interest, owing to the crucial role that effective handling of infrequent words can play in accurate semantic understanding…
Unseen Word Representation by Aligning Heterogeneous Lexical Semantic Spaces
Victor Prokhorov, Mohammad Taher Pilehvar, Dimitri Kartsaklis +2
Word embedding techniques heavily rely on the abundance of training data for individual words. Given the Zipfian distribution of words in natural language texts, a large number of…
On the Importance of the Kullback-Leibler Divergence Term in Variational Autoencoders for Text Generation
Victor Prokhorov, Ehsan Shareghi, Yingzhen Li +2
Variational Autoencoders (VAEs) are known to suffer from learning uninformative latent representation of the input due to issues such as approximated posterior collapse, or entangl…
A Benchmark for Deep Information Synthesis
Debjit Paul, Daniel Murphy, Milan Gritta +14
Large language model (LLM)-based agents are increasingly used to solve complex tasks involving tool use, such as web browsing, code execution, and data analysis. However, current e…
Autoencoding Conditional Neural Processes for Representation Learning
Victor Prokhorov, Ivan Titov, N. Siddharth
Conditional neural processes (CNPs) are a flexible and efficient family of models that learn to learn a stochastic process from data. They have seen particular application in conte…
Unsupervised Representation Disentanglement of Text: An Evaluation on Synthetic Datasets
Lan Zhang, Victor Prokhorov, Ehsan Shareghi
To highlight the challenges of achieving representation disentanglement for text domain in an unsupervised setting, in this paper we select a representative set of successfully app…