papers

Publications (10)

cs.CL2018

Card-660: Cambridge Rare Word Dataset - a Reliable Benchmark for Infrequent Word Representation Models

Mohammad Taher Pilehvar, Dimitri Kartsaklis, Victor Prokhorov +1

Rare word representation has recently enjoyed a surge of interest, owing to the crucial role that effective handling of infrequent words can play in accurate semantic understanding…

cs.CL2018

Unseen Word Representation by Aligning Heterogeneous Lexical Semantic Spaces

Victor Prokhorov, Mohammad Taher Pilehvar, Dimitri Kartsaklis +2

Word embedding techniques heavily rely on the abundance of training data for individual words. Given the Zipfian distribution of words in natural language texts, a large number of…

cs.CL2019

On the Importance of the Kullback-Leibler Divergence Term in Variational Autoencoders for Text Generation

Victor Prokhorov, Ehsan Shareghi, Yingzhen Li +2

Variational Autoencoders (VAEs) are known to suffer from learning uninformative latent representation of the input due to issues such as approximated posterior collapse, or entangl…

cs.AI2026

A Benchmark for Deep Information Synthesis

Debjit Paul, Daniel Murphy, Milan Gritta +14

Large language model (LLM)-based agents are increasingly used to solve complex tasks involving tool use, such as web browsing, code execution, and data analysis. However, current e…

cs.LG2024

Autoencoding Conditional Neural Processes for Representation Learning

Victor Prokhorov, Ivan Titov, N. Siddharth

Conditional neural processes (CNPs) are a flexible and efficient family of models that learn to learn a stochastic process from data. They have seen particular application in conte…

cs.CL2021

Unsupervised Representation Disentanglement of Text: An Evaluation on Synthetic Datasets

Lan Zhang, Victor Prokhorov, Ehsan Shareghi

To highlight the challenges of achieving representation disentanglement for text domain in an unsupervised setting, in this paper we select a representative set of successfully app…