2 papers
cs.CL2024
A Simple Joint Model for Improved Contextual Neural Lemmatization
Chaitanya Malaviya, Shijie Wu, Ryan Cotterell
English verbs have multiple forms. For instance, talk may also appear as talks, talked or talking, depending on the context. The NLP task of lemmatization seeks to map these divers…
cs.CL2024
MixCE: Training Autoregressive Language Models by Mixing Forward and Reverse Cross-Entropies
Shiyue Zhang, Shijie Wu, Ozan Irsoy +4
Autoregressive language models are trained by minimizing the cross-entropy of the model distribution Q relative to the data distribution P -- that is, minimizing the forward cross-…