activity
20172026
most citedOvercoming Multi-Model Forgetting

12 citations · 21 across the 15 of their papers we have counts for

collaborators
Showing cs.CLShow all

17 papers · 1 filter

cs.CL2026

Learning to Learn from Language Feedback with Social Meta-Learning

Jonathan Cook, Diego Antognini, Martin Klissarov +2

Large language models (LLMs) often struggle to learn from corrective feedback within a conversational context. They are rarely proactive in soliciting this feedback, even when face…

cs.CL2020

Control, Generate, Augment: A Scalable Framework for Multi-Attribute Text Generation

Giuseppe Russo, Nora Hollenstein, Claudiu Musat +1

We introduce CGA, a conditional VAE architecture, to control, generate, and augment text. CGA is able to generate natural English sentences controlling multiple semantic and syntac…

cs.CL2020

A Swiss German Dictionary: Variation in Speech and Writing

Larissa Schmidt, Lucy Linder, Sandra Djambazovska +3

We introduce a dictionary containing forms of common words in various Swiss German dialects normalized into High German. As Swiss German is, for now, a predominantly spoken languag…

cs.CL2020

Fast Cross-domain Data Augmentation through Neural Sentence Editing

Guillaume Raille, Sandra Djambazovska, Claudiu Musat

Data augmentation promises to alleviate data scarcity. This is most important in cases where the initial data is in short supply. This is, for existing methods, also where augmenti…

cs.CL2019

Automatic Creation of Text Corpora for Low-Resource Languages from the Internet: The Case of Swiss German

Lucy Linder, Michael Jungo, Jean Hennebert +2

This paper presents SwissCrawl, the largest Swiss German text corpus to date. Composed of more than half a million sentences, it was generated using a customized web scraping tool…

cs.CL2019

Alleviating Sequence Information Loss with Data Overlapping and Prime Batch Sizes

Noémien Kocher, Christian Scuito, Lorenzo Tarantino +3

In sequence modeling tasks the token order matters, but this information can be partially lost due to the discretization of the sequence into data points. In this paper, we study t…