Publications (58)
Cross-lingual Text Classification Transfer: The Case of Ukrainian
Daryna Dementieva, Valeriia Khylenko, Georg Groh
Despite the extensive amount of labeled datasets in the NLP text classification field, the persistent imbalance in data availability across various languages remains evident. To su…
It Depends: Resolving Referential Ambiguity in Minimal Contexts with Commonsense Knowledge
Lukas Ellinger, Georg Groh
Ambiguous words or underspecified references require interlocutors to resolve them, often by relying on shared context and commonsense knowledge. Therefore, we systematically inves…
DIALECTIC: A Multi-Agent System for Startup Evaluation
Jae Yoon Bae, Simon Malberg, Joyce Galang +2
Venture capital (VC) investors face a large number of investment opportunities but only invest in few of these, with even fewer ending up successful. Early-stage screening of oppor…
End-to-End Annotator Bias Approximation on Crowdsourced Single-Label Sentiment Analysis
Gerhard Johann Hagerer, David Szabo, Andreas Koch +5
Sentiment analysis is often a crowdsourcing task prone to subjective labels given by many annotators. It is not yet fully understood how the annotation bias of each annotator can b…
Language Models for German Text Simplification: Overcoming Parallel Data Scarcity through Style-specific Pre-training
Miriam Anschütz, Joshua Oehms, Thomas Wimmer +2
Automatic text simplification systems help to reduce textual information barriers on the internet. However, for languages other than English, only few parallel data to train these…
German4All -- A Dataset and Model for Readability-Controlled Paraphrasing in German
Miriam Anschütz, Thanh Mai Pham, Eslam Nasrallah +3
The ability to paraphrase texts across different complexity levels is essential for creating accessible texts that can be tailored toward diverse reader groups. Thus, we introduce…