643 citations
- Ludwig-Maximilians-Universität MünchenDE328 papers
- Centre National de la Recherche ScientifiqueFR76 papers
- Instituto de Astrofísica de CanariasES74 papers
- University College LondonGB67 papers
- University of BonnDE66 papers
- Universidad de La LagunaES64 papers
- California Institute of TechnologyUS62 papers
- Centro de Investigaciones Energéticas, Medioambientales y TecnológicasES62 papers
- Institute for High Energy PhysicsES62 papers
- Trieste Astronomical ObservatoryIT62 papers
- Universidad Autónoma de MadridES62 papers
- Université Bourgogne Franche-ComtéFR62 papers
10 papers · 1 filter
Rethinking Ground Truth: A Case Study on Human Label Variation in MLLM Benchmarking
Tomas Ruiz, Tanalp Agustoslu, Carsten Schwemmer
Human Label Variation (HLV), i.e. systematic differences among annotators' judgments, remains underexplored in benchmarks despite rapid progress in large language model (LLM) devel…
Survey Response Generation: Generating Closed-Ended Survey Responses In-Silico with Large Language Models
Georg Ahnert, Anna-Carolina Haensch, Barbara Plank +1
Many in-silico simulations of human survey responses with large language models (LLMs) focus on generating closed-ended survey responses, whereas LLMs are typically trained to gene…
AIn't Nothing But a Survey? Using Large Language Models for Coding German Open-Ended Survey Responses on Survey Motivation
Leah von der Heyde, Anna-Carolina Haensch, Bernd Weiß +1
The recent development and wider accessibility of LLMs have spurred discussions about how they can be used in survey research, including classifying open-ended survey responses. Du…
CRAFT Your Dataset: Task-Specific Synthetic Dataset Generation Through Corpus Retrieval and Augmentation
Ingo Ziegler, Abdullatif Köksal, Desmond Elliott +1
Building high-quality datasets for specialized tasks is a time-consuming and resource-intensive process that often requires specialized domain knowledge. We propose Corpus Retrieva…
Applying QNLP to sentiment analysis in finance
Jonas Stein, Ivo Christ, Nicolas Kraus +3
As an application domain where the slightest qualitative improvements can yield immense value, finance is a promising candidate for early quantum advantage. Focusing on the rapidly…
Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages
Ayyoob Imani, Peiqin Lin, Amir Hossein Kargaran +8
The NLP community has mainly focused on scaling Large Language Models (LLMs) vertically, i.e., making them better for about 100 languages. We instead scale LLMs horizontally: we cr…