1 citations · 1 across the 2 of their papers we have counts for
5 papers
Read the Paper, Write the Code: Agentic Reproduction of Social-Science Results
Benjamin Kohler, David Zollikofer, Johanna Einsiedler +2
Recent work has used LLM agents to reproduce empirical social science results with access to both the data and code. We broaden this scope by asking: Can they reproduce results giv…
Measuring Scalar Constructs in Social Science with LLMs
Hauke Licht, Rupak Sarkar, Patrick Y. Wu +4
Many constructs that characterize language, like its complexity or emotionality, have a naturally continuous semantic structure; a public speech is not just "simple" or "complex,"…
Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
Project Apertus, Alejandro Hernández-Cano, Alexander Hägele +100
We present Apertus, a fully open suite of large language models (LLMs) designed to address two systemic shortcomings in today's open model ecosystem: data compliance and multilingu…
Co-DETECT: Collaborative Discovery of Edge Cases in Text Classification
Chenfei Xiong, Jingwei Ni, Yu Fan +10
We introduce Co-DETECT (Collaborative Discovery of Edge cases in TExt ClassificaTion), a novel mixed-initiative annotation framework that integrates human expertise with automatic…
The Medium Is Not the Message: Deconfounding Document Embeddings via Linear Concept Erasure
Yu Fan, Yang Tian, Shauli Ravfogel +3
Embedding-based similarity metrics between text sequences can be influenced not just by the content dimensions we most care about, but can also be biased by spurious attributes lik…