most citedAILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

4 citations · 4 across the 4 of their papers we have counts for

collaborators

7 papers

cs.AI2026

SWE-chat: Coding Agent Interactions From Real Users in the Wild

Joachim Baumann, Vishakh Padmakumar, Xiang Li +3

AI coding agents are being adopted at scale, yet we lack empirical evidence on how people actually use them and how much of their output is useful in practice. We present SWE-chat,…

cs.IR2026

GoogleTrendArchive: A Year-Long Archive of Real-Time Web Search Trends Worldwide

Aleksandra Urman, Anikó Hannák, Joachim Baumann

GoogleTrendArchive is a comprehensive archive of Google Trending Now data spanning over one year (from November 28, 2024 to January 3, 2026) across 125 countries and 1,358 location…

cs.CL2025

Auditing Google's AI Overviews and Featured Snippets: A Case Study on Baby Care and Pregnancy

Desheng Hu, Joachim Baumann, Aleksandra Urman +4

Google Search increasingly surfaces AI-generated content through features like AI Overviews (AIO) and Featured Snippets (FS), which users frequently rely on despite having no contr…

cs.AI2025

Reduced AI Acceptance After the Generative AI Boom: Evidence From a Two-Wave Survey Study

Joachim Baumann, Aleksandra Urman, Ulrich Leicht-Deobald +3

The rapid adoption of generative artificial intelligence (GenAI) technologies has led many organizations to integrate AI into their products and services, often without considering…

cs.CL2025

Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation

Joachim Baumann, Paul Röttger, Aleksandra Urman +4

Large language models are rapidly transforming social science research by enabling the automation of labor-intensive tasks like data annotation and text analysis. However, LLM outp…

cs.CL2025

SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors

Tiancheng Hu, Joachim Baumann, Lorenzo Lupo +3

Large language model (LLM) simulations of human behavior have the potential to revolutionize the social and behavioral sciences, if and only if they faithfully reflect real human b…