5 papers · 1 filter
Humanly: A Configurable and Traceable Environment for Human-AI Collaborative Writing
Shenzhe Zhu, Haoqian Zhang, Xu Yang +7
Teachers, conference chairs, and public readers all judge writing from limited evidence, seeing only a finished document and not the process that produced it. Final text alone cann…
Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability
Alicia Parrish, Rajat Shinde, Sanket Badhe +57
Current AI safety evaluation and benchmarking frameworks predominantly rely on Western-centric culture-agnostic defaults that mask critical regional laws, socio-linguistic nuances,…
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
Tiancheng Hu, Joachim Baumann, Lorenzo Lupo +3
Large language model (LLM) simulations of human behavior have the potential to revolutionize the social and behavioral sciences, if and only if they faithfully reflect real human b…
Auditing Google's AI Overviews and Featured Snippets: A Case Study on Baby Care and Pregnancy
Desheng Hu, Joachim Baumann, Aleksandra Urman +4
Google Search increasingly surfaces AI-generated content through features like AI Overviews (AIO) and Featured Snippets (FS), which users frequently rely on despite having no contr…
Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation
Joachim Baumann, Paul Röttger, Aleksandra Urman +4
Large language models are rapidly transforming social science research by enabling the automation of labor-intensive tasks like data annotation and text analysis. However, LLM outp…