2 papers
cs.AI2026
Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results
Jan Batzner, Sree Harsha Nelaturu, Damian Stachura +45
AI evaluations are widely used for testing and understanding progress. However, the diverse evaluators bring with them inconsistencies that challenge analysis and comparison. First…
cs.SI2026
@GrokSet: multi-party Human-LLM Interactions in Social Media
Matteo Migliarini, Berat Ercevik, Oluwagbemike Olowe +5
Large Language Models (LLMs) are increasingly deployed as active participants on public social media platforms, yet their behavior in these unconstrained social environments remain…