3 papers
cs.AI2026
AttuneBench: A Conversation-Based Benchmark for LLM Emotional Intelligence
Kate M. Lubrano, Faisal Sayed, Ankita Rathod +4
Emotional intelligence (EI), the ability to perceive, understand, and respond appropriately to others' emotional states, is central to human communication, and increasingly importa…
cs.SD2026
PitchBench: Measuring Pitch Hearing in Audio-Language Models
Milan Liessens Dujardin, Song-Ze Yu, Craver Corbyn Thomas-Smith +2
Audio-language models (ALMs) are increasingly used in real-world applications that require understanding music, from music tutoring and transcription to captioning, recommendation…
cs.SE2026
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
Ben Rank, Hardik Bhatnagar, Ameya Prabhu +4
AI agents have become surprisingly proficient at software engineering over the past year, largely due to improvements in reasoning capabilities. This raises a deeper question: can…