2 papers
cs.CV2026
v-HUB: A Benchmark for Video Humor Understanding from Vision and Sound
Zhengpeng Shi, Yanpeng Zhao, Jianqun Zhou +6
AI models capable of comprehending humor hold real-world promise -- for example, enhancing engagement in human-machine interactions. To gauge and diagnose the capacity of multimoda…
cs.CL2026
Multicultural Spyfall: Assessing LLMs through Dynamic Multilingual Social Deduction Game
Haryo Akbarianto Wibowo, Alaa Elsetohy, Qinrong Cui +1
The rapid advancement of Large Language Models (LLMs) has necessitated more robust evaluation methods that go beyond static benchmarks, which are increasingly prone to data saturat…