3 papers
cs.CL2026
SEA-SpeechBench: A Large-Scale Multitask Benchmark for Speech Understanding Across Southeast Asia
Jingyi Liao, Wenyu Zhang, Zhuohan Liu +6
The rapid advancement of audio and multimodal large language models has unlocked transformative speech understanding capabilities, yet evaluation frameworks remain predominantly En…
cs.CL2026
Automated Creativity Evaluation of Language Models Across Open-Ended Tasks
Min Sen Tan, Zachary Kit Chun Choy, Syed Ali Redha Alsagoff +4
Large language models (LLMs) have achieved remarkable progress in language understanding, reasoning, and generation, sparking growing interest in their creative potential. Realizin…
cs.CL2026
Does Less Hallucination Mean Less Creativity? An Empirical Investigation in LLMs
Mohor Banerjee, Nadya Yuki Wangsajaya, Syed Ali Redha Alsagoff +3
Large Language Models (LLMs) exhibit remarkable capabilities in natural language understanding and reasoning, but suffer from hallucination: the generation of factually incorrect c…