1 paper · 1 filter
Yongjian Chen, Pengfei Wei, Yiqun Sun +2
Multimodal Large Language Models (MLLMs) process speech and text jointly, yet whether they exploit prosodic cues for pragmatic inference or rely on surface acoustic patterns has re…