1 paper · 1 filter
Mingchao Liu, Yu Sun, Ruixiao Sun +5
Multimodal large language models (MLLMs) are effective at capturing the semantics of short video content; however, they often fail to attend to the policy-specific details required…