1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Mingchao Liu, Yu Sun, Ruixiao Sun +5
Multimodal large language models (MLLMs) are effective at capturing the semantics of short video content; however, they often fail to attend to the policy-specific details required…