1 paper · 1 filter
Siddhant Sukhani, Yash Bhardwaj, Riya Bhadani +3
We evaluate multimodal large language models (MLLMs) for topic-aligned captioning in financial short-form videos (SVs) by testing joint reasoning over transcripts (T), audio (A), a…