6 citations · 18 across the 22 of their papers we have counts for
1 paper · 1 filter
Rohit Saxena, Aryo Pradipta Gema, Pasquale Minervini
Understanding time from visual representations is a fundamental cognitive skill, yet it remains a challenge for multimodal large language models (MLLMs). In this work, we investiga…