1 paper
Ji Soo Lee, Jongha Kim, Jeehye Na +2
Despite the advancements of Video Large Language Models (VideoLLMs) in various tasks, they struggle with fine-grained temporal understanding, such as Dense Video Captioning (DVC).…