1 paper
Yuanhao Zou, Arthad Kulkarni, Lucas Tonanez +12
Multimodal large language models have made rapid progress in video temporal grounding, yet real-world applications routinely require localizing every event that satisfies compositi…