4 papers · 1 filter
CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts
Lianyu Hu, Shengqian Qin, Zeqin Liao +4
Chain-of-thought (CoT) reasoning has enabled multi-modal large language models (MLLMs) to tackle complex visual reasoning tasks by generating explicit intermediate reasoning steps…
Temporal-Aware Reasoning Optimization for Video Temporal Grounding
Minghang Zheng, Zihao Yin, Yi Yang +2
Multi-modal Large Language Models (MLLMs) have achieved remarkable progress in video temporal grounding with reinforcement learning for generating reasoning paths. However, existin…
TVI-CoT: Text-Visual Interleaved Chain-of-Thought Reasoning for Multimodal Understanding
Lianyu Hu, Xiaoyu Ma, Zeqin Liao +1
Chain-of-thought (CoT) reasoning has proven effective for enhancing problem-solving in large language models. However, when applied to multimodal LLMs (MLLMs), existing CoT approac…
DynaTokens: Controlling Token Dynamics for Continual Video-Language Understanding
Toan Nguyen, Yang Liu, Celso De Melo +1
Continual VideoQA with multimodal LLMs remains challenging because sequential adaptation induces task interference, while storing task-specific prompts becomes impractical as task…