Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
TVI-CoT: Text-Visual Interleaved Chain-of-Thought Reasoning for Multimodal Understanding
Lianyu Hu, Xiaoyu Ma, Zeqin Liao +1
Chain-of-thought (CoT) reasoning has proven effective for enhancing problem-solving in large language models. However, when applied to multimodal LLMs (MLLMs), existing CoT approac…
cs.CV2025
Mixup Helps Understanding Multimodal Video Better
Xiaoyu Ma, Ding Ding, Hao Chen
Multimodal video understanding plays a crucial role in tasks such as action recognition and emotion classification by combining information from different modalities. However, mult…