1 citations · 1 across the 3 of their papers we have counts for
5 papers
Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning
Wenxi Gao, Guanxi Lu, Didi Zhu +5
Unified multimodal models (UMMs) with interleaved reasoning, which generate both textual and visual steps as part of intermediate reasoning traces, have demonstrated great potentia…
VisualFLIP: Do Predictions Depend on Task-Critical Visual Evidence in Multimodal Reasoning?
Didi Zhu, Changrui Chen, Stefanos Zafeiriou +1
When a multimodal large language model answers a visual reasoning question correctly, is the prediction actually supported by the task-critical visual evidence? Correct answers can…
Watch Wider and Think Deeper: Collaborative Cross-modal Chain-of-Thought for Complex Visual Reasoning
Wenting Lu, Didi Zhu, Tao Shen +3
Multi-modal reasoning requires the seamless integration of visual and linguistic cues, yet existing Chain-of-Thought methods suffer from two critical limitations in cross-modal sce…
Keeping Yourself is Important in Downstream Tuning Multimodal Large Language Model
Wenke Huang, Jian Liang, Xianda Guo +14
Multi-modal Large Language Models (MLLMs) integrate visual and linguistic reasoning to address complex tasks such as image captioning and visual question answering. While MLLMs dem…
Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning
Wenke Huang, Jian Liang, Zekun Shi +6
Multimodal Large Language Model (MLLM) have demonstrated strong generalization capabilities across diverse distributions and tasks, largely due to extensive pre-training datasets.…