31 papers
Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning
Yiyang Fang, Pei Fu, Jinjie Li +7
Multimodal Large Language Models (MLLMs) often follow a fixed Think-then-Answer paradigm, which is inefficient in heterogeneous multitask settings because simple inputs may not req…
DeltaV: Thinking with Visual State Updates in Unified Large Multimodal Models
Pengjie Wang, Linger Deng, Zujia Zhang +6
Current Unified Large Multimodal Models (ULMMs) support interleaved multimodal reasoning through textual reasoning and intermediate visual states, but typically generate each visua…
EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models
Yiyang Fang, Wenke Huang, Pei Fu +5
Multimodal Large Language Models (MLLMs) have shown remarkable progress in visual reasoning and understanding tasks but still struggle to capture the complexity and subjectivity of…
ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval
Yuhan Liu, Pei Fu, Hang Li +8
Leveraging Multimodal Large Language Models (MLLMs) via contrastive learning has become a mainstream paradigm for improving the performance of Universal Multimodal Retrieval (UMR).…
UniTranslator: A Unified Multi-modal Framework for End-to-end In-Image Machine Translation
Jiahao Lyu, Pei Fu, Zhenhang Li +6
In-Image Machine Translation (IIMT) aims to translate scene text in an image and render the translated text back into the original regions while preserving the overall visual appea…
RS-Gen: A Multi-Stage Agentic Framework for Reasoning and Search-Augmented Image Generation
Feifei Bian, Zhimin Zheng, Wei Deng +2
Recent years have witnessed remarkable progress in image generation and editing, particularly regarding instruction following and visual fidelity. However, when handling ambiguous…