3 papers
cs.CV2026
UM-Text: A Unified Multimodal Model for Image Understanding and Visual Text Editing
Lichen Ma, Xiaolong Fu, Gaojing Zhou +6
With the rapid advancement of image generation, visual text editing using natural language instructions has received increasing attention. The main challenge of this task is to ful…
cs.CL2026
C2F-Thinker: Coarse-to-Fine Reasoning with Hint-Guided Reinforcement Learning for Multimodal Sentiment Analysis
Miaosen Luo, Zhenhao Yang, Jieshen Long +3
Multimodal sentiment analysis aims to integrate textual, acoustic, and visual information for deep emotional understanding. Despite the progress of multimodal large language models…
cs.CV2025
Decoupled Seg Tokens Make Stronger Reasoning Video Segmenter and Grounder
Dang Jisheng, Wu Xudong, Wang Bimei +7
Existing video segmenter and grounder approaches, exemplified by Sa2VA, directly fuse features within segmentation models. This often results in an undesirable entanglement of dyna…