5 papers
DeltaV: Thinking with Visual State Updates in Unified Large Multimodal Models
Pengjie Wang, Linger Deng, Zujia Zhang +6
Current Unified Large Multimodal Models (ULMMs) support interleaved multimodal reasoning through textual reasoning and intermediate visual states, but typically generate each visua…
GeoFocus: Blending Efficient Global-to-Local Perception for Multimodal Geometry Problem-Solving
Linger Deng, Yuliang Liu, Wenwen Yu +4
Geometry problem-solving remains a significant challenge for Large Multimodal Models (LMMs), requiring not only global shape recognition but also attention to intricate local relat…
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance
Zhang Li, Biao Yang, Qiang Liu +7
While large multi-modal models (LMMs) demonstrate promising capabilities in segmentation and comprehension, they still struggle with two limitations: inaccurate segmentation and ha…
Theorem-Validated Reverse Chain-of-Thought Problem Generation for Geometric Reasoning
Linger Deng, Linghao Zhu, Yuliang Liu +6
Large Multimodal Models (LMMs) face limitations in geometric reasoning due to insufficient Chain of Thought (CoT) image-text training data. While existing approaches leverage templ…
VimTS: A Unified Video and Image Text Spotter for Enhancing the Cross-domain Generalization
Yuliang Liu, Mingxin Huang, Hao Yan +6
Text spotting, a task involving the extraction of textual information from image or video sequences, faces challenges in cross-domain adaption, such as image-to-image and image-to-…