7 papers
Bridging Vision, Language, and Mathematics: Pictographic Character Reconstruction with Bézier Curves
Zihao Wan, Pau Tong Lin Xu, Fuwen Luo +3
While Vision-language Models (VLMs) have demonstrated strong semantic capabilities, their ability to interpret the underlying geometric structure of visual information is less expl…
MUCAR: Benchmarking Multilingual Cross-Modal Ambiguity Resolution for Multimodal Large Language Models
Xiaolong Wang, Zhaolu Kang, Wangyuxuan Zhai +8
Multimodal Large Language Models (MLLMs) have demonstrated significant advances across numerous vision-language tasks. MLLMs have shown promising capability in aligning visual and…
Agent-Environment Alignment via Automated Interface Generation
Kaiming Liu, Xuanyu Lei, Ziyue Wang +2
Large language model (LLM) agents have shown impressive reasoning capabilities in interactive decision-making tasks. These agents interact with environment through intermediate int…
Thinking with Visual Abstract: Enhancing Multimodal Reasoning via Visual Abstraction
Dairu Liu, Ziyue Wang, Minyuan Ruan +4
Images usually convey richer detail than text, but often include redundant information, which potentially downgrades multimodal reasoning performance. When faced with lengthy or co…
EscapeCraft: A 3D Room Escape Environment for Benchmarking Complex Multimodal Reasoning Ability
Ziyue Wang, Yurui Dong, Fuwen Luo +5
The rapid advancing of Multimodal Large Language Models (MLLMs) has spurred interest in complex multimodal reasoning tasks in the real-world and virtual environment, which require…
DongbaMIE: A Multimodal Information Extraction Dataset for Evaluating Semantic Understanding of Dongba Pictograms
Xiaojun Bi, Shuo Li, Junyao Xing +7
Dongba pictographic is the only pictographic script still in use in the world. Its pictorial ideographic features carry rich cultural and contextual information. However, due to th…