collaborators

5 papers

cs.LG2025

Bridging Vision, Language, and Mathematics: Pictographic Character Reconstruction with Bézier Curves

Zihao Wan, Pau Tong Lin Xu, Fuwen Luo +3

While Vision-language Models (VLMs) have demonstrated strong semantic capabilities, their ability to interpret the underlying geometric structure of visual information is less expl…

cs.CL2025

MUCAR: Benchmarking Multilingual Cross-Modal Ambiguity Resolution for Multimodal Large Language Models

Xiaolong Wang, Zhaolu Kang, Wangyuxuan Zhai +8

Multimodal Large Language Models (MLLMs) have demonstrated significant advances across numerous vision-language tasks. MLLMs have shown promising capability in aligning visual and…

cs.CV2025

EscapeCraft: A 3D Room Escape Environment for Benchmarking Complex Multimodal Reasoning Ability

Ziyue Wang, Yurui Dong, Fuwen Luo +5

The rapid advancing of Multimodal Large Language Models (MLLMs) has spurred interest in complex multimodal reasoning tasks in the real-world and virtual environment, which require…

cs.CV2025

DongbaMIE: A Multimodal Information Extraction Dataset for Evaluating Semantic Understanding of Dongba Pictograms

Xiaojun Bi, Shuo Li, Junyao Xing +7

Dongba pictographic is the only pictographic script still in use in the world. Its pictorial ideographic features carry rich cultural and contextual information. However, due to th…

cs.CL2025

Perspective Transition of Large Language Models for Solving Subjective Tasks

Xiaolong Wang, Yuanchi Zhang, Ziyue Wang +5

Large language models (LLMs) have revolutionized the field of natural language processing, enabling remarkable progress in various tasks. Different from objective tasks such as com…