collaborators

6 papers

cs.LG2025

Bridging Vision, Language, and Mathematics: Pictographic Character Reconstruction with Bézier Curves

Zihao Wan, Pau Tong Lin Xu, Fuwen Luo +3

While Vision-language Models (VLMs) have demonstrated strong semantic capabilities, their ability to interpret the underlying geometric structure of visual information is less expl…

cs.CL2025

Thinking with Visual Abstract: Enhancing Multimodal Reasoning via Visual Abstraction

Dairu Liu, Ziyue Wang, Minyuan Ruan +4

Images usually convey richer detail than text, but often include redundant information, which potentially downgrades multimodal reasoning performance. When faced with lengthy or co…

cs.CV2025

EscapeCraft: A 3D Room Escape Environment for Benchmarking Complex Multimodal Reasoning Ability

Ziyue Wang, Yurui Dong, Fuwen Luo +5

The rapid advancing of Multimodal Large Language Models (MLLMs) has spurred interest in complex multimodal reasoning tasks in the real-world and virtual environment, which require…

cs.CV2025

DongbaMIE: A Multimodal Information Extraction Dataset for Evaluating Semantic Understanding of Dongba Pictograms

Xiaojun Bi, Shuo Li, Junyao Xing +7

Dongba pictographic is the only pictographic script still in use in the world. Its pictorial ideographic features carry rich cultural and contextual information. However, due to th…

cs.CL2025

Perspective Transition of Large Language Models for Solving Subjective Tasks

Xiaolong Wang, Yuanchi Zhang, Ziyue Wang +5

Large language models (LLMs) have revolutionized the field of natural language processing, enabling remarkable progress in various tasks. Different from objective tasks such as com…

cs.CV2024

StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Junming Lin, Zheng Fang, Chi Chen +5

The rapid development of Multimodal Large Language Models (MLLMs) has expanded their capabilities from image comprehension to video understanding. However, most of these MLLMs focu…