collaborators

9 papers

cs.LG2025

Bridging Vision, Language, and Mathematics: Pictographic Character Reconstruction with Bézier Curves

Zihao Wan, Pau Tong Lin Xu, Fuwen Luo +3

While Vision-language Models (VLMs) have demonstrated strong semantic capabilities, their ability to interpret the underlying geometric structure of visual information is less expl…

cs.AI2025

Patient-Zero: Scaling Synthetic Patient Agents to Real-World Distributions without Real Patient Data

Yunghwei Lai, Ziyue Wang, Weizhi Ma +1

Synthetic data generation with Large Language Models (LLMs) has emerged as a promising solution in the medical domain to mitigate data scarcity and privacy constraints. However, ex…

cs.CL2025

MUCAR: Benchmarking Multilingual Cross-Modal Ambiguity Resolution for Multimodal Large Language Models

Xiaolong Wang, Zhaolu Kang, Wangyuxuan Zhai +8

Multimodal Large Language Models (MLLMs) have demonstrated significant advances across numerous vision-language tasks. MLLMs have shown promising capability in aligning visual and…

cs.AI2025

Agent-Environment Alignment via Automated Interface Generation

Kaiming Liu, Xuanyu Lei, Ziyue Wang +2

Large language model (LLM) agents have shown impressive reasoning capabilities in interactive decision-making tasks. These agents interact with environment through intermediate int…

cs.CL2025

Thinking with Visual Abstract: Enhancing Multimodal Reasoning via Visual Abstraction

Dairu Liu, Ziyue Wang, Minyuan Ruan +4

Images usually convey richer detail than text, but often include redundant information, which potentially downgrades multimodal reasoning performance. When faced with lengthy or co…

cs.CV2025

CoSpace: Benchmarking Continuous Space Perception Ability for Vision-Language Models

Yiqi Zhu, Ziyue Wang, Can Zhang +2

Vision-Language Models (VLMs) have recently witnessed significant progress in visual comprehension. As the permitting length of image context grows, VLMs can now comprehend a broad…