works on

From the 1 of 11 linked papers with an AI index.

collaborators

11 papers

cs.LG2026

Consistency-Driven Co-Evolution for Self-Supervised Cross-Representation Learning

Xuehang Guo, Pengyuan Li, Tom Hope +3

As chart images, tabular data, and visualization code play increasingly important roles across diverse domains, cross-representation understanding across these modalities poses fun…

cs.CV2026

CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning

Xuehang Guo, Pingyue Zhang, Ruiyi Zhang +6

Chart question answering (CQA) requires multimodal large language models (MLLMs) to integrate visual comprehension with logical reasoning, yet current models struggle with accurate…

cs.CV2026

Long-Horizon Embodied Decision-Making via Multimodal Memory Compression

Bingxuan Li, Rui Yang, Cheng Qian +6

Agents are increasingly expected to act not only as task executors, but also as decision-makers on behalf of human users. This shift requires agents to accumulate evidence over lon…

cs.CV2026

HumanCLAW: Can Vision-Language Models Act Through a Body?

Siyao Li, Li Siyao, Jiawei Gu +16

The paper introduces HumanCLAW, a framework that separates decision making of vision‑language models from low‑level motor execution, allowing evaluation of a model's action intelli…

cs.CL2026

SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents

Ziyi Wang, Yuxuan Lu, Yimeng Zhang +8

Language model agents are increasingly effective in solving realistic tasks through multi-turn tool use. However, training reliable tool-using agents remains challenging in practic…

cs.AI2026

Brick-Composer: Using MLLMs for Assembly with Diverse Bricks

Jiateng Liu, Bingxuan Li, Zhenhailong Wang +8

We dream of AI agents that can read arbitrary designs and construct real-world objects from reusable building blocks. As a first step toward this vision, we study whether multimoda…