works on

From the 3 of 15 linked papers with an AI index.

collaborators

15 papers

cs.AI2026

MAG: A Web-Agent Benchmark and Harness for Multimodal Action and Guide Generation

Chengguang Gan, Hanjun Wei, Yunhao Liang +3

The paper presents MAG, a benchmark and harness that combine web‑agent action execution and guide text generation into a single multimodal task using screenshot‑based grounding, an…

cs.AI2026

A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its Mechanism

Chengguang Gan, Zhixi Cai, Yunhao Liang +3

The paper evaluates whether Group Relative Policy Optimization (GRPO) improves the performance of small (4‑8 B parameter) language and vision‑language web agents and finds that it…

cs.AI2026

What We Talk About When We Talk About LLM Planning: Evidence for Two Distinct Planning Abilities

Sukai Huang, Chenyuan Zhang, Fucai Ke +4

The paper investigates whether large language models (LLMs) have distinct planning abilities by applying multidimensional item response theory to benchmark data, uncovering two sep…

cs.CV2026

Explain Before You Answer: A Survey on Compositional Visual Reasoning

Fucai Ke, Joy Hsu, Zhixi Cai +10

Compositional visual reasoning has emerged as a key research frontier in multimodal AI, aiming to endow machines with the human-like ability to decompose visual scenes, ground inte…

cs.AI2026

Xetrieval: Mechanistically Explaining Dense Retrieval

Zhixin Cai, Jun Bai, Yang Liu +7

Explaining why dense retrievers assign high relevance scores remains challenging because retrieval decisions are made through opaque high-dimensional embeddings. Existing explanati…

cs.AI2026

Mini-BEHAVIOR-Gran: Revealing U-Shaped Effects of Instruction Granularity on Language-Guided Embodied Agents

Sukai Huang, Chenyuan Zhang, Fucai Ke +4

Instruction granularity is an important yet poorly controlled variable in language-guided embodied AI. Existing benchmarks typically pair each task with a single static instruction…