From the 3 of 15 linked papers with an AI index.
15 papers
MAG: A Web-Agent Benchmark and Harness for Multimodal Action and Guide Generation
Chengguang Gan, Hanjun Wei, Yunhao Liang +3
The paper presents MAG, a benchmark and harness that combine web‑agent action execution and guide text generation into a single multimodal task using screenshot‑based grounding, an…
A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its Mechanism
Chengguang Gan, Zhixi Cai, Yunhao Liang +3
The paper evaluates whether Group Relative Policy Optimization (GRPO) improves the performance of small (4‑8 B parameter) language and vision‑language web agents and finds that it…
What We Talk About When We Talk About LLM Planning: Evidence for Two Distinct Planning Abilities
Sukai Huang, Chenyuan Zhang, Fucai Ke +4
The paper investigates whether large language models (LLMs) have distinct planning abilities by applying multidimensional item response theory to benchmark data, uncovering two sep…
Explain Before You Answer: A Survey on Compositional Visual Reasoning
Fucai Ke, Joy Hsu, Zhixi Cai +10
Compositional visual reasoning has emerged as a key research frontier in multimodal AI, aiming to endow machines with the human-like ability to decompose visual scenes, ground inte…
Xetrieval: Mechanistically Explaining Dense Retrieval
Zhixin Cai, Jun Bai, Yang Liu +7
Explaining why dense retrievers assign high relevance scores remains challenging because retrieval decisions are made through opaque high-dimensional embeddings. Existing explanati…
Mini-BEHAVIOR-Gran: Revealing U-Shaped Effects of Instruction Granularity on Language-Guided Embodied Agents
Sukai Huang, Chenyuan Zhang, Fucai Ke +4
Instruction granularity is an important yet poorly controlled variable in language-guided embodied AI. Existing benchmarks typically pair each task with a single static instruction…