benchmark 1benchmark evaluation 1code language models 1group relative policy optimization 1guide generation 1language models 1learning rate effects 1model scaling 1multimodal learning 1prompt engineering 1representation analysis 1screenshot grounding 1
From the 3 of 7 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
MAG: A Web-Agent Benchmark and Harness for Multimodal Action and Guide Generation
Chengguang Gan, Hanjun Wei, Yunhao Liang +3
The paper presents MAG, a benchmark and harness that combine web‑agent action execution and guide text generation into a single multimodal task using screenshot‑based grounding, an…
cs.AI2026
A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its Mechanism
Chengguang Gan, Zhixi Cai, Yunhao Liang +3
The paper evaluates whether Group Relative Policy Optimization (GRPO) improves the performance of small (4‑8 B parameter) language and vision‑language web agents and finds that it…