10 citations · 12 across the 11 of their papers we have counts for
4 papers · 1 filter
How Output Format Confounds Data Quality and Capability in Instruction Tuning
Chengguang Gan, Hanjun Wei, Yunhao Liang +3
Instruction-tuning data are judged by quality metrics, and tuned models are judged by benchmarks, but both judgments pass through an output interface: the surface format in which a…
Joint Training Is Not Enough: Conditioned Cross-Granularity Training for Multimodal Document Understanding
Chengguang Gan, Yunhao Liang, Hanjun Wei +2
The Mutual Reinforcement Effect (MRE) asks whether a fine, span-level and a coarse, document-level task help each other when one model handles both. We test it in multimodal docume…
MAG: A Web-Agent Benchmark and Harness for Multimodal Action and Guide Generation
Chengguang Gan, Hanjun Wei, Yunhao Liang +3
Digital Adoption Platforms (DAPs) are embedded overlays widely used on web systems to guide users through operations inside a page, helping them get started with unfamiliar interfa…
A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its Mechanism
Chengguang Gan, Zhixi Cai, Yunhao Liang +3
Reinforcement learning with verifiable rewards, and Group Relative Policy Optimization (GRPO) in particular, is now run routinely on a supervised checkpoint in the hope of producin…