2 citations · 4 across the 14 of their papers we have counts for
18 papers
CircuitLens: Reasoning Circuits as Data Selection Signals for Reinforcement Learning with Verifiable Rewards
Zhuofan Chen, Ziqian Jiao, Yikai Cui +3
Reinforcement learning with verifiable rewards (RLVR) is sensitive to which problems a model trains on, yet existing selection criteria--difficulty filtering, hand-curation, reward…
How Output Format Confounds Data Quality and Capability in Instruction Tuning
Chengguang Gan, Hanjun Wei, Yunhao Liang +3
Instruction-tuning data are judged by quality metrics, and tuned models are judged by benchmarks, but both judgments pass through an output interface: the surface format in which a…
MAG: A Web-Agent Benchmark and Harness for Multimodal Action and Guide Generation
Chengguang Gan, Hanjun Wei, Yunhao Liang +3
Digital Adoption Platforms (DAPs) are embedded overlays widely used on web systems to guide users through operations inside a page, helping them get started with unfamiliar interfa…
A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its Mechanism
Chengguang Gan, Zhixi Cai, Yunhao Liang +3
Reinforcement learning with verifiable rewards, and Group Relative Policy Optimization (GRPO) in particular, is now run routinely on a supervised checkpoint in the hope of producin…
What We Talk About When We Talk About LLM Planning: Evidence for Two Distinct Planning Abilities
Sukai Huang, Chenyuan Zhang, Fucai Ke +4
When LLMs exhibit uneven performance across planning tasks, these gaps are often attributed to task difficulty. We argue that this explanation is incomplete, as task-level variatio…
Xetrieval: Mechanistically Explaining Dense Retrieval
Zhixin Cai, Jun Bai, Yang Liu +7
Explaining why dense retrievers assign high relevance scores remains challenging because retrieval decisions are made through opaque high-dimensional embeddings. Existing explanati…