collaborators

5 papers

cs.LG2025

Critique to Verify: Accurate and Honest Test-Time Scaling with RL-Trained Verifiers

Zhicheng Yang, Zhijiang Guo, Yinya Huang +4

Test-time scaling via solution sampling and aggregation has become a key paradigm for improving the reasoning performance of Large Language Models (LLMs). While reward model select…

cs.LG2025

When Inverse Data Outperforms: Exploring the Pitfalls of Mixed Data in Multi-Stage Fine-Tuning

Mengyi Deng, Xin Li, Tingyu Zhu +3

Existing work has shown that o1-level performance can be achieved with limited data distillation, but most existing methods focus on unidirectional supervised fine-tuning (SFT), ov…

cs.CL2025

Understanding GUI Agent Localization Biases through Logit Sharpness

Xingjian Tao, Yiwei Wang, Yujun Cai +2

Multimodal large language models (MLLMs) have enabled GUI agents to interact with operating systems by grounding language into spatial actions. Despite their promising performance,…

cs.LG2025

TreeRPO: Tree Relative Policy Optimization

Zhicheng Yang, Zhijiang Guo, Yinya Huang +3

Large Language Models (LLMs) have shown remarkable reasoning capabilities through Reinforcement Learning with Verifiable Rewards (RLVR) methods. However, a key limitation of existi…

cs.CL2024

Are LLMs Really Not Knowledgeable? Mining the Submerged Knowledge in LLMs' Memory

Xingjian Tao, Yiwei Wang, Yujun Cai +2

Large language models (LLMs) have shown promise as parametric knowledge bases, but often underperform on question answering (QA) tasks due to hallucinations and uncertainty. While…