collaborators

6 papers

cs.AI2025

Scaling Spatial Reasoning in MLLMs through Programmatic Data Synthesis

Zhi Helu, Huang Jingjing, Xu Wang +9

Embodied intelligence, a grand challenge in artificial intelligence, is fundamentally constrained by the limited spatial understanding and reasoning capabilities of current models.…

cs.CL2025

LaSeR: Reinforcement Learning with Last-Token Self-Rewarding

Wenkai Yang, Weijie Liu, Ruobing Xie +4

Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as a core paradigm for enhancing the reasoning capabilities of Large Language Models (LLMs). To address t…

cs.CL2025

Beyond the Surface: Measuring Self-Preference in LLM Judgments

Zhi-Yuan Chen, Hao Wang, Xinyu Zhang +2

Recent studies show that large language models (LLMs) exhibit self-preference bias when serving as judges, meaning they tend to favor their own responses over those generated by ot…

cs.CL2025

Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning

Wenkai Yang, Shuming Ma, Yankai Lin +1

Recent studies have shown that making a model spend more time thinking through longer Chain of Thoughts (CoTs) enables it to gain significant improvements in complex reasoning task…

cs.LG2025

Revisiting Weak-to-Strong Generalization in Theory and Practice: Reverse KL vs. Forward KL

Wei Yao, Wenkai Yang, Ziqiao Wang +2

As large language models advance toward superhuman performance, ensuring their alignment with human values and abilities grows increasingly complex. Weak-to-strong generalization o…

cs.LG2025

The Capabilities and Limitations of Weak-to-Strong Generalization: Generalization and Calibration

Wei Yao, Wenkai Yang, Gengze Xu +3

Weak-to-strong generalization, where weakly supervised strong models outperform their weaker teachers, offers a promising approach to aligning superhuman models with human values.…