activity
20242026
most citedScaling over Scaling: Exploring Test-Time Scaling Plateau in Large Reasoning Models

1 citations · 1 across the 4 of their papers we have counts for

collaborators

9 papers

cs.AI2026

Finding RELIEF: Shaping Reasoning Behavior without Reasoning Supervision via Belief Engineering

Chak Tou Leong, Dingwei Chen, Heming Xia +4

Large reasoning models (LRMs) have achieved remarkable success in complex problem-solving, yet they often suffer from computational redundancy or reasoning unfaithfulness. Current…

cs.AI20251 cited

Scaling over Scaling: Exploring Test-Time Scaling Plateau in Large Reasoning Models

Jian Wang, Boyan Zhu, Chak Tou Leong +2

Large reasoning models (LRMs) have exhibited the capacity of enhancing reasoning performance via internal test-time scaling. Building upon this, a promising direction is to further…

cs.CL2025

SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution

Hanlin Wang, Chak Tou Leong, Jiashuo Wang +2

Reinforcement learning (RL) holds significant promise for training LLM agents to handle complex, goal-oriented tasks that require multi-step interactions with external environments…

cs.CL2025

Towards Harmless Multimodal Assistants with Blind Preference Optimization

Yongqi Li, Lu Yang, Jian Wang +3

Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in multimodal understanding, reasoning, and interaction. Given the extensive applications of MLLM…

cs.CL2025

Integrating Chain-of-Thought for Multimodal Alignment: A Study on 3D Vision-Language Learning

Yanjun Chen, Yirong Sun, Xinghao Chen +4

Chain-of-Thought (CoT) reasoning has proven effective in natural language tasks but remains underexplored in multimodal alignment. This study investigates its integration into 3D v…

cs.LG2025

STeCa: Step-level Trajectory Calibration for LLM Agent Learning

Hanlin Wang, Jian Wang, Chak Tou Leong +1

Large language model (LLM)-based agents have shown promise in tackling complex tasks by interacting dynamically with the environment. Existing work primarily focuses on behavior cl…