1 citations · 1 across the 4 of their papers we have counts for
9 papers
Finding RELIEF: Shaping Reasoning Behavior without Reasoning Supervision via Belief Engineering
Chak Tou Leong, Dingwei Chen, Heming Xia +4
Large reasoning models (LRMs) have achieved remarkable success in complex problem-solving, yet they often suffer from computational redundancy or reasoning unfaithfulness. Current…
Scaling over Scaling: Exploring Test-Time Scaling Plateau in Large Reasoning Models
Jian Wang, Boyan Zhu, Chak Tou Leong +2
Large reasoning models (LRMs) have exhibited the capacity of enhancing reasoning performance via internal test-time scaling. Building upon this, a promising direction is to further…
SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Hanlin Wang, Chak Tou Leong, Jiashuo Wang +2
Reinforcement learning (RL) holds significant promise for training LLM agents to handle complex, goal-oriented tasks that require multi-step interactions with external environments…
Towards Harmless Multimodal Assistants with Blind Preference Optimization
Yongqi Li, Lu Yang, Jian Wang +3
Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in multimodal understanding, reasoning, and interaction. Given the extensive applications of MLLM…
Integrating Chain-of-Thought for Multimodal Alignment: A Study on 3D Vision-Language Learning
Yanjun Chen, Yirong Sun, Xinghao Chen +4
Chain-of-Thought (CoT) reasoning has proven effective in natural language tasks but remains underexplored in multimodal alignment. This study investigates its integration into 3D v…
STeCa: Step-level Trajectory Calibration for LLM Agent Learning
Hanlin Wang, Jian Wang, Chak Tou Leong +1
Large language model (LLM)-based agents have shown promise in tackling complex tasks by interacting dynamically with the environment. Existing work primarily focuses on behavior cl…