activity
20242026
most citedScaling over Scaling: Exploring Test-Time Scaling Plateau in Large Reasoning Models

1 citations · 1 across the 6 of their papers we have counts for

collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL2026

Seeing Isn't Believing: Mitigating Belief Inertia via Active Intervention in Embodied Agents

Hanlin Wang, Chak Tou Leong, Jian Wang +1

Recent advancements in large language models (LLMs) have enabled agents to tackle complex embodied tasks through environmental interaction. However, these agents still make subopti…

cs.CL2026

Foresight Optimization for Strategic Reasoning in Large Language Models

Jiashuo Wang, Jiawen Duan, Jian Wang +6

Reasoning capabilities in large language models (LLMs) have generally advanced significantly. However, it is still challenging for existing reasoning-based LLMs to perform effectiv…

cs.CL2025

SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution

Hanlin Wang, Chak Tou Leong, Jiashuo Wang +2

Reinforcement learning (RL) holds significant promise for training LLM agents to handle complex, goal-oriented tasks that require multi-step interactions with external environments…

cs.CL2025

Towards Harmless Multimodal Assistants with Blind Preference Optimization

Yongqi Li, Lu Yang, Jian Wang +3

Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in multimodal understanding, reasoning, and interaction. Given the extensive applications of MLLM…

cs.CL2025

Integrating Chain-of-Thought for Multimodal Alignment: A Study on 3D Vision-Language Learning

Yanjun Chen, Yirong Sun, Xinghao Chen +4

Chain-of-Thought (CoT) reasoning has proven effective in natural language tasks but remains underexplored in multimodal alignment. This study investigates its integration into 3D v…

cs.CL2025

Why Safeguarded Ships Run Aground? Aligned Large Language Models' Safety Mechanisms Tend to Be Anchored in The Template Region

Chak Tou Leong, Qingyu Yin, Jian Wang +1

The safety alignment of large language models (LLMs) remains vulnerable, as their initial behavior can be easily jailbroken by even relatively simple attacks. Since infilling a fix…