5 papers
Self-Boosting Vision-Language Models with Noisy Student On-Policy Self-Distillation
Shuai Wang, Daoan Zhang, Zhe Tang +2
Post-training enables vision-language models (VLMs) to understand human instructions and perform various downstream tasks. Current post-training methods usually rely on human-annot…
GeoFaith: A Spatio-Temporal Dual View of Faithful Chain-of-Thought
Weijiang Lv, Wentong Zhao, Jiayu Wang +3
Chain-of-Thought (CoT) reasoning has advanced large language models (LLMs), but outcome-based supervision leads to pervasive post-hoc rationalization, producing plausible yet unfai…
Athena: Enhancing Multimodal Reasoning with Data-efficient Process Reward Models
Shuai Wang, Zhenhua Liu, Jiaheng Wei +3
We present Athena-PRM, a multimodal process reward model (PRM) designed to evaluate the reward score for each step in solving complex reasoning problems. Developing high-performanc…
AgenticMath: Enhancing LLM Reasoning via Agentic-based Math Data Generation
Xianyang Liu, Yilin Liu, Shuai Wang +5
The creation of high-quality datasets to improve Large Language Model (LLM) reasoning remains a significant challenge, as current methods often suffer from generating low-quality/i…
LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
Shuai Wang, Daoan Zhang, Tianyi Bai +3
Humans can perceive and understand 3D space and long videos from sequential visual observations. But do vision-language models (VLMs) can? Recent work demonstrates that even state-…