activity
20242026
collaborators

7 papers

cs.LG2026

Self-Boosting Vision-Language Models with Noisy Student On-Policy Self-Distillation

Shuai Wang, Daoan Zhang, Zhe Tang +2

Post-training enables vision-language models (VLMs) to understand human instructions and perform various downstream tasks. Current post-training methods usually rely on human-annot…

cs.LG2026

Athena: Enhancing Multimodal Reasoning with Data-efficient Process Reward Models

Shuai Wang, Zhenhua Liu, Jiaheng Wei +3

We present Athena-PRM, a multimodal process reward model (PRM) designed to evaluate the reward score for each step in solving complex reasoning problems. Developing high-performanc…

cs.CL2026

AgenticMath: Enhancing LLM Reasoning via Agentic-based Math Data Generation

Xianyang Liu, Yilin Liu, Shuai Wang +5

The creation of high-quality datasets to improve Large Language Model (LLM) reasoning remains a significant challenge, as current methods often suffer from generating low-quality/i…

cs.CV2025

LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models

Shuai Wang, Daoan Zhang, Tianyi Bai +3

Humans can perceive and understand 3D space and long videos from sequential visual observations. But do vision-language models (VLMs) can? Recent work demonstrates that even state-…

cs.CV2025

FlexAC: Towards Flexible Control of Associative Reasoning in Multimodal Large Language Models

Shengming Yuan, Xinyu Lyu, Shuailong Wang +3

Multimodal large language models (MLLMs) face an inherent trade-off between faithfulness and creativity, as different tasks require varying degrees of associative reasoning. Howeve…

cs.CY2025

Seeing The Words: Evaluating AI-generated Biblical Art

Hidde Makimei, Shuai Wang, Willem van Peursen

The past years witnessed a significant amount of Artificial Intelligence (AI) tools that can generate images from texts. This triggers the discussion of whether AI can generate acc…