collaborators

8 papers

cs.CV2025

ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning

Shengyuan Ding, Xinyu Fang, Ziyu Liu +10

Reward models are critical for aligning vision-language systems with human preferences, yet current approaches suffer from hallucination, weak visual grounding, and an inability to…

cs.AI2025

NP-Engine: Empowering Optimization Reasoning in Large Language Models with Verifiable Synthetic NP Problems

Xiaozhe Li, Xinyu Fang, Shengyuan Ding +4

Large Language Models (LLMs) have shown strong reasoning capabilities, with models like OpenAI's O-series and DeepSeek R1 excelling at tasks such as mathematics, coding, logic, and…

cs.CV2025

SPARK: Synergistic Policy And Reward Co-Evolving Framework

Ziyu Liu, Yuhang Zang, Shengyuan Ding +5

Recent Large Language Models (LLMs) and Large Vision-Language Models (LVLMs) increasingly use Reinforcement Learning (RL) for post-pretraining, such as RL with Verifiable Rewards (…

cs.AI2025

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems

Xiaozhe Li, Jixuan Chen, Xinyu Fang +4

Large Language Models (LLMs) have shown remarkable capabilities in solving diverse tasks. However, their proficiency in iteratively optimizing complex solutions through learning fr…

cs.CV2025

MM-IFEngine: Towards Multimodal Instruction Following

Shengyuan Ding, Shenxi Wu, Xiangyu Zhao +7

The Instruction Following (IF) ability measures how well Multi-modal Large Language Models (MLLMs) understand exactly what users are telling them and whether they are doing it righ…

cs.CV2025

Creation-MMBench: Assessing Context-Aware Creative Intelligence in MLLM

Xinyu Fang, Zhijian Chen, Kai Lan +10

Creativity is a fundamental aspect of intelligence, involving the ability to generate novel and appropriate solutions across diverse contexts. While Large Language Models (LLMs) ha…