4 papers · 1 filter
SLPO: Scaling Latent Reasoning via a Surrogate Policy
Runyang You, Zhiyuan Liu, Yongqi Li +1
Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thought reasoners. Yet this scaling path remai…
Parallel Test-Time Scaling for Latent Reasoning Models
Runyang You, Yongqi Li, Meng Liu +3
Parallel test-time scaling (TTS) is a pivotal approach for enhancing large language models (LLMs), typically by sampling multiple token-based chains-of-thought in parallel and aggr…
Agent-as-a-Judge
Runyang You, Hongru Cai, Caiqi Zhang +5
LLM-as-a-Judge has revolutionized AI evaluation by leveraging large language models for scalable assessments. However, as evaluands become increasingly complex, specialized, and mu…
Towards Harmless Multimodal Assistants with Blind Preference Optimization
Yongqi Li, Lu Yang, Jian Wang +3
Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in multimodal understanding, reasoning, and interaction. Given the extensive applications of MLLM…