4 papers · 1 filter
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling
Peng Kuang, Haibo Jin, Xiaoyu Han +5
Process Reward Models (PRMs) have been proven to be highly effective in guiding test-time scaling (TTS) methods, which significantly boost the capabilities of LLM-based multi-agent…
Closing the Loop on Latent Reasoning via Test-Time Reconstruction
Xiaopeng Yuan, Haibo Jin, Ye Yu +4
Recent work moves intermediate reasoning from natural-language traces into latent or cache-level representations to reduce token overhead and avoid a discrete communication bottlen…
Learning to Communicate: Toward End-to-End Optimization of Multi-Agent Language Systems
Ye Yu, Heming Liu, Haibo Jin +3
Multi-agent systems built on large language models have shown strong performance on complex reasoning tasks, yet most work focuses on agent roles and orchestration while treating i…
TIM-PRM: Verifying multimodal reasoning with Tool-Integrated PRM
Peng Kuang, Xiangxiang Wang, Wentao Liu +2
Multimodal Large Language Models (MLLMs) have achieved impressive performances in mathematical reasoning, yet they remain vulnerable to visual hallucinations and logical inconsiste…