4 papers
ECHO-2: A Large-Scale Distributed Rollout Framework for Cost-Efficient Reinforcement Learning
Jingwei Song, Meng Chen, Jie Xiao +15
Reinforcement learning (RL) is a critical stage in post-training large language models (LLMs), involving repeated interaction between rollout generation, reward evaluation, and cen…
Data Verification is the Future of Quantum Computing Copilots
Junhao Song, Ziqian Bi, Xinliang Chia +2
Quantum program generation demands a level of precision that may not be compatible with the statistical reasoning carried out in the inference of large language models (LLMs). Hall…
PRIME: Policy-Reinforced Iterative Multi-agent Execution for Algorithmic Reasoning in Large Language Models
Jiawei Xu, Zhenyu Yu, Ziqian Bi +3
Large language models have demonstrated remarkable capabilities across diverse reasoning tasks, yet their performance on algorithmic reasoning remains limited. To handle this limit…
Adaptive Detector-Verifier Framework for Zero-Shot Polyp Detection in Open-World Settings
Shengkai Xu, Hsiang Lun Kao, Tianxiang Xu +9
Polyp detectors trained on clean datasets often underperform in real-world endoscopy, where illumination changes, motion blur, and occlusions degrade image quality. Existing approa…