7 papers
RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments
Zhiyuan Zeng, Hamish Ivison, Yiping Wang +14
We introduce Reinforcement Learning (RL) with Adaptive Verifiable Environments (RLVE), an approach using verifiable environments that procedurally generate problems and provide alg…
Dynamic Thinking-Token Selection for Efficient Reasoning in Large Reasoning Models
Zhenyuan Guo, Tong Chen, Wenlong Meng +4
Large Reasoning Models (LRMs) excel at solving complex problems by explicitly generating a reasoning trace before deriving the final answer. However, these extended generations inc…
StreamVLA: Breaking the Reason-Act Cycle via Completion-State Gating
Tongqing Chen, Hang Wu, Jiasen Wang +2
Long-horizon robotic manipulation requires bridging the gap between high-level planning (System 2) and low-level control (System 1). Current Vision-Language-Action (VLA) models oft…
PAMAS: Self-Adaptive Multi-Agent System with Perspective Aggregation for Misinformation Detection
Zongwei Wang, Min Gao, Junliang Yu +2
Misinformation on social media poses a critical threat to information credibility, as its diverse and context-dependent nature complicates detection. Large language model-empowered…
Trading-R1: Financial Trading with LLM Reasoning via Reinforcement Learning
Yijia Xiao, Edward Sun, Tong Chen +3
Developing professional, structured reasoning on par with human financial analysts and traders remains a central challenge in AI for finance, where markets demand interpretability…
Beyond Sequential Reranking: Reranker-Guided Search Improves Reasoning Intensive Retrieval
Haike Xu, Tong Chen
The widely used retrieve-and-rerank pipeline faces two critical limitations: they are constrained by the initial retrieval quality of the top-k documents, and the growing computati…