2 papers
cs.AI2025
MarsRL: Advancing Multi-Agent Reasoning System via Reinforcement Learning with Agentic Pipeline Parallelism
Shulin Liu, Dong Du, Tao Yang +2
Recent progress in large language models (LLMs) has been propelled by reinforcement learning with verifiable rewards (RLVR) and test-time scaling. However, the limited output lengt…
cs.CL2025
Baichuan4-Finance Technical Report
Hanyu Zhang, Boyu Qiu, Yuhao Feng +6
Large language models (LLMs) have demonstrated strong capabilities in language understanding, generation, and reasoning, yet their potential in finance remains underexplored due to…