7 papers
DocPO: Advancing Document Policy Optimization via Tailored Step-Aware Rewards
Yunhao Wang, Binghong Wu, Zhenyu Huang +4
Reinforcement learning (RL) for document parsing often relies on reference-based rewards rooted in edit distance (e.g., tree edit distance), yet it remains hard to optimize in the…
Deep Research Pretraining via Predictive Navigation
Jiang Zhou, Zhiyuan Fan, Xing Wu +3
Deep research agents are often trained on expensive, environment-grounded tool-use trajectories that require repeated retrieval, document inspection, and report evaluation. We intr…
MORE: A Multilingual Document Parsing Benchmark and Evaluation
Long Xu, Binghong Wu, Tinghao Yu +6
Multilingual documents encapsulate rich regional cultures, scientific discoveries, and historical records. Parsing this content into structured, machine-readable formats is critica…
Toward Scalable Terminal Task Synthesis via Skill Graphs
Zhiyuan Fan, Tinghao Yu, Yuanjun Cai +8
Terminal agents have demonstrated strong potential for autonomous command-line execution, yet their training remains constrained by the scarcity of high-quality and diverse executi…
PolicyLong: Towards On-Policy Context Extension
Junlong Jia, Ziyang Chen, Xing Wu +4
Extending LLM context windows is hindered by scarce high-quality long-context data. Recent methods synthesize data with genuine long-range dependencies via information-theoretic ve…
Hunyuan-TurboS: Advancing Large Language Models through Mamba-Transformer Synergy and Adaptive Chain-of-Thought
Tencent Hunyuan Team, Ao Liu, Botong Zhou +248
As Large Language Models (LLMs) rapidly advance, we introduce Hunyuan-TurboS, a novel large hybrid Transformer-Mamba Mixture of Experts (MoE) model. It synergistically combines Mam…