collaborators

7 papers

cs.CV2026

DocPO: Advancing Document Policy Optimization via Tailored Step-Aware Rewards

Yunhao Wang, Binghong Wu, Zhenyu Huang +4

Reinforcement learning (RL) for document parsing often relies on reference-based rewards rooted in edit distance (e.g., tree edit distance), yet it remains hard to optimize in the…

cs.CL2026

Deep Research Pretraining via Predictive Navigation

Jiang Zhou, Zhiyuan Fan, Xing Wu +3

Deep research agents are often trained on expensive, environment-grounded tool-use trajectories that require repeated retrieval, document inspection, and report evaluation. We intr…

cs.CV2026

MORE: A Multilingual Document Parsing Benchmark and Evaluation

Long Xu, Binghong Wu, Tinghao Yu +6

Multilingual documents encapsulate rich regional cultures, scientific discoveries, and historical records. Parsing this content into structured, machine-readable formats is critica…

cs.AI2026

Toward Scalable Terminal Task Synthesis via Skill Graphs

Zhiyuan Fan, Tinghao Yu, Yuanjun Cai +8

Terminal agents have demonstrated strong potential for autonomous command-line execution, yet their training remains constrained by the scarcity of high-quality and diverse executi…

cs.LG2026

PolicyLong: Towards On-Policy Context Extension

Junlong Jia, Ziyang Chen, Xing Wu +4

Extending LLM context windows is hindered by scarce high-quality long-context data. Recent methods synthesize data with genuine long-range dependencies via information-theoretic ve…

cs.CL2025

Hunyuan-TurboS: Advancing Large Language Models through Mamba-Transformer Synergy and Adaptive Chain-of-Thought

Tencent Hunyuan Team, Ao Liu, Botong Zhou +248

As Large Language Models (LLMs) rapidly advance, we introduce Hunyuan-TurboS, a novel large hybrid Transformer-Mamba Mixture of Experts (MoE) model. It synergistically combines Mam…