10 papers
DocPO: Advancing Document Policy Optimization via Tailored Step-Aware Rewards
Yunhao Wang, Binghong Wu, Zhenyu Huang +4
Reinforcement learning (RL) for document parsing often relies on reference-based rewards rooted in edit distance (e.g., tree edit distance), yet it remains hard to optimize in the…
HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better
Gengluo Li, Xingyu Wan, Shangpin Peng +20
We present HunyuanOCR-1.5, a lightweight end-to-end OCR-specialized vision-language model. HunyuanOCR unifies document parsing, text spotting, information extraction, text-image tr…
ChartArena: Benchmarking Chart Parsing across Languages, Scenarios, and Formats
Shangpin Peng, Gengluo Li, Xingyu Wan +10
Charts are a primary medium for conveying quantitative and relational information, yet systematically evaluating chart parsing models remains difficult. Existing benchmarks focus o…
MORE: A Multilingual Document Parsing Benchmark and Evaluation
Long Xu, Binghong Wu, Tinghao Yu +6
Multilingual documents encapsulate rich regional cultures, scientific discoveries, and historical records. Parsing this content into structured, machine-readable formats is critica…
StrucTab: A Structured Optimization Framework for Table Parsing
Gengluo Li, Shangpin Peng, Chengquan Zhang +10
Table parsing aims to convert table images into structured, machine-readable representations, a task requiring the joint perception of complex spatial layouts and textual content.…
Hy-MT2: A Family of Fast, Efficient and Powerful Multilingual Translation Models in the Wild
Mao Zheng, Zheng Li, Tao Chen +10
Hy-MT2 is a family of fast-thinking multilingual translation models designed for complex real-world scenarios. It includes three model sizes: 1.8B, 7B, and 30B-A3B (MoE), all of wh…