5 citations · 6 across the 8 of their papers we have counts for
4 papers · 1 filter
HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better
Gengluo Li, Xingyu Wan, Shangpin Peng +20
We present HunyuanOCR-1.5, a lightweight end-to-end OCR-specialized vision-language model. HunyuanOCR unifies document parsing, text spotting, information extraction, text-image tr…
E-VAds: An E-commerce Short Videos Understanding Benchmark for MLLMs
Xianjie Liu, Yiman Hu, Liang Wu +4
E-commerce short videos represent a high-revenue segment of the online video industry characterized by a goal-driven format and dense multi-modal signals. Current models often stru…
HunyuanOCR Technical Report
Hunyuan Vision Team, Pengyuan Lyu, Xingyu Wan +29
This paper presents HunyuanOCR, a commercial-grade, open-source, and lightweight (1B parameters) Vision-Language Model (VLM) dedicated to OCR tasks. The architecture comprises a Na…
HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling
Xianjie Liu, Yiman Hu, Yixiong Zou +3
Multimodal Large Language Models (MLLMs) have made significant strides in visual understanding tasks. However, their performance on high-resolution images remains suboptimal. While…