From the 1 of 22 linked papers with an AI index.
22 papers
Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs
Xinming Wang, Weinong Wang, Hongming Yang +13
Hybrid-thinking multimodal large language models (MLLMs) allow a single model to alternate between deliberative thinking and latency-efficient non-thinking inference. Although thes…
HyMobileAgent: Data-Environment Co-Scaling for Efficient GUI Agents
Hy Vision Team, Huawen Shen, Zhengyang Tang +20
The paper introduces HyMobileAgent, a vision-native mobile GUI agent that combines large multimodal models with a co-scaling framework for data and environments to enable precise p…
HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better
Gengluo Li, Xingyu Wan, Shangpin Peng +20
We present HunyuanOCR-1.5, a lightweight end-to-end OCR-specialized vision-language model. HunyuanOCR unifies document parsing, text spotting, information extraction, text-image tr…
Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe
Wenjin Hou, Shangpin Peng, Weinong Wang +13
On-policy distillation (OPD) has recently emerged as an effective post-training paradigm for consolidating the capabilities of specialized expert models into a single student model…
ChartArena: Benchmarking Chart Parsing across Languages, Scenarios, and Formats
Shangpin Peng, Gengluo Li, Xingyu Wan +10
Charts are a primary medium for conveying quantitative and relational information, yet systematically evaluating chart parsing models remains difficult. Existing benchmarks focus o…
StrucTab: A Structured Optimization Framework for Table Parsing
Gengluo Li, Shangpin Peng, Chengquan Zhang +10
Table parsing aims to convert table images into structured, machine-readable representations, a task requiring the joint perception of complex spatial layouts and textual content.…