works on

From the 1 of 14 linked papers with an AI index.

activity
20242026
collaborators

17 papers

cs.CV2026

HyMobileAgent: Data-Environment Co-Scaling for Efficient GUI Agents

Hy Vision Team, Huawen Shen, Zhengyang Tang +20

The paper introduces HyMobileAgent, a vision-native mobile GUI agent that combines large multimodal models with a co-scaling framework for data and environments to enable precise p…

cs.CV2026

HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better

Gengluo Li, Xingyu Wan, Shangpin Peng +20

We present HunyuanOCR-1.5, a lightweight end-to-end OCR-specialized vision-language model. HunyuanOCR unifies document parsing, text spotting, information extraction, text-image tr…

cs.CV2026

ChartArena: Benchmarking Chart Parsing across Languages, Scenarios, and Formats

Shangpin Peng, Gengluo Li, Xingyu Wan +10

Charts are a primary medium for conveying quantitative and relational information, yet systematically evaluating chart parsing models remains difficult. Existing benchmarks focus o…

cs.CV2026

StrucTab: A Structured Optimization Framework for Table Parsing

Gengluo Li, Shangpin Peng, Chengquan Zhang +10

Table parsing aims to convert table images into structured, machine-readable representations, a task requiring the joint perception of complex spatial layouts and textual content.…

cs.AI2026

Does the Question Really Matter? Training-Free Data Selection for Vision-Language SFT

Peng Sun, Yi Yang, Huawen Shen +4

Visual instruction tuning is crucial for improving vision-language large models (VLLMs). However, many samples can be solved via linguistic patterns or common-sense shortcuts, with…

cs.AI2026

DRS-GUI: Dynamic Region Search for Training-Free GUI Grounding

Yichao Liu, Huawen Shen, Liu Yu +3

GUI agents powered by Multimodal Large Language Models (MLLMs) have demonstrated impressive capability in understanding and executing user instructions. However, accurately groundi…