26 citations · 26 across the 6 of their papers we have counts for
10 papers
Dolphin-v2: Universal Document Parsing via Scalable Anchor Prompting
Hao Feng, Wei Shi, Ke Zhang +9
Document parsing has garnered widespread attention as vision-language models (VLMs) advance OCR capabilities. However, the field remains fragmented across dozens of specialized mod…
Complex Mathematical Expression Recognition: Benchmark, Large-Scale Dataset and Strong Baseline
Weikang Bai, Yongkun Du, Yuchen Su +2
Mathematical Expression Recognition (MER) has made significant progress in recognizing simple expressions, but the robust recognition of complex mathematical expressions with many…
MDiff4STR: Mask Diffusion Model for Scene Text Recognition
Yongkun Du, Miaomiao Zhao, Songlin Fan +3
Mask Diffusion Models (MDMs) have recently emerged as a promising alternative to auto-regressive models (ARMs) for vision-language tasks, owing to their flexible balance of efficie…
DocPTBench: Benchmarking End-to-End Photographed Document Parsing and Translation
Yongkun Du, Pinxuan Chen, Xuye Ying +1
The advent of Multimodal Large Language Models (MLLMs) has unlocked the potential for end-to-end document parsing and translation. However, prevailing benchmarks such as OmniDocBen…
LRANet++: Low-Rank Approximation Network for Accurate and Efficient Text Spotting
Yuchen Su, Zhineng Chen, Yongkun Du +3
End-to-end text spotting aims to jointly optimize text detection and recognition within a unified framework. Despite significant progress, designing an accurate and efficient end-t…
Test-Time Reinforcement Learning for GUI Grounding via Region Consistency
Yong Du, Yuchen Yan, Fei Tang +5
Graphical User Interface (GUI) grounding, the task of mapping natural language instructions to precise screen coordinates, is fundamental to autonomous GUI agents. While existing m…