activity
20242026
collaborators
Showing 2025Show all

7 papers · 1 filter

cs.CV2025

Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing

Baode Wang, Biao Wu, Weizhen Li +8

Automated parsing of scanned documents into richly structured, machine-readable formats remains a critical bottleneck in Document AI, as traditional multi-stage pipelines suffer fr…

cs.CL2025

Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing

Baode Wang, Biao Wu, Weizhen Li +8

Document parsing from scanned images into structured formats remains a significant challenge due to its complexly intertwined elements such as text paragraphs, figures, formulas, a…

cs.CV2025

UniVid: The Open-Source Unified Video Model

Jiabin Luo, Junhui Lin, Zeyu Zhang +4

Unified video modeling that combines generation and understanding capabilities is increasingly important but faces two key challenges: maintaining semantic faithfulness during flow…

cs.RO2025

Automotive-ENV: Benchmarking Multimodal Agents in Vehicle Interface Systems

Junfeng Yan, Biao Wu, Meng Fang +1

Multimodal agents have demonstrated strong performance in general GUI interactions, but their application in automotive systems has been largely unexplored. In-vehicle GUIs present…

cs.AI2025

Foundations and Recent Trends in Multimodal Mobile Agents: A Survey

Biao Wu, Yanda Li, Zhiwei Zhang +3

Mobile agents are essential for automating tasks in complex and dynamic mobile environments. As foundation models evolve, the demands for agents that can adapt in real-time and pro…

cs.IR2025

WebArXiv: Evaluating Multimodal Agents on Time-Invariant arXiv Tasks

Zihao Sun, Ling Chen

Recent progress in large language models (LLMs) has enabled the development of autonomous web agents capable of navigating and interacting with real websites. However, evaluating s…