From the 1 of 10 linked papers with an AI index.
10 papers
OvisOCR2 Technical Report
Shiyin Lu, Yinglun Li, Yu Xia +10
OvisOCR2 is a 0.8 B parameter end‑to‑end model that converts document page images into Markdown, handling text, formulas, tables, and visual regions, and achieves state‑of‑the‑art…
SnapAudit: Active Auditing of Differentially Private In-Context Learning via Snapshot-Based Simulation
Yuyang Xia, Ruixuan Liu, Li Xiong
In-context learning (ICL) allows LLMs to adapt to new tasks via a few demonstrations, but those demonstrations may contain sensitive data. Differentially private (DP) ICL mechanism…
LPO: Towards Accurate GUI Agent Interaction via Location Preference Optimization
Jiaqi Tang, Yu Xia, Yi-Feng Wu +9
The advent of autonomous agents is transforming interactions with Graphical User Interfaces (GUIs) by employing natural language as a powerful intermediary. Despite the predominanc…
NL2SpaTiaL: Generating Geometric Spatio-Temporal Logic Specifications from Natural Language for Manipulation Tasks
Licheng Luo, Kaier Liang, Yu Xia +1
While Temporal Logic provides a rigorous verification framework for robotics, it typically operates on trajectory-level signals and does not natively represent the object-centric g…
MSRAMIE: Multimodal Structured Reasoning Agent for Multi-instruction Image Editing
Zhaoyuan Qiu, Ken Chen, Xiangwei Wang +3
Existing instruction-based image editing models perform well with simple, single-step instructions but degrade in realistic scenarios that involve multiple, lengthy, and interdepen…
Building Autonomous GUI Navigation via Agentic-Q Estimation and Step-Wise Policy Optimization
Yibo Wang, Guangda Huzhang, Yuwei Hu +7
Recent advances in Multimodal Large Language Models (MLLMs) have substantially driven the progress of autonomous agents for Graphical User Interface (GUI). Nevertheless, in real-wo…