collaborators

10 papers

cs.CV2026

OvisOCR2 Technical Report

Shiyin Lu, Yinglun Li, Yu Xia +10

OvisOCR2 is a 0.8 B parameter end‑to‑end model that converts document page images into Markdown, handling text, formulas, tables, and visual regions, and achieves state‑of‑the‑art…

cs.CR2026

SnapAudit: Active Auditing of Differentially Private In-Context Learning via Snapshot-Based Simulation

Yuyang Xia, Ruixuan Liu, Li Xiong

In-context learning (ICL) allows LLMs to adapt to new tasks via a few demonstrations, but those demonstrations may contain sensitive data. Differentially private (DP) ICL mechanism…

cs.LG2026

LPO: Towards Accurate GUI Agent Interaction via Location Preference Optimization

Jiaqi Tang, Yu Xia, Yi-Feng Wu +9

The advent of autonomous agents is transforming interactions with Graphical User Interfaces (GUIs) by employing natural language as a powerful intermediary. Despite the predominanc…

cs.RO2026

NL2SpaTiaL: Generating Geometric Spatio-Temporal Logic Specifications from Natural Language for Manipulation Tasks

Licheng Luo, Kaier Liang, Yu Xia +1

While Temporal Logic provides a rigorous verification framework for robotics, it typically operates on trajectory-level signals and does not natively represent the object-centric g…

cs.CV2026

MSRAMIE: Multimodal Structured Reasoning Agent for Multi-instruction Image Editing

Zhaoyuan Qiu, Ken Chen, Xiangwei Wang +3

Existing instruction-based image editing models perform well with simple, single-step instructions but degrade in realistic scenarios that involve multiple, lengthy, and interdepen…

cs.AI2026

Building Autonomous GUI Navigation via Agentic-Q Estimation and Step-Wise Policy Optimization

Yibo Wang, Guangda Huzhang, Yuwei Hu +7

Recent advances in Multimodal Large Language Models (MLLMs) have substantially driven the progress of autonomous agents for Graphical User Interface (GUI). Nevertheless, in real-wo…