8 papers
The Next Screenshot Knows: Gated Hindsight Distillation for Mobile GUI Agents
Weiwei Li, Junzhuo Liu, Tong Chu +2
GUI agents are commonly trained offline from successful interaction trajectories. Standard training decomposes each trajectory into prefix-action pairs: the agent predicts an actio…
When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents
Junzhuo Liu, Weiwei Li, Jun Ling +1
Privileged on-policy distillation provides dense supervision for multi-turn agents by allowing a synchronized teacher to re-score the student's response at every turn with access t…
GOTS: Greedy Orthogonal Token Selection for High-Resolution Vision-Language Models
Jun Ling, Tao Huang, Junzhuo Liu +2
Modern vision-language models (VLMs) increasingly rely on dynamic or high-resolution visual encoding, producing thousands of visual tokens that substantially increase downstream la…
DOne: Decoupling Structure and Rendering for High-Fidelity Design-to-Code Generation
Xinhao Huang, Jinke Yu, Wenhao Xu +5
While Vision Language Models (VLMs) have shown promise in Design-to-Code generation, they suffer from a "holistic bottleneck-failing to reconcile high-level structural hierarchy wi…
Let Samples Speak: Mitigating Spurious Correlation by Exploiting the Clusterness of Samples
Weiwei Li, Junzhuo Liu, Yuanyuan Ren +3
Deep learning models are known to often learn features that spuriously correlate with the class label during training but are irrelevant to the prediction task. Existing methods ty…
GeoRef: Referring Expressions in Geometry via Task Formulation, Synthetic Supervision, and Reinforced MLLM-based Solutions
Bing Liu, Wenqiang Yv, Xuzheng Yang +6
AI-driven geometric problem solving is a complex vision-language task that requires accurate diagram interpretation, mathematical reasoning, and robust cross-modal grounding. A fou…