collaborators

7 papers

cs.LG2026

Blockwise Policy-Drift Gating for On-Policy Distillation

Liwen Zheng, Haiyun Jiang

On-policy distillation (OPD) trains a student policy using teacher signals computed on trajectories sampled by the student itself. Recent work shows that sampled-token OPD can be f…

cs.CV2026

Vision-EKIPL: External Knowledge-Infused Policy Learning for Visual Reasoning

Chaoyang Wang, Zeyu Zhang, Meng Meng +2

Visual reasoning is crucial for understanding complex multimodal data and advancing Artificial General Intelligence. Existing methods enhance the reasoning capability of Multimodal…

cs.CV2026

Where and How to Prune: An Empirical Study of Visual Token Pruning for GUI Agent Navigation

Daiqiang Li, Zihao Pan, Zeyu Zhang +8

In recent years, GUI agents have demonstrated strong potential in navigation tasks. However, preserving complete historical screenshots introduces substantial computational overhea…

cs.AI2026

MaLoRA: Gated Modality LoRA for Key-Space Alignment in Multimodal LLM Fine-Tuning

Xinhan Zheng, Huyu Wu, Xueting Wang +2

Multimodal large language models (MLLMs) exhibit a pronounced preference for textual inputs when processing vision-language data, limiting their ability to reason effectively from…

cs.IR2026

ARK: Answer-Centric Retriever Tuning via KG-augmented Curriculum Learning

Hang Ding, Jiawei Zhou, Haiyun Jiang

Retrieval-Augmented Generation (RAG) has emerged as a powerful framework for knowledge-intensive tasks, yet its effectiveness in long-context scenarios is often bottlenecked by the…

cs.RO2025

AerialMind: Towards Referring Multi-Object Tracking in UAV Scenarios

Chenglizhao Chen, Shaofeng Liang, Runwei Guan +6

Referring Multi-Object Tracking (RMOT) aims to achieve precise object detection and tracking through natural language instructions, representing a fundamental capability for intell…