7 papers
Blockwise Policy-Drift Gating for On-Policy Distillation
Liwen Zheng, Haiyun Jiang
On-policy distillation (OPD) trains a student policy using teacher signals computed on trajectories sampled by the student itself. Recent work shows that sampled-token OPD can be f…
Vision-EKIPL: External Knowledge-Infused Policy Learning for Visual Reasoning
Chaoyang Wang, Zeyu Zhang, Meng Meng +2
Visual reasoning is crucial for understanding complex multimodal data and advancing Artificial General Intelligence. Existing methods enhance the reasoning capability of Multimodal…
Where and How to Prune: An Empirical Study of Visual Token Pruning for GUI Agent Navigation
Daiqiang Li, Zihao Pan, Zeyu Zhang +8
In recent years, GUI agents have demonstrated strong potential in navigation tasks. However, preserving complete historical screenshots introduces substantial computational overhea…
MaLoRA: Gated Modality LoRA for Key-Space Alignment in Multimodal LLM Fine-Tuning
Xinhan Zheng, Huyu Wu, Xueting Wang +2
Multimodal large language models (MLLMs) exhibit a pronounced preference for textual inputs when processing vision-language data, limiting their ability to reason effectively from…
ARK: Answer-Centric Retriever Tuning via KG-augmented Curriculum Learning
Hang Ding, Jiawei Zhou, Haiyun Jiang
Retrieval-Augmented Generation (RAG) has emerged as a powerful framework for knowledge-intensive tasks, yet its effectiveness in long-context scenarios is often bottlenecked by the…
AerialMind: Towards Referring Multi-Object Tracking in UAV Scenarios
Chenglizhao Chen, Shaofeng Liang, Runwei Guan +6
Referring Multi-Object Tracking (RMOT) aims to achieve precise object detection and tracking through natural language instructions, representing a fundamental capability for intell…