11 papers
Balancing Efficiency and Efficacy: Training-Free Attention-Guided Switching Between Explicit and Latent Thoughts for MLLMs
Haoqian Kang, Liupeng Li, Kuofeng Gao +5
Reasoning in Multimodal Large Language Models (MLLMs) requires both fine-grained visual perception and rigorous logical deduction. Explicit text-based Chain-of-Thought (CoT) is com…
UI-MOPD: Multi-Platform On-Policy Distillation for Unified GUI Agents
Niu Lian, Tongbo Chen, Alan Chen +10
Recent advances in multimodal foundation models and agent systems have driven GUI agents from single-platform task execution toward cross-platform interaction. However, unified mul…
Unlocking Proactivity in Task-Oriented Dialogue
Azure Zhang, Ning Gao, Yuqin Dai +7
Proactive task-oriented dialogue (TOD), such as outbound sales, demands a persuasive agent that actively probes the user's concerns and steers the conversation toward acceptance wi…
CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception
Liupeng Li, Haoqian Kang, Zhenyu Lu +4
High-resolution (HR) image perception presents a key bottleneck for multimodal large language models (MLLMs). While visual search offers a promising solution, existing methods stru…
SegCompass: Exploring Interpretable Alignment with Sparse Autoencoders for Enhanced Reasoning Segmentation
Zhenyu Lu, Liupeng Li, Jinpeng Wang +4
While large language models provide strong compositional reasoning, existing reasoning segmentation pipelines fail to transparently connect this reasoning to visual perception. Cur…
Tailoring Teaching to Aptitude: Direction-Adaptive Self-Distillation for LLM Reasoning
Hongbin Zhang, Chaozheng Wang, Kehai Chen +4
On-policy self-distillation (OPSD) is an emerging LLM post-training paradigm in which the model serves as its own teacher: conditioned on privileged information such as a reference…