2 papers
cs.AI2026
WeAgent-MMSearch: Native Text-Vision Interaction for Multimodal Search Agents
Zongkai Liu, Hui Zhang, Liqiang Niu +7
Multimodal search agents extend parametric knowledge with newly emerging and long-tail evidence from the open web. Yet many existing agentic search environments often expose retrie…
cs.CV2025
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models
Juntao Liu, Liqiang Niu, Wenchao Chen +2
Existing visual token compression methods for Multimodal Large Language Models (MLLMs) predominantly operate as post-encoder modules, limiting their potential for efficiency gains.…