#multimodal large language models
38 papers · 1 filter
Beacon: Knowing When and How to Perform Agentic Visual Reasoning
Qixun Wang, Yang Shi, Letian Cheng +11
The paper introduces Beacon, an agentic visual reasoning system that learns when to invoke external tools and how to use them effectively, improving multimodal large language model…
Capturing Token Tendencies for Training-Free Token Pruning in Multimodal Large Language Models
Jie Ma, Zhike Qiu, Jie Gao +4
The paper introduces Trend-aware Pruning, a training‑free method that models the temporal dynamics of attention to selectively keep visual tokens that become important in deeper la…
PathView-Bench: Can Multimodal Large Language Models Achieve Fine-grained Multiscale Understanding of Pathology Images?
Zongyi Chen, Yu Liang, Jie Lin +1
The paper presents PathVU, a benchmark that tests multimodal large language models on fine-grained, multiscale visual understanding of pathology images using region- and slide-leve…
FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image Retrieval
Bohan Hou, Haoqiang Lin, Xuemeng Song +4
The paper introduces an automated pipeline to create a fine-grained multimodal dataset and a two-stage fine-tuning strategy that improves multimodal large language models' ability…
LAST: The Last Query Token Guides Visual Token Pruning for Edge-Cloud Collaborative MLLM Inference
Feng Yang, Xinrui Ju, Keyang Zhang +6
The paper introduces LAST, a training‑free method that uses the attention of the last query token to prune visual tokens on edge devices before sending them to a cloud multimodal L…
VizPilot: Automated Onboarding for SVG-based Composite Visualizations using Multimodal LLMs
Nishaanthini Gnanavel, Yong Wang
VizPilot is a browser extension that automatically creates interactive onboarding guides for SVG-based composite visualizations by analyzing their structure with multimodal large l…