3 papers
cs.AI2026
Reading is not Reasoning: Bridging the Agentic Policy Gap in Vision-Text Compression
Cheng Fan, Junyi Zhou, Tingzhang Luo +5
Multi-step language-model agents repeatedly process growing interaction histories, leading to substantial context costs. Vision--text compression reduces these costs by rendering h…
cs.CV2026
RoRA: Role-Oriented Regional Allocation for Visual Token Pruning in MLLMs
Qiyanhui Lu, Han Wu, Rongjian Xu +6
Multimodal large language models (MLLMs) encode images as long visual token sequences, making prefilling and KV-cache storage expensive. Existing training-free pruning methods sele…
cs.CV2026
CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation
Tingzhang Luo, Ruizhong Liu, Yichao Liu +3
Referring Remote Sensing Image Segmentation (RRSIS) has achieved significant progress through the integration of VLMs and the Segment Anything Model (SAM). However, this progress l…