4 papers · 1 filter
OccamToken: Efficient VLM Inference with Training-Free and Budget-Adaptive Token Pruning
Geng Li, Guohao Chen, Ting Chen +6
Vision-language models (VLMs) rely on long visual token sequences for visual understanding, making the prefill stage expensive in both computation and memory. Most existing pruning…
Rethinking VLM Representation for VLA Initialization
Weifeng Lin, Siyuan Huang, Hao Li +5
Vision-Language-Action (VLA) models widely adopt pretrained Vision-Language Models (VLMs) as policy backbones, yet it remains unclear what kind of pretrained VLM representation is…
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
Weifeng Lin, Xinyu Wei, Ruichuan An +7
We present Perceive Anything Model (PAM), a conceptually straightforward and efficient framework for comprehensive region-level visual understanding in images and videos. Our appro…
Layered Diffusion Model for One-Shot High Resolution Text-to-Image Synthesis
Emaad Khwaja, Abdullah Rashwan, Ting Chen +3
We present a one-shot text-to-image diffusion model that can generate high-resolution images from natural language descriptions. Our model employs a layered U-Net architecture that…