1 citations · 1 across the 16 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
OccamToken: Efficient VLM Inference with Training-Free and Budget-Adaptive Token Pruning
Geng Li, Guohao Chen, Ting Chen +6
Vision-language models (VLMs) rely on long visual token sequences for visual understanding, making the prefill stage expensive in both computation and memory. Most existing pruning…
cs.CV2026
CompassAD: Intent-Driven 3D Affordance Grounding in Functionally Competing Objects
Jingliang Li, Jindou Jia, Tuo An +7
When told to "cut the cake," a robot must choose the knife over nearby scissors, despite both objects affording the same cutting function. In real-world scenes, multiple objects ma…
cs.CV2025
Dual Enhancement on 3D Vision-Language Perception for Monocular 3D Visual Grounding
Yuzhen Li, Min Liu, Yuan Bian +4
Monocular 3D visual grounding is a novel task that aims to locate 3D objects in RGB images using text descriptions with explicit geometry information. Despite the inclusion of geom…