Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Dual-Adaptive SAM3: Hierarchical Routing over Low-Rank Expert Layers for Parameter-Efficient Medical Image Segmentation
Ying Chen, Jinyue Li, Kun Wang +2
The Segment Anything Model with Concepts (SAM3) heralds a new paradigm for open-vocabulary segmentation through natural language interaction, offering significant potential for med…
cs.CV2026
CircuitProbe: Tracing Visual Temporal Evidence Flow in Video Language Models
Yiming Zhang, Zhuokai Zhao, Chengzhang Yu +8
Autoregressive large vision--language models (LVLMs) interface video and language by projecting video features into the LLM's embedding space as continuous visual token embeddings.…
cs.CV2026
Granulon: Awakening Pixel-Level Visual Encoders with Adaptive Multi-Granularity Semantics for MLLM
Junyuan Mao, Qiankun Li, Linghao Meng +5
Recent advances in multimodal large language models largely rely on CLIP-based visual encoders, which emphasize global semantic alignment but struggle with fine-grained visual unde…