3 papers
cs.CV2026
ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding
Jovana Kondic, Pengyuan Li, Dhiraj Joshi +24
Understanding charts requires models to jointly reason over geometric visual patterns, structured numerical data, and natural language -- a capability where current vision-language…
cs.CV2025
Q-SAM2: Accurate Quantization for Segment Anything Model 2
Nicola Farronato, Florian Scheidegger, Mattia Rigotti +3
The Segment Anything Model 2 (SAM2) is a powerful foundation model for promptable segmentation. However, its high computational and memory costs are a major barrier to deployment o…
cs.CV2025
VP Lab: a PEFT-Enabled Visual Prompting Laboratory for Semantic Segmentation
Niccolo Avogaro, Thomas Frick, Yagmur G. Cinar +12
Large-scale pretrained vision backbones have transformed computer vision by providing powerful feature extractors that enable various downstream tasks, including training-free appr…