Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
A-MAR: Agent-based Multimodal Art Retrieval for Fine-Grained Artwork Understanding
Shuai Wang, Hongyi Zhu, Jia-Hong Huang +6
Understanding artworks requires multi-step reasoning over visual content and cultural, historical, and stylistic context. While recent multimodal large language models show promise…
cs.AI2026
SAM3-LiteText: An Anatomical Study of the SAM3 Text Encoder for Efficient Vision-Language Segmentation
Chengxi Zeng, Yuxuan Jiang, Ge Gao +6
Vision-language segmentation models such as SAM3 enable flexible, prompt-driven visual grounding, but inherit large, general-purpose text encoders originally designed for open-ende…