5 citations · 12 across the 21 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
WISE: Weak-Supervision-Guided Step-by-Step Explanations for Multimodal LLMs in Image Classification
Yiwen Jiang, Deval Mehta, Siyuan Yan +3
Multimodal Large Language Models (MLLMs) have shown promise in visual-textual reasoning, with Multimodal Chain-of-Thought (MCoT) prompting significantly enhancing interpretability.…
cs.CV2025
FTCFormer: Fuzzy Token Clustering Transformer for Image Classification
Muyi Bao, Changyu Zeng, Yifan Wang +5
Transformer-based deep neural networks have achieved remarkable success across various computer vision tasks, largely attributed to their long-range self-attention mechanism and sc…