2 papers
cs.AI2025
NUMINA: A Natural Understanding Benchmark for Multi-dimensional Intelligence and Numerical Reasoning Abilities
Changyu Zeng, Yifan Wang, Zimu Wang +6
Recent advancements in 2D multimodal large language models (MLLMs) have significantly improved performance in vision-language tasks. However, extending these capabilities to 3D env…
cs.CV2025
FTCFormer: Fuzzy Token Clustering Transformer for Image Classification
Muyi Bao, Changyu Zeng, Yifan Wang +5
Transformer-based deep neural networks have achieved remarkable success across various computer vision tasks, largely attributed to their long-range self-attention mechanism and sc…