From the 1 of 10 linked papers with an AI index.
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
VT-Bench: A Unified Benchmark for Visual-Tabular Multi-Modal Learning
Zi-Yi Jia, Zi-Jian Cheng, Xin-Yue Zhang +4
Multi-model learning has attracted great attention in visual-text tasks. However, visual-tabular data, which plays a pivotal role in high-stakes domains like healthcare and industr…
cs.CV2026
LAST: Leveraging Tools as Hints to Enhance Spatial Reasoning for Multimodal Large Language Models
Shi-Yu Tian, Zhi Zhou, Kun-Yang Yu +5
Spatial reasoning is a cornerstone capability for intelligent systems to perceive and interact with the physical world. However, multimodal large language models (MLLMs) frequently…