latent world models 1multimodal prediction 1physical parameter identification 1representation learning 1robotics simulation 1
From the 1 of 5 linked papers with an AI index.
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
CREG: Compass Relational Evidence Graph for Characterizing Directional Structure in VLM Spatial-Reasoning Attribution
Kaizhen Tan, Yang Feng, Heqing Du
Standard attribution heatmaps show where a vision-language model (VLM) focuses, but they do not reveal whether the recovered evidence is organized by the queried spatial relation o…
cs.CV2026
When Does Visual Token Pruning Improve Calibration? The Role of Evidence Coverage in MLLMs
Kaizhen Tan, Yang Feng, Heqing Du +3
Visual token pruning is widely used to reduce the inference cost of multimodal large language models (MLLMs), but it is usually evaluated only by accuracy. We study how pruning aff…