Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Teaching Vision-Language Models to Use the Scale They Are Given: Label-Free Equivariance Training for Metric Physical Reasoning
Kaizhen Tan, Yang Feng, Heqing Du +3
Metric questions about video require vision-language models to use supplied real-world references to convert visual measurements into physical units. Yet we find that current model…
cs.CV2026
GeoFidelity-Bench: Evaluating Segment-Level Geographic Fidelity in Text-to-Image Street-View Generation
Kaizhen Tan, Hanzhe Hong, Siru Tao
Text-to-image models can generate visually plausible city streets, but whether their outputs correspond to a requested road segment rather than a generic city prior remains unclear…
cs.CV2026
When Does Visual Token Pruning Improve Calibration? The Role of Evidence Coverage in MLLMs
Kaizhen Tan, Yang Feng, Heqing Du +3
Visual token pruning is widely used to reduce the inference cost of multimodal large language models (MLLMs), but it is usually evaluated only by accuracy. We study how pruning aff…