Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
GeoTikzBridge: Advancing Multimodal Code Generation for Geometric Perception and Reasoning
Jiayin Sun, Caixia Sun, Boyu Yang +7
Multimodal Large Language Models (MLLMs) have recently demonstrated remarkable perceptual and reasoning abilities. However, they struggle to perceive fine-grained geometric structu…
cs.CV2024
Uni-AdaFocus: Spatial-temporal Dynamic Computation for Video Recognition
Yulin Wang, Haoji Zhang, Yang Yue +4
This paper presents a comprehensive exploration of the phenomenon of data redundancy in video understanding, with the aim to improve computational efficiency. Our investigation com…