2 citations · 7 across the 32 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Teaching Vision-Language Models to Use the Scale They Are Given: Label-Free Equivariance Training for Metric Physical Reasoning
Kaizhen Tan, Yang Feng, Heqing Du +3
Metric questions about video require vision-language models to use supplied real-world references to convert visual measurements into physical units. Yet we find that current model…
cs.CV2026
From Attenuation to Attention: Variational Information Flow Manipulation for Fine-Grained Visual Perception
Jilong Zhu, Yang Feng
While Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in general visual understanding, they frequently falter in fine-grained perception tasks th…
cs.CV2025
LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Shaolei Zhang, Qingkai Fang, Zhe Yang +1
The advent of real-time large multimodal models (LMMs) like GPT-4o has sparked considerable interest in efficient LMMs. LMM frameworks typically encode visual inputs into vision to…