Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs
Jihwan Kim, Nikhil Parthasarathy, Danfeng Qin +5
The fundamental challenge in scaling Video Large Language Models (Video LLMs) to long-form video lies in managing the explosion of visual-token context length. Existing strategies…
cs.CV2024
MobileNetV4 -- Universal Models for the Mobile Ecosystem
Danfeng Qin, Chas Leichner, Manolis Delakis +11
We present the latest generation of MobileNets, known as MobileNetV4 (MNv4), featuring universally efficient architecture designs for mobile devices. At its core, we introduce the…