2 papers
cs.CV2026
VisionZip: Longer is Better but Not Necessary in Vision Language Models
Senqiao Yang, Yukang Chen, Zhuotao Tian +4
Recent advancements in vision-language models have enhanced performance by increasing the length of visual tokens, making them much longer than text tokens and significantly raisin…
cs.CV2024
Decoupled Kullback-Leibler Divergence Loss
Jiequan Cui, Zhuotao Tian, Zhisheng Zhong +3
In this paper, we delve deeper into the Kullback-Leibler (KL) Divergence loss and mathematically prove that it is equivalent to the Decoupled Kullback-Leibler (DKL) Divergence loss…