Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
Negative Entity Suppression for Zero-Shot Captioning with Synthetic Images
Zimao Lu, Hui Xu, Bing Liu +1
Text-only training provides an attractive approach to address data scarcity challenges in zero-shot image captioning (ZIC), avoiding the expense of collecting paired image-text ann…
cs.CV2024
LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information
Ke Wang, Hong Xuan
Multi-modal large language models (MLLMs) utilizing instruction-following data, such as LLaVA, have achieved great progress in the industry. A major limitation in these models is t…
cs.CV2024
E-ANT: A Large-Scale Dataset for Efficient Automatic GUI NavigaTion
Ke Wang, Tianyu Xia, Zhangxuan Gu +5
Online GUI navigation on mobile devices has driven a lot of attention recent years since it contributes to many real-world applications. With the rapid development of large languag…