2 papers
cs.CV2025
Adaptively Clustering Neighbor Elements for Image-Text Generation
Zihua Wang, Xu Yang, Hanwang Zhang +4
We propose a novel Transformer-based image-to-text generation model termed as \textbf{ACF} that adaptively clusters vision patches into object regions and language words into phras…
cs.CV2024
Lever LM: Configuring In-Context Sequence to Lever Large Vision Language Models
Xu Yang, Yingzhe Peng, Haoxuan Ma +4
As Archimedes famously said, ``Give me a lever long enough and a fulcrum on which to place it, and I shall move the world'', in this study, we propose to use a tiny Language Model…