3 papers
cs.CV2024
Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction
Shiyu Zhao, Zhenting Wang, Felix Juefei-Xu +7
Prevailing Multimodal Large Language Models (MLLMs) encode the input image(s) as vision tokens and feed them into the language backbone, similar to how Large Language Models (LLMs)…
cs.CV2024
Unleashing In-context Learning of Autoregressive Models for Few-shot Image Manipulation
Bolin Lai, Felix Juefei-Xu, Miao Liu +8
Text-guided image manipulation has experienced notable advancement in recent years. In order to mitigate linguistic ambiguity, few-shot learning with visual examples has been appli…
cond-mat.str-el2024
Chemical versus physical pressure effects on the structure transition of bilayer nickelates
Gang Wang, Ningning Wang, Tenglong Lu +11
The observation of high- superconductivity (HTSC) in concomitant with pressure-induced orthorhombic-tetragonal structural transition in the bilayer LaNiO has spa…