1 citations · 1 across the 5 of their papers we have counts for
4 papers · 1 filter
Kwai Keye-VL 1.5 Technical Report
Biao Yang, Bin Wen, Boyang Ding +58
In recent years, the development of Large Language Models (LLMs) has significantly advanced, extending their capabilities to multimodal tasks through Multimodal Large Language Mode…
Kwai Keye-VL Technical Report
Kwai Keye Team, Biao Yang, Bin Wen +57
While Multimodal Large Language Models (MLLMs) demonstrate remarkable capabilities on static images, they often fall short in comprehending dynamic, information-dense short-form vi…
EVLM: An Efficient Vision-Language Model for Visual Understanding
Kaibing Chen, Dong Shen, Hanwen Zhong +14
In the field of multi-modal language models, the majority of methods are built on an architecture similar to LLaVA. These models use a single-layer ViT feature as a visual prompt,…
Block-wise LoRA: Revisiting Fine-grained LoRA for Effective Personalization and Stylization in Text-to-Image Generation
Likun Li, Haoqi Zeng, Changpeng Yang +2
The objective of personalization and stylization in text-to-image is to instruct a pre-trained diffusion model to analyze new concepts introduced by users and incorporate them into…