Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
Kimi-VL Technical Report
Kimi Team, Angang Du, Bohong Yin +92
We present Kimi-VL, an efficient open-source Mixture-of-Experts (MoE) vision-language model (VLM) that offers advanced multimodal reasoning, long-context understanding, and strong…
cs.CV2024
Benchmarking Large Language Models for Image Classification of Marine Mammals
Yijiashun Qi, Shuzhang Cai, Zunduo Zhao +3
As Artificial Intelligence (AI) has developed rapidly over the past few decades, the new generation of AI, Large Language Models (LLMs) trained on massive datasets, has achieved gr…
cs.CV2024
Instruction-Guided Visual Masking
Jinliang Zheng, Jianxiong Li, Sijie Cheng +6
Instruction following is crucial in contemporary LLM. However, when extended to multimodal setting, it often suffers from misalignment between specific textual instruction and targ…