14 citations · 30 across the 10 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning
Liang Chen, Hongcheng Gao, Tianyu Liu +5
Vision-Language Models (VLMs) excel in many direct multimodal tasks but struggle to translate this prowess into effective decision-making within interactive, visually rich environm…
cs.CV2025★ 1 cited
Kimi-VL Technical Report
Kimi Team, Angang Du, Bohong Yin +92
We present Kimi-VL, an efficient open-source Mixture-of-Experts (MoE) vision-language model (VLM) that offers advanced multimodal reasoning, long-context understanding, and strong…