5 papers · 1 filter
Rethinking Token-Level Policy Optimization for Multimodal Chain-of-Thought
Yunheng Li, Hangyi Kuang, Hengrui Zhang +4
Multimodal Chain-of-Thought (CoT) reasoning requires large vision-language models to construct reasoning trajectories that interleave perceptual grounding with multi-step inference…
CREM: Compression-Driven Representation Enhancement for Multimodal Retrieval and Comprehension
Lihao Liu, Yan Wang, Biao Yang +10
Multimodal Large Language Models (MLLMs) have shown remarkable success in comprehension tasks such as visual description and visual question answering. However, their direct applic…
Kelix Technical Report
Boyang Ding, Chenglong Chu, Dunju Zang +28
Autoregressive large language models (LLMs) scale well by expressing diverse tasks as sequences of discrete natural-language tokens and training with next-token prediction, which u…
Kwai Keye-VL 1.5 Technical Report
Biao Yang, Bin Wen, Boyang Ding +58
In recent years, the development of Large Language Models (LLMs) has significantly advanced, extending their capabilities to multimodal tasks through Multimodal Large Language Mode…
Kwai Keye-VL Technical Report
Kwai Keye Team, Biao Yang, Bin Wen +57
While Multimodal Large Language Models (MLLMs) demonstrate remarkable capabilities on static images, they often fall short in comprehending dynamic, information-dense short-form vi…