collaborators

6 papers

cs.CV2025

Kwai Keye-VL 1.5 Technical Report

Biao Yang, Bin Wen, Boyang Ding +58

In recent years, the development of Large Language Models (LLMs) has significantly advanced, extending their capabilities to multimodal tasks through Multimodal Large Language Mode…

cs.CV2025

Thyme: Think Beyond Images

Yi-Fan Zhang, Xingyu Lu, Shukang Yin +17

Following OpenAI's introduction of the ``thinking with images'' concept, recent efforts have explored stimulating the use of visual information in the reasoning process to enhance…

cs.RO2025

A Navigation Framework Utilizing Vision-Language Models

Yicheng Duan, Kaiyu tang

Vision-and-Language Navigation (VLN) presents a complex challenge in embodied AI, requiring agents to interpret natural language instructions and navigate through visually rich, un…

cs.CV2025

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Yi-Fan Zhang, Xingyu Lu, Xiao Hu +13

Multimodal Reward Models (MRMs) play a crucial role in enhancing the performance of Multimodal Large Language Models (MLLMs). While recent advancements have primarily focused on im…

cs.SI2025

VLM as Policy: Common-Law Content Moderation Framework for Short Video Platform

Xingyu Lu, Tianke Zhang, Chang Meng +14

Exponentially growing short video platforms (SVPs) face significant challenges in moderating content detrimental to users' mental health, particularly for minors. The dissemination…

cs.CL2024

Kwai-STaR: Transform LLMs into State-Transition Reasoners

Xingyu Lu, Yuhang Hu, Changyi Liu +12

Mathematical reasoning presents a significant challenge to the cognitive capabilities of LLMs. Various methods have been proposed to enhance the mathematical ability of LLMs. Howev…