6 papers
Kimi K2.5: Visual Agentic Intelligence
Kimi Team, Tongtong Bai, Yifan Bai +339
We introduce Kimi K2.5, an open-source multimodal agentic model designed to advance general agentic intelligence. K2.5 emphasizes the joint optimization of text and vision so that…
Parallels Between VLA Model Post-Training and Human Motor Learning: Progress, Challenges, and Trends
Tian-Yu Xiang, Ao-Qun Jin, Xiao-Hu Zhou +11
Vision-language-action (VLA) models extend vision-language models (VLM) by integrating action generation modules for robotic manipulation. Leveraging the strengths of VLM in vision…
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
Ziyu Guo, Renrui Zhang, Hongyu Li +6
Recent advances in visual generation have increasingly explored the integration of reasoning capabilities. They incorporate textual reasoning, i.e., think, either before (as pre-pl…
VLA Model Post-Training via Action-Chunked PPO and Self Behavior Cloning
Si-Cheng Wang, Tian-Yu Xiang, Xiao-Hu Zhou +6
Reinforcement learning (RL) is a promising avenue for post-training vision-language-action (VLA) models, but practical deployment is hindered by sparse rewards and unstable trainin…
DVLTA-VQA: Decoupled Vision-Language Modeling with Text-Guided Adaptation for Blind Video Quality Assessment
Li Yu, Situo Wang, Wei Zhou +1
Inspired by the dual-stream theory of the human visual system (HVS) - where the ventral stream is responsible for object recognition and detail analysis, while the dorsal stream fo…
VLA Model-Expert Collaboration for Bi-directional Manipulation Learning
Tian-Yu Xiang, Ao-Qun Jin, Xiao-Hu Zhou +8
The emergence of vision-language-action (VLA) models has given rise to foundation models for robot manipulation. Although these models have achieved significant improvements, their…