5 papers
Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models
Peiqi Jia, Haonan Jia, Ziqi Miao +3
With the widespread deployment of Multimodal Large Language Models (MLLMs) in social interaction, understanding and controlling their behavior under complex personality conditions…
Kwai Keye-VL-2.0 Technical Report
Kwai Keye Team, Bin Wen, Changyi Liu +50
We introduce Kwai Keye-VL-2.0-30B-A3B, an open-source Mixture-of-Experts (MoE) multimodal foundation model designed to advance long-video understanding and agentic intelligence. To…
Seeing with You: Perception-Reasoning Coevolution for Multimodal Reasoning
Ziqi Miao, Haonan Jia, Lijun Li +4
Reinforcement learning with verifiable rewards (RLVR) has substantially enhanced the reasoning capabilities of multimodal large language models (MLLMs). However, existing RLVR appr…
Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters
Ailin Huang, Ang Li, Aobo Kong +213
We introduce Step 3.5 Flash, a sparse Mixture-of-Experts (MoE) model that bridges frontier-level agentic intelligence and computational efficiency. We focus on what matters most wh…
TabSieve: Explicit In-Table Evidence Selection for Tabular Prediction
Yongyao Wang, Ziqi Miao, Lu Yang +4
Tabular prediction can benefit from in-table rows as few-shot evidence, yet existing tabular models typically perform instance-wise inference and LLM-based prompting is often britt…