10 papers
Kwai Keye-VL-2.0 Technical Report
Kwai Keye Team, Bin Wen, Changyi Liu +50
We introduce Kwai Keye-VL-2.0-30B-A3B, an open-source Mixture-of-Experts (MoE) multimodal foundation model designed to advance long-video understanding and agentic intelligence. To…
KairosAgent: Agentic Time Series Forecasting with Fused Semantic Reasoning
Kun Feng, Ziwei Shan, Yuchen Fang +6
Cross-domain multimodal time series forecasting is a challenging task, requiring models to integrate precise numerical comprehension, cross-domain semantic understanding, and effec…
Does The Way You Plan Matter? An Empirical Study of Planning Representations for LLM Web Agents
Alejandra Zambrano, Sara Vera Marjanovic, Imene Kerboua +2
Despite recent advances, LLM-based web agents still struggle with limited exploration, omission of critical steps, and sensitivity to task constraints. Prior work suggests that man…
OptSkills: Learning Generalizable Optimization Skills from Problem Archetypes via Cluster-Based Distillation
Haochen Yang, Ke Zhao, Mengyuan Ma +3
Leveraging Large Language Models (LLMs) to automatically formulate and solve optimization problems from natural language has emerged as an efficient paradigm for automated optimiza…
VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning
Xingyu Lu, Jinpeng Wang, Yi-Fan Zhang +13
Visual captioning requires models to capture visual content faithfully while minimizing both omission and hallucination. As the dominant paradigm for captioning, MLLMs have achieve…
ContextRL: Enhancing MLLM's Knowledge Discovery Efficiency with Context-Augmented RL
Xingyu Lu, Jinpeng Wang, YiFan Zhang +12
We propose ContextRL, a novel framework that leverages context augmentation to overcome these bottlenecks. Specifically, to enhance Identifiability, we provide the reward model wit…