1 citations · 1 across the 2 of their papers we have counts for
3 papers
cs.AI2026
Hindsight Hint Distillation: Scaffolded Reasoning for SWE Agents from CoT-free Answers
Shengjie Wang, Guanghe Li, Zonghan Yang +1
Solving complex long-horizon tasks requires strong planning and reasoning capabilities. Although datasets with explicit chain-of-thought (CoT) rationales can substantially benefit…
cs.CL2026
Kimi K2.5: Visual Agentic Intelligence
Kimi Team, Tongtong Bai, Yifan Bai +333
We introduce Kimi K2.5, an open-source multimodal agentic model designed to advance general agentic intelligence. K2.5 emphasizes the joint optimization of text and vision so that…
cs.LG2024★ 1 cited
DiffStitch: Boosting Offline Reinforcement Learning with Diffusion-based Trajectory Stitching
Guanghe Li, Yixiang Shan, Zhengbang Zhu +2
In offline reinforcement learning (RL), the performance of the learned policy highly depends on the quality of offline datasets. However, in many cases, the offline dataset contain…