1 citations · 1 across the 4 of their papers we have counts for
8 papers
Kimi K3: Open Frontier Intelligence
Kimi Team, Tongtong Bai, Yifan Bai +398
We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is…
Quantizing Intent: Cross-Domain Semantic IDs from Organic Activity for Industrial Ranking
Julie Choi, Haoran Ye, Zhiwei Ding +3
Ads click-through rate (CTR) prediction is constrained by sparse user supervision: most users engage with ads infrequently while generating dense behavioral evidence in organic sur…
AutoEP: LLMs-Driven Automation of Hyperparameter Evolution for Metaheuristic Algorithms
Zhenxing Xu, Yizhe Zhang, Weidong Bao +6
Dynamically configuring algorithm hyperparameters is a fundamental challenge in computational intelligence. While learning-based methods offer automation, they suffer from prohibit…
Meta Context Engineering via Agentic Skill Evolution
Haoran Ye, Xuning He, Vincent Arak +2
The operational efficacy of large language models relies heavily on their inference-time context. This has established Context Engineering (CE) as a formal discipline for optimizin…
Meta-R1: Empowering Large Reasoning Models with Metacognition
Haonan Dong, Haoran Ye, Wenhao Zhu +2
Large Reasoning Models (LRMs) demonstrate remarkable capabilities on complex tasks, exhibiting emergent, human-like thinking patterns. Despite their advances, we identify a fundame…
Breaking the Block: Preserving Data Continuity to Train Superior SAEs for Instruct Models
Jiaming Li, Haoran Ye, Yukun Chen +5
Sparse Autoencoders (SAEs) are a cornerstone of mechanistic interpretability. Existing training methods inherit the Block Training paradigm from LLM pre-training, which introduces…