18 papers
Solar Open 2 Technical Report
Sungrae Park, Sanghoon Kim, Gyoungjin Gim +50
We present Solar Open 2, a 250B-A15B Mixture-of-Experts language model built for long-horizon agentic tasks, scaled up from Solar Open 1 (Solar Open 100B). To hold entire agent tra…
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop
Yining Hong, Jiageng Liu, Han Yin +5
Spatial intelligence unfolds through a perception-action loop: agents act to acquire observations, and reason about how observations vary as a function of action. Rather than passi…
Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention
Ali Hatamizadeh, Yejin Choi, Jan Kautz
Linear attention replaces the unbounded cache of softmax attention with a fixed-size recurrent state, reducing sequence mixing to linear time and decoding to constant memory. The h…
Introspective X Training: Feedback Conditioning Improves Scaling Across all LLM Training Stages
Brandon Cui, Ximing Lu, Jaehun Jung +7
We tackle the question of how to scale more efficiently across the many, ever-growing stages of current LLM training pipelines. Our guiding intuition stems from the fact that the d…
DeltaPrompts: Escaping the Zero-Delta Trap in Multimodal Distillation
Jaehun Jung, Hyunwoo Kim, Brandon Cui +4
Distillation enables compact Vision-Language Models (VLMs) to obtain strong reasoning capabilities, yet the prompts driving this process are typically chosen via simple heuristics…
How to Instruct Your Robot: Dense Language Annotations Power Robot Policy Learning
Bosung Kim, Ruiyi Wang, David Acuna +5
Scaling robot policy learning is bottlenecked by the cost of collecting demonstrations, while language annotations for existing demonstrations are comparatively cheap. We study lan…