2 papers
cs.LG2026
In-Cell Learning: Language Models That Update Their Own Weights in Sequence Without Changing the File They Ship
Zifeng Liu, Yaxin Lu, Xuanhan Wu +6
A 4-bit quantized weight specifies a rounding cell rather than a single full-precision value. We introduce in-cell learning, a paradigm for writing new knowledge only within these…
cs.LG2026
Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO
Yunpeng Ba, Zhi Zheng, Yue Xie +7
Evolution Strategies (ES) have recently emerged as a memory-efficient post-training paradigm for LLM reasoning. However, the optimization behavior of ES remains understudied, makin…