2 papers
cs.AI2026
General learned delegation by clones
Darren Li, Meiqi Chen, Chenze Shao +2
Frontier language models improve with additional test-time computation, but serial reasoning or uncoordinated parallel sampling can be compute-inefficient under fixed inference bud…
cs.CL2025
Continuous Autoregressive Language Models
Chenze Shao, Darren Li, Fandong Meng +1
The efficiency of large language models (LLMs) is fundamentally limited by their sequential, token-by-token generation process. We argue that overcoming this bottleneck requires a…