distributed execution 1GPU memory efficiency 1large language model training 1long-context reinforcement learning 1policy optimization 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.LG2026
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA
Mind Lab, :, Vin Bo +74
Macaron-V1 is an open agent-model family for experiential intelligence: learning from experience in real environments and continuing to learn after deployment. It is organized arou…
cs.LG2026
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
Changhai Zhou, Kieran Liu, Yuhua Zhou +17
LongStraw introduces an execution framework that enables reinforcement‑learning post‑training on million‑token prompts using a fixed GPU budget by separating prompt evaluation from…
cs.HC2026
Macaron-A2UI: A Model for Generative UI in Personal Agents
Fancy Kong, Congjie Zheng, Murphy Zhuang +8
As personal agents evolve to handle complex, user-centric tasks, static plain-text chat is rapidly becoming a bottleneck. Generative UI emerges as the necessary new interface layer…