2 papers
cs.LG2026
How to Compress KV Cache in RL Post-Training? Shadow Mask Distillation for Memory-Efficient Alignment
Rui Zhu, Weiheng Bai, Qiushi Wu +3
Reinforcement Learning (RL) has emerged as a crucial paradigm for unlocking the advanced reasoning capabilities of Large Language Models (LLMs), encompassing frameworks like RLHF a…
astro-ph.IM2025
BREAKFAST: A Framework for general joint BA duty and follow-up guidance of multiple -ray monitors
Chen-Wei Wang, Peng Zhang, Shao-Lin Xiong +30
With the growing number of gamma-ray monitors in operation, several research teams have adopted a strategy of joint operation and scientific duty to improve efficiency. A successfu…