4 citations · 12 across the 22 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs
Jingtan Wang, Arun Verma, Xiaoqiang Lin +4
How to divide a fixed annotation budget between supervised fine-tuning (SFT) and reinforcement learning (RL) during LLM post-training remains an open problem. Existing work charact…
cs.CL2025
MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents
Zijian Zhou, Ao Qu, Zhaoxuan Wu +6
Modern language agents must operate over long-horizon, multi-turn interactions, where they retrieve external information, adapt to observations, and answer interdependent queries.…
cs.CL2025
TETRIS: Optimal Draft Token Selection for Batch Speculative Decoding
Zhaoxuan Wu, Zijian Zhou, Arun Verma +3
We propose TETRIS, a novel method that optimizes the total throughput of batch speculative decoding in multi-request settings. Unlike existing methods that optimize for a single re…