3 papers
cs.LG2026
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
Changhai Zhou, Kieran Liu, Yuhua Zhou +17
Long-context RL post-training is constrained by the lifetime of state and gradients, not attention cost alone. In GRPO, one multi-million-token prompt must serve old-policy and ref…
cs.LG2025
ProtTeX-CC: Activating In-Context Learning in Protein LLM via Two-Stage Instruction Compression
Chuanliu Fan, Zicheng Ma, Jun Gao +5
Recent advances in protein large language models, such as ProtTeX, represent both side-chain amino acids and backbone structure as discrete token sequences of residue length. While…
cs.CV2024
Interleaved-Modal Chain-of-Thought
Jun Gao, Yongqi Li, Ziqiang Cao +1
Chain-of-Thought (CoT) prompting elicits large language models (LLMs) to produce a series of intermediate reasoning steps before arriving at the final answer. However, when transit…