Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs
Yue Cheng, Jiajun Zhang, Xiaohui Gao +3
Reinforcement Learning with Verifiable Reward (RLVR) is empirically shown to notably enhance the reasoning performance of large language models (LLMs), particularly in mathematics…
cs.AI2026
Benchmarking the Limits of In-Context Reinforcement Learning for Ad-Hoc Teamwork
Yuheng Jing, Kai Li, Ziwen Zhang +8
In-Context Reinforcement Learning (ICRL) has enabled foundation agents to adapt instantaneously to novel tasks, yet its efficacy in Ad-Hoc Teamwork (AHT)-where coordination with un…