1 paper · 1 filter
Chenxing Wei, Hong Wang, Ying He +4
Test-time policy adaptation for multi-turn interactions (T2PAM) is essential for aligning Large Language Models (LLMs) with dynamic user needs during inference time. However, exist…