8 papers
STU: Stateful Test-Time Unlearning via Restricted Knowledge Boundary Control
Xunlei Chen, Qinghui Gong, Ruini Xue +3
Controlling restricted knowledge in large language models is essential for model alignment and safe deployment. Test-time unlearning avoids costly retraining and parameter updates…
UAM: A Dual-Stream Perspective on Forgetting in VLA Training
Jianke Zhang, Yuanfei Luo, Yucheng Hu +6
Vision--language--action (VLA) models are typically built by fine-tuning a pretrained vision--language model (VLM) on action data. However, we show that this standard recipe system…
Improving Multi-turn Dialogue Consistency with Self-Recall Thinking
Renning Pang, Tian Lan, Leyuan Liu +3
Large language model (LLM) based multi-turn dialogue systems often struggle to track dependencies across non-adjacent turns, undermining both consistency and scalability. As conver…
Case-Based Calibration of Adaptive Reasoning and Execution for LLM Tool Use
Renning Pang, Tian Lan, Leyuan Liu +3
Tool use extends large language models beyond parametric knowledge, but reliable execution requires balancing appropriate reasoning depth with strict structural validity. We approa…
Matrix-Space Reinforcement Learning for Reusing Local Transition Geometry
Zuyuan Zhang, Carlee Joe-Wong, Tian Lan
Compositional generalization in sequential decision-making requires identifying which parts of prior rollouts remain useful for new tasks. Existing methods reuse skills or predicti…
Interactive Critique-Revision Training for Reliable Structured LLM Generation
Fei Xu Yu, Zuyuan Zhang, Mahdi Imani +2
In structured decision-making workflows such as form filling, compliance checking, and maintenance reporting, LLM outputs must be locally correct, globally consistent, and auditabl…