6 papers
UAM: A Dual-Stream Perspective on Forgetting in VLA Training
Jianke Zhang, Yuanfei Luo, Yucheng Hu +6
Vision--language--action (VLA) models are typically built by fine-tuning a pretrained vision--language model (VLM) on action data. However, we show that this standard recipe system…
Improving Multi-turn Dialogue Consistency with Self-Recall Thinking
Renning Pang, Tian Lan, Leyuan Liu +3
Large language model (LLM) based multi-turn dialogue systems often struggle to track dependencies across non-adjacent turns, undermining both consistency and scalability. As conver…
Case-Based Calibration of Adaptive Reasoning and Execution for LLM Tool Use
Renning Pang, Tian Lan, Leyuan Liu +3
Tool use extends large language models beyond parametric knowledge, but reliable execution requires balancing appropriate reasoning depth with strict structural validity. We approa…
Matrix-Space Reinforcement Learning for Reusing Local Transition Geometry
Zuyuan Zhang, Carlee Joe-Wong, Tian Lan
Compositional generalization in sequential decision-making requires identifying which parts of prior rollouts remain useful for new tasks. Existing methods reuse skills or predicti…
Interactive Critique-Revision Training for Reliable Structured LLM Generation
Fei Xu Yu, Zuyuan Zhang, Mahdi Imani +2
In structured decision-making workflows such as form filling, compliance checking, and maintenance reporting, LLM outputs must be locally correct, globally consistent, and auditabl…
Diffusion-Based Cross-Modal Feature Extraction for Multi-Label Classification
Tian Lan, Yiming Zheng, Jianxin Yin
Multi-label classification has broad applications and depends on powerful representations capable of capturing multi-label interactions. We introduce \textit{Diff-Feat}, a simple b…