42 papers
ExToken: Structured Exploration for Efficient Vision-Language-Action Reinforcement Fine-tuning
Yilun Kong, Yunpeng Qing, Guozheng Ma +4
The paper proposes ExToken, a framework that conditions vision‑language‑action policies on discrete behavioral tokens derived from offline demonstrations to promote diverse, struct…
From Hindsight to Foresight: Self-Encouraged Hindsight Distillation for Knowledge-based Visual Question Answering
Yu Zhao, Ying Zhang, Xuhui Sui +4
The paper introduces a teacher‑student framework called Hindsight Distillation (HinD) that uses privileged answer information to generate reasoning trajectories for a multimodal LL…
Hyper-modal Imputation Diffusion Embedding with Dual-Distillation for Federated Multimodal Knowledge Graph Completion
Ying Zhang, Yu Zhao, Xuhui Sui +5
The paper introduces a federated learning framework for multimodal knowledge graph completion that recovers missing multimodal information with a hyper-modal imputation diffusion e…
Language-based Trial and Error Falls Behind in the Era of Experience
Haoyu Wang, Guozheng Ma, Shugang Cui +7
While Large Language Models (LLMs) excel in language-based agentic tasks, their applicability to unseen, nonlinguistic environments (e.g., symbolic or spatial tasks) remains limite…
Closed-Form Spectral Regularization for Multi-Task Model Merging
Yongxian Wei, Runxi Cheng, Xingxuan Zhang +4
Model merging combines several independently fine-tuned experts into a single multi-task model without any training data, reducing the storage, serving, and decentralized-developme…
Convergent Differential Privacy Analysis for General Federated Learning
Yan Sun, Qixin Zhang, Li Shen +1
The powerful cooperation of federated learning (FL) and differential privacy~(DP) provides a promising paradigm for the large-scale private clients. However, existing analyses in F…