8 papers
Towards Generalizable Reasoning: Group Causal Counterfactual Policy Optimization for LLM Reasoning
Jingyao Wang, Peizheng Guo, Wenwen Qiang +4
Large language models (LLMs) excel at complex tasks with advances in reasoning capabilities. However, existing reward mechanisms remain tightly coupled to final correctness and pay…
Test-Time Perturbation Learning with Delayed Feedback for Vision-Language-Action Models
Zehua Zang, Xi Wang, Fuchun Sun +4
Vision-Language-Action models (VLAs) achieve remarkable performance in sequential decision-making but remain fragile to subtle environmental shifts, such as small changes in object…
CAMD: Coverage-Aware Multimodal Decoding for Efficient Reasoning of Multimodal Large Language Models
Huijie Guo, Jingyao Wang, Lingyu Si +3
Recent advances in Multimodal Large Language Models (MLLMs) have shown impressive reasoning capabilities across vision-language tasks, yet still face the challenge of compute-diffi…
C^2Prompt: Class-aware Client Knowledge Interaction for Federated Continual Learning
Kunlun Xu, Yibo Feng, Jiangmeng Li +2
Federated continual learning (FCL) tackles scenarios of learning from continuously emerging task data across distributed clients, where the key challenge lies in addressing both te…
On the Plasticity and Stability for Post-Training Large Language Models
Wenwen Qiang, Ziyin Gu, Jiahuan Zhou +4
Training stability remains a critical bottleneck for Group Relative Policy Optimization (GRPO), often manifesting as a trade-off between reasoning plasticity and general capability…
Doubly Debiased Test-Time Prompt Tuning for Vision-Language Models
Fei Song, Yi Li, Rui Wang +3
Test-time prompt tuning for vision-language models has demonstrated impressive generalization capabilities under zero-shot settings. However, tuning the learnable prompts solely ba…