collaborators

8 papers

cs.LG2026

Towards Generalizable Reasoning: Group Causal Counterfactual Policy Optimization for LLM Reasoning

Jingyao Wang, Peizheng Guo, Wenwen Qiang +4

Large language models (LLMs) excel at complex tasks with advances in reasoning capabilities. However, existing reward mechanisms remain tightly coupled to final correctness and pay…

cs.CV2026

Test-Time Perturbation Learning with Delayed Feedback for Vision-Language-Action Models

Zehua Zang, Xi Wang, Fuchun Sun +4

Vision-Language-Action models (VLAs) achieve remarkable performance in sequential decision-making but remain fragile to subtle environmental shifts, such as small changes in object…

cs.LG2026

CAMD: Coverage-Aware Multimodal Decoding for Efficient Reasoning of Multimodal Large Language Models

Huijie Guo, Jingyao Wang, Lingyu Si +3

Recent advances in Multimodal Large Language Models (MLLMs) have shown impressive reasoning capabilities across vision-language tasks, yet still face the challenge of compute-diffi…

cs.LG2026

C^2Prompt: Class-aware Client Knowledge Interaction for Federated Continual Learning

Kunlun Xu, Yibo Feng, Jiangmeng Li +2

Federated continual learning (FCL) tackles scenarios of learning from continuously emerging task data across distributed clients, where the key challenge lies in addressing both te…

cs.LG2026

On the Plasticity and Stability for Post-Training Large Language Models

Wenwen Qiang, Ziyin Gu, Jiahuan Zhou +4

Training stability remains a critical bottleneck for Group Relative Policy Optimization (GRPO), often manifesting as a trade-off between reasoning plasticity and general capability…

cs.LG2025

Doubly Debiased Test-Time Prompt Tuning for Vision-Language Models

Fei Song, Yi Li, Rui Wang +3

Test-time prompt tuning for vision-language models has demonstrated impressive generalization capabilities under zero-shot settings. However, tuning the learnable prompts solely ba…