collaborators

11 papers

cs.LG2026

T-POP: Test-Time Personalization with Online Preference Feedback

Zikun Qu, Min Zhang, Mingze Kong +7

Personalizing large language models (LLMs) to individual user preferences is a critical step beyond generating generically helpful responses. However, current personalization metho…

cs.CL2026

Self-Reflective Generation at Test Time

Jian Mu, Qixin Zhang, Zhiyong Wang +5

Large language models (LLMs) increasingly solve complex reasoning tasks via long chain-of-thought, but their forward-only autoregressive generation process is fragile; early token…

cs.RO2026

Can Vision-Language-Action Models Learn from Real-World Data Continually without Forgetting?

Jiarun Zhu, Yijun Hong, Xiaoquan Sun +7

Vision-Language-Action (VLA) models provide a promising foundation for general-purpose robotics, yet their real-world deployment demands the ability to continually acquire new skil…

cs.LG2026

Linear and Neural Dueling Bandits with Delayed Feedback

Xiangyi Wang, Pingchen Lu, Jie Mao +4

Contextual dueling bandits form a cornerstone of preference-based decision-making, with critical applications in recommender systems and large language model alignment. However, st…

cs.CL2026

FERA: Uncertainty-Aware Federated Reasoning for Large Language Models

Ruhan Wang, Chengkai Huang, Zhiyong Wang +6

Large language models (LLMs) exhibit strong reasoning capabilities when guided by high-quality demonstrations, yet such data is often distributed across organizations that cannot c…

cs.AI2026

Workflow-R1: Group Sub-sequence Policy Optimization for Multi-turn Workflow Construction

Mingze Kong, Zikun Qu, Zhongquan Zhou +7

The rapid evolution of agentic workflows has demonstrated strong performance of LLM-based agents in addressing complex reasoning tasks. However, existing workflow optimization meth…