11 papers
T-POP: Test-Time Personalization with Online Preference Feedback
Zikun Qu, Min Zhang, Mingze Kong +7
Personalizing large language models (LLMs) to individual user preferences is a critical step beyond generating generically helpful responses. However, current personalization metho…
Self-Reflective Generation at Test Time
Jian Mu, Qixin Zhang, Zhiyong Wang +5
Large language models (LLMs) increasingly solve complex reasoning tasks via long chain-of-thought, but their forward-only autoregressive generation process is fragile; early token…
Can Vision-Language-Action Models Learn from Real-World Data Continually without Forgetting?
Jiarun Zhu, Yijun Hong, Xiaoquan Sun +7
Vision-Language-Action (VLA) models provide a promising foundation for general-purpose robotics, yet their real-world deployment demands the ability to continually acquire new skil…
Linear and Neural Dueling Bandits with Delayed Feedback
Xiangyi Wang, Pingchen Lu, Jie Mao +4
Contextual dueling bandits form a cornerstone of preference-based decision-making, with critical applications in recommender systems and large language model alignment. However, st…
FERA: Uncertainty-Aware Federated Reasoning for Large Language Models
Ruhan Wang, Chengkai Huang, Zhiyong Wang +6
Large language models (LLMs) exhibit strong reasoning capabilities when guided by high-quality demonstrations, yet such data is often distributed across organizations that cannot c…
Workflow-R1: Group Sub-sequence Policy Optimization for Multi-turn Workflow Construction
Mingze Kong, Zikun Qu, Zhongquan Zhou +7
The rapid evolution of agentic workflows has demonstrated strong performance of LLM-based agents in addressing complex reasoning tasks. However, existing workflow optimization meth…