collaborators

11 papers

cs.LG2026

Linear and Neural Dueling Bandits with Delayed Feedback

Xiangyi Wang, Pingchen Lu, Jie Mao +4

Contextual dueling bandits form a cornerstone of preference-based decision-making, with critical applications in recommender systems and large language model alignment. However, st…

cs.CL2026

FERA: Uncertainty-Aware Federated Reasoning for Large Language Models

Ruhan Wang, Chengkai Huang, Zhiyong Wang +6

Large language models (LLMs) exhibit strong reasoning capabilities when guided by high-quality demonstrations, yet such data is often distributed across organizations that cannot c…

cs.RO2026

Can Vision-Language-Action Models Learn from Real-World Data Continually without Forgetting?

Jiarun Zhu, Yijun Hong, Xiaoquan Sun +7

Vision-Language-Action (VLA) models provide a promising foundation for general-purpose robotics, yet their real-world deployment demands the ability to continually acquire new skil…

cs.AI2026

Workflow-R1: Group Sub-sequence Policy Optimization for Multi-turn Workflow Construction

Mingze Kong, Zikun Qu, Zhongquan Zhou +7

The rapid evolution of agentic workflows has demonstrated strong performance of LLM-based agents in addressing complex reasoning tasks. However, existing workflow optimization meth…

cs.CL2025

Self-Reflective Generation at Test Time

Jian Mu, Qixin Zhang, Zhiyong Wang +5

Large language models (LLMs) increasingly solve complex reasoning tasks via long chain-of-thought, but their forward-only autoregressive generation process is fragile; early token…

cs.LG2025

FedPOB: Sample-Efficient Federated Prompt Optimization via Bandits

Pingchen Lu, Zhi Hong, Zhiwei Shang +6

The performance of large language models (LLMs) is highly sensitive to the input prompt, making prompt optimization a critical task. However, real-world application is hindered by…