5 papers
Distilling LLM Reasoning into an Interpretable Policy Tree for Human-AI Collaboration
Beiwen Zhang, Yongheng Liang, Guowei Zou +2
Constructing efficient and reliable policies to assist humans is indispensable for human-AI collaboration. Existing methods mainly follow two lines of work. Most prior work relies…
CoFlow: Coordinated Few-Step Flow for Offline Multi-Agent Decision Making
Guowei Zou, Haitao Wang, Beiwen Zhang +2
Generative models have emerged as a promising paradigm for offline multi-agent reinforcement learning (MARL), but existing approaches require many iterative sampling steps. Recent…
PACT: Phenotype-Aware Contrastive Team Representation for Multi-Phenotype Grouped Ad Hoc Teamwork
Beiwen Zhang, Yongheng Liang, Hejun Wu +3
Learning to collaborate with various unfamiliar teammates poses a great challenge in the domain of multi-agent systems. Existing ad hoc teamwork methods typically drive controlled…
Towards Analyzing and Understanding the Limitations of VAPO: A Theoretical Perspective
Jintian Shao, Yiming Cheng, Hongyi Huang +4
The VAPO framework has demonstrated significant empirical success in enhancing the efficiency and reliability of reinforcement learning for long chain-of-thought (CoT) reasoning ta…
ComplexFormer: Disruptively Advancing Transformer Inference Ability via Head-Specific Complex Vector Attention
Jintian Shao, Hongyi Huang, Jiayi Wu +4
Transformer models rely on self-attention to capture token dependencies but face challenges in effectively integrating positional information while allowing multi-head attention (M…