48 citations · 51 across the 5 of their papers we have counts for
5 papers
LongVQ: Long Sequence Modeling with Vector Quantization on Structured Memory
Zicheng Liu, Li Wang, Siyuan Li +3
Transformer models have been successful in various sequence processing tasks, but the self-attention mechanism's computational cost limits its practicality for long sequences. Alth…
Rethinking Memory and Communication Cost for Efficient Large Language Model Training
Chan Wu, Hanxiao Zhang, Lin Ju +8
Recently, various distributed strategies for large language model training have been proposed. However, these methods provided limited solutions for the trade-off between memory co…
Robust Visual Imitation Learning with Inverse Dynamics Representations
Siyuan Li, Xun Wang, Rongchang Zuo +5
Imitation learning (IL) has achieved considerable success in solving complex sequential decision-making problems. However, current IL methods mainly assume that the environment for…
Behavior Contrastive Learning for Unsupervised Skill Discovery
Rushuai Yang, Chenjia Bai, Hongyi Guo +5
In reinforcement learning, unsupervised skill discovery aims to learn diverse skills without extrinsic rewards. Previous methods discover skills by maximizing the mutual informatio…
InstructUIE: Multi-task Instruction Tuning for Unified Information Extraction
Xiao Wang, Weikang Zhou, Can Zu +11
Large language models have unlocked strong multi-task capabilities from reading instructive prompts. However, recent studies have shown that existing large models still have diffic…