4 papers
Recursively Summarizing Enables Long-Term Dialogue Memory in Large Language Models
Qingyue Wang, Yanhe Fu, Yanan Cao +3
Recently, large language models (LLMs), such as GPT-4, stand out remarkable conversational abilities, enabling them to engage in dynamic and contextually relevant dialogues across…
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism
Zheng Zhang, Donglin Yang, Yaqi Xia +4
Recently, Mixture-of-Experts (MoE) has become one of the most popular techniques to scale pre-trained models to extraordinarily large sizes. Dynamic activation of experts allows fo…
A Comprehensive Survey of Data Augmentation in Visual Reinforcement Learning
Guozheng Ma, Zhen Wang, Zhecheng Yuan +3
Visual reinforcement learning (RL), which makes decisions directly from high-dimensional visual inputs, has demonstrated significant potential in various domains. However, deployin…
Error Analysis Prompting Enables Human-Like Translation Evaluation in Large Language Models
Qingyu Lu, Baopu Qiu, Liang Ding +3
Generative large language models (LLMs), e.g., ChatGPT, have demonstrated remarkable proficiency across several NLP tasks, such as machine translation, text summarization. Recent r…