4 papers
TEON: Tensorized Orthonormalization Beyond Layer-Wise Muon for Large Language Model Pre-Training
Ruijie Zhang, Yequan Zhao, Ziyue Liu +5
The Muon optimizer has demonstrated strong empirical performance in pre-training large language models by performing matrix-level gradient (or momentum) orthogonalization in each l…
QueueEDIT: Structural Self-Correction for Sequential Model Editing in LLMs
Taolin Zhang, Haidong Kang, Dongyang Li +3
Recently, large language models (LLMs) have demonstrated impressive results but still suffer from hallucinations. Model editing has been proposed to correct factual inaccuracies in…
BELLE: A Bi-Level Multi-Agent Reasoning Framework for Multi-Hop Question Answering
Taolin Zhang, Dongyang Li, Qizhou Chen +2
Multi-hop question answering (QA) involves finding multiple relevant passages and performing step-by-step reasoning to answer complex questions. Previous works on multi-hop QA empl…
Concept Based Continuous Prompts for Interpretable Text Classification
Qian Chen, Dongyang Li, Xiaofeng He
Continuous prompts have become widely adopted for augmenting performance across a wide range of natural language tasks. However, the underlying mechanism of this enhancement remain…