7 papers
Decomposing LLM Self-Correction: The Accuracy-Correction Paradox and Error Depth Hypothesis
Yin Li
Large Language Models (LLMs) are widely believed to possess self-correction capabilities, yet recent studies suggest that intrinsic self-correction--where models correct their own…
AI Pangaea: Unifying Intelligence Islands for Adapting Myriad Tasks
Jianlong Chang, Haixin Wang, Zhiyuan Dang +11
The pursuit of artificial general intelligence continuously demands generalization in one model across myriad tasks, even those not seen before. However, current AI models are isol…
Theoretical Analysis of Positional Encodings in Transformer Models: Impact on Expressiveness and Generalization
Yin Li
Positional encodings are a core part of transformer-based models, enabling processing of sequential data without recurrence. This paper presents a theoretical framework to analyze…
Pangu Ultra MoE: How to Train Your Big MoE on Ascend NPUs
Yehui Tang, Yichun Yin, Yaoyuan Wang +71
Sparse large language models (LLMs) with Mixture of Experts (MoE) and close to a trillion parameters are dominating the realm of most capable language models. However, the massive…
PoEmotion: Can AI Utilize Chinese Calligraphy to Express Emotion from Poems?
Tiancheng Liu, Anqi Wang, Xinda Chen +4
This paper presents PoEmotion, an approach to visualizing emotions in poetry with Chinese calligraphy strokes. Traditional textual emotion analysis often lacks emotional resonance…
Air Quality Prediction with A Meteorology-Guided Modality-Decoupled Spatio-Temporal Network
Hang Yin, Yan-Ming Zhang, Jian Xu +3
Air quality prediction plays a crucial role in public health and environmental protection. Accurate air quality prediction is a complex multivariate spatiotemporal problem, that in…