4 papers
How Transformers Utilize Multi-Head Attention in In-Context Learning? A Case Study on Sparse Linear Regression
Xingwu Chen, Lei Zhao, Difan Zou
Despite the remarkable success of transformer-based models in various real-world tasks, their underlying mechanisms remain poorly understood. Recent studies have suggested that tra…
GuLu XuanYuan , a biomimetic Transformer that intergrates humanoid MIP, reptile UGV, and bird UAV
Le Chen, Jie Yu, XingWu Chen
This article proposes a multi habitat bio-mimetic robot, named as GuLu XuanYuan.It combines all common types of mobile robots, namely humanoid MIP, unmanned ground vehicle, and unm…
What Can Transformer Learn with Varying Depth? Case Studies on Sequence Learning Tasks
Xingwu Chen, Difan Zou
We study the capabilities of the transformer architecture with varying depth. Specifically, we designed a novel set of sequence learning tasks to systematically evaluate and compre…
Orbits and tsectors in irregular exceptional directions of full-null degenerate singular point
Jun Zhang, Xingwu Chen, Weinian Zhang
Near full-null degenerate singular points of analytic vector fields, asymptotic behaviors of orbits are not given by eigenvectors but totally decided by nonlinearities. Especially,…