7 papers
DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling
Yixiao Qian, Song Chen, Pengkai Wang +3
Modern language models are built primarily from Transformers, recurrent models, and their hybrid architectures. Transformers rely on token-level attention memories, while recurrent…
A Continuous-Time Analysis of Smoothed Matrix-Polar Spectral Gradient Flows for Muon-Type Optimization
Jinlin Liu, Song Chen, Jiaxu Liu +1
This paper studies smoothed matrix-polar spectral gradient flows for unconstrained matrix-valued optimization.The canonical polar-factor map loses smoothness at rank-deficient matr…
DSSMs: State Space Models with Explicit Memory via Delay Differential Equations
Yixiao Qian, Song Chen, Jiaxu Liu +2
State Space Models (SSMs) have emerged as a powerful paradigm for efficient long-sequence modeling, offering parallel training and fast linear-time recurrent inference. However, li…
A Backstepping Framework for Unconstrained Accelerated Optimization Algorithms
Song Chen, Jiaxu Liu, Chao Xu
This paper introduces a control-theoretic perspective on unconstrained optimization algorithms using the backstepping methods. We model the optimization process as an augmented str…
Swarm-Inspired Generation of Collective Behaviors in Graph Dynamical Systems
Ji Chen, Song Chen, Chengzhang Gong +2
Collective behavior arises when locally interacting units produce coordinated global organization, from synchronization in dynamical systems to task-relevant information flow on gr…
Distributed physics-informed neural networks via domain decomposition for fast flow reconstruction
Yixiao Qian, Jiaxu Liu, Zewei Xia +3
Physics-Informed Neural Networks (PINNs) offer a powerful paradigm for flow reconstruction, seamlessly integrating sparse velocity measurements with the governing Navier-Stokes equ…