4 papers
How Vision Becomes Language: A Layer-wise Information-Theoretic Analysis of Multimodal Reasoning
Hongxuan Wu, Yukun Zhang, Xueqing Zhou
When a multimodal Transformer answers a visual question, is the prediction driven by visual evidence, linguistic reasoning, or genuinely fused cross-modal computation -- and how do…
Where to Add PDE Diffusion in Transformers
Yukun Zhang, Xueqing Zhou
Transformers enable powerful content-based global routing via self-attention, but they lack an explicit local geometric prior along the sequence axis. As a result, the placement of…
Understanding Transformer Architecture through Continuous Dynamics: A Partial Differential Equation Perspective
Yukun Zhang, Xueqing Zhou
The Transformer architecture has revolutionized artificial intelligence, yet a principled theoretical understanding of its internal mechanisms remains elusive. This paper introduce…
Continuous-Time Attention: PDE-Guided Mechanisms for Long-Sequence Transformers
Yukun Zhang, Xueqing Zhou
We propose a novel framework, Continuous_Time Attention, which infuses partial differential equations (PDEs) into the Transformer's attention mechanism to address the challenges of…