3 papers
cs.AI2026
How Vision Becomes Language: A Layer-wise Information-Theoretic Analysis of Multimodal Reasoning
Hongxuan Wu, Yukun Zhang, Xueqing Zhou
When a multimodal Transformer answers a visual question, is the prediction driven by visual evidence, linguistic reasoning, or genuinely fused cross-modal computation -- and how do…
cs.LG2025
Where to Add PDE Diffusion in Transformers
Yukun Zhang, Xueqing Zhou
Transformers enable powerful content-based global routing via self-attention, but they lack an explicit local geometric prior along the sequence axis. As a result, the placement of…
cs.LG2025
Continuous-Time Attention: PDE-Guided Mechanisms for Long-Sequence Transformers
Yukun Zhang, Xueqing Zhou
We propose a novel framework, Continuous_Time Attention, which infuses partial differential equations (PDEs) into the Transformer's attention mechanism to address the challenges of…