7 papers
Rethinking Structure Preservation in Text-Guided Image Editing with Visual Autoregressive Models
Tao Xia, Jiawei Liu, Yukun Zhang +3
Visual autoregressive (VAR) models have recently emerged as a promising family of generative models, enabling a wide range of downstream vision tasks such as text-guided image edit…
How Vision Becomes Language: A Layer-wise Information-Theoretic Analysis of Multimodal Reasoning
Hongxuan Wu, Yukun Zhang, Xueqing Zhou
When a multimodal Transformer answers a visual question, is the prediction driven by visual evidence, linguistic reasoning, or genuinely fused cross-modal computation -- and how do…
Where to Add PDE Diffusion in Transformers
Yukun Zhang, Xueqing Zhou
Transformers enable powerful content-based global routing via self-attention, but they lack an explicit local geometric prior along the sequence axis. As a result, the placement of…
Integrating Domain Knowledge for Financial QA: A Multi-Retriever RAG Approach with LLMs
Yukun Zhang, Stefan Elbl Droguett, Samyak Jain
This research project addresses the errors of financial numerical reasoning Question Answering (QA) tasks due to the lack of domain knowledge in finance. Despite recent advances in…
Understanding Transformer Architecture through Continuous Dynamics: A Partial Differential Equation Perspective
Yukun Zhang, Xueqing Zhou
The Transformer architecture has revolutionized artificial intelligence, yet a principled theoretical understanding of its internal mechanisms remains elusive. This paper introduce…
Continuous-Time Attention: PDE-Guided Mechanisms for Long-Sequence Transformers
Yukun Zhang, Xueqing Zhou
We propose a novel framework, Continuous_Time Attention, which infuses partial differential equations (PDEs) into the Transformer's attention mechanism to address the challenges of…