activity
20242026
collaborators

7 papers

cs.LG2026

Visual Autoregressive Transformers Must Use Memory

Yang Cao, Xiaoyu Li, Yekun Ke +3

A fundamental challenge in Visual Autoregressive models is the substantial memory overhead required during inference to store previously generated representations. Despite various…

cs.LG2025

Provable Failure of Language Models in Learning Majority Boolean Logic via Gradient Descent

Bo Chen, Zhenmei Shi, Zhao Song +1

Recent advancements in Transformer-based architectures have led to impressive breakthroughs in natural language processing tasks, with models such as GPT-4, Claude, and Gemini demo…

cs.LG2025

Bypassing the Exponential Dependency: Looped Transformers Efficiently Learn In-context by Multi-step Gradient Descent

Bo Chen, Xiaoyu Li, Yingyu Liang +2

In-context learning has been recognized as a key factor in the success of Large Language Models (LLMs). It refers to the model's ability to learn patterns on the fly from provided…

cs.LG2025

HSR-Enhanced Sparse Attention Acceleration

Bo Chen, Yingyu Liang, Zhizhou Sha +2

Large Language Models (LLMs) have demonstrated remarkable capabilities across various applications, but their performance on long-context tasks is often limited by the computationa…

cs.LG2025

Force Matching with Relativistic Constraints: A Physics-Inspired Approach to Stable and Efficient Generative Modeling

Yang Cao, Bo Chen, Xiaoyu Li +5

This paper introduces Force Matching (ForM), a novel framework for generative modeling that represents an initial exploration into leveraging special relativistic mechanics to enha…

cs.CV2025

High-Order Matching for One-Step Shortcut Diffusion Models

Bo Chen, Chengyue Gong, Xiaoyu Li +5

One-step shortcut diffusion models [Frans, Hafner, Levine and Abbeel, ICLR 2025] have shown potential in vision generation, but their reliance on first-order trajectory supervision…