collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2025

Provable Failure of Language Models in Learning Majority Boolean Logic via Gradient Descent

Bo Chen, Zhenmei Shi, Zhao Song +1

Recent advancements in Transformer-based architectures have led to impressive breakthroughs in natural language processing tasks, with models such as GPT-4, Claude, and Gemini demo…

cs.LG2025

Visual Autoregressive Transformers Must Use Memory

Yang Cao, Xiaoyu Li, Yekun Ke +3

A fundamental challenge in Visual Autoregressive models is the substantial memory overhead required during inference to store previously generated representations. Despite various…

cs.LG2025

Force Matching with Relativistic Constraints: A Physics-Inspired Approach to Stable and Efficient Generative Modeling

Yang Cao, Bo Chen, Xiaoyu Li +5

This paper introduces Force Matching (ForM), a novel framework for generative modeling that represents an initial exploration into leveraging special relativistic mechanics to enha…

cs.LG2024

Circuit Complexity Bounds for RoPE-based Transformer Architecture

Bo Chen, Xiaoyu Li, Yingyu Liang +3

Characterizing the express power of the Transformer architecture is critical to understanding its capacity limits and scaling law. Recent works provide the circuit complexity bound…

cs.LG2024

Bypassing the Exponential Dependency: Looped Transformers Efficiently Learn In-context by Multi-step Gradient Descent

Bo Chen, Xiaoyu Li, Yingyu Liang +2

In-context learning has been recognized as a key factor in the success of Large Language Models (LLMs). It refers to the model's ability to learn patterns on the fly from provided…

cs.LG2024

HSR-Enhanced Sparse Attention Acceleration

Bo Chen, Yingyu Liang, Zhizhou Sha +2

Large Language Models (LLMs) have demonstrated remarkable capabilities across various applications, but their performance on long-context tasks is often limited by the computationa…