#transformers

try —

20 papers match

cs.LG2026

TopoFormer: Topology Meets Attention for Graph Learning

Md Joshem Uddin, Astrit Tola, Cuneyt Gurcan Akcora +1

The paper introduces TopoFormer, a framework that converts graph topology into ordered token sequences via a Topo-Scan module and processes them with a Transformer to obtain effici…

#graph representation learning#topological data analysis#transformers#graph neural networks
cs.LG2026

ClockRoPE: Random Fourier Rotations for Temporal Routine Modeling

Yiwen Chen, Joshua Ainslie, Krzysztof Choromanski +4

The paper introduces ClockRoPE, a method that uses random Fourier rotations to create periodic position embeddings for transformers, improving temporal routine modeling in sequenti…

#transformers#position embeddings#sequential recommendation#temporal modeling
cs.NE2026

The Sparsity Ceiling: Where Spiking Networks Can and Cannot Trade Activity for Energy

Zeyu Wang

The paper studies how much spiking neural networks can lower their firing activity without losing performance, showing that the achievable sparsity depends on the task and architec…

#spiking neural networks#sparsity#energy efficiency#recurrent models
cs.FL2026

A Compositional Theory of Causally Masked Transformers

Franz Nowak, Ryan Cotterell, Reda Boumasmoud

The paper develops an algebraic framework to characterize what decision problems finite‑precision, causally masked transformers can solve, linking attention mechanisms to memory re…

#transformers#expressivity#finite precision#attention mechanisms
cs.CV2026

SEMA: a Scalable and Efficient Mamba like Attention via Token Localization and Averaging

Nhat Thanh Tran, Fanghui Xue, Shuai Zhang +4

The paper introduces SEMA, a new attention mechanism for vision transformers that combines token localization with arithmetic averaging to avoid the dispersion problem of linear at…

#attention mechanisms#transformers#vision models#token localization
cs.RO2026

Beyond Implicit Force: Evaluating Explicit Force-Torque Proxies in Action Chunking with Transformers

King Hang Wong, Lingqiao Liu, Feras Dayoub

The paper investigates whether explicit joint‑torque signals can replace the implicit force cues present in leader‑follower teleoperation for transformer‑based action‑chunking poli…

#contact-rich manipulation#action chunking#transformers#force-torque sensing
cs.CV2026

Post-Training Pruning for Diffusion Transformers

Chengzhi Hu, Xuewen Liu, Jing Zhang +3

The paper introduces DiT-Pruning, a post‑training pruning method tailored for Diffusion Transformers that uses a new energy‑based saliency metric and clustering‑aware granularity t…

#diffusion models#transformers#model pruning#post‑training compression
cs.AI2026

Per-Token Fixed-Point Convergence in Depth-Recurrent Transformers

Joe Logan

The paper investigates how depth‑recurrent transformers reach a per‑token fixed point during repeated application of a weight‑tied core, showing that most tokens converge quickly w…

#transformers#depth recurrence#fixed point convergence#token-level analysis
econ.EM2026

From Vector Autoregressions to AI-based Time Series Forecasting: A Review

Likai Chen, Weining Wang

The paper reviews recent AI-driven time‑series forecasting methods—including transformers, large pretrained zero‑shot models, and diffusion‑based forecasters—and relates them to tr…

#time series forecasting#transformers#large pretrained models#diffusion models
cs.LG2026

Hierarchical Self-Supervised Representation Learning Framework for Multivariate Time Series Grounded in ECG Analysis

Siwon Kim

The paper introduces ER-JEPA, a lightweight self‑supervised learning framework that builds hierarchical representations for multivariate time‑series data, demonstrated on 12‑lead E…

#self-supervised learning#time series#electrocardiogram#hierarchical models
cs.LG2026

DeepLoop: Depth Scaling for Looped Transformers

Shuzhen Li, Yifan Zhang, Jiacheng Guo +2

DeepLoop reuses a compact stack of transformer blocks across multiple passes to increase model depth without adding parameters, and introduces new residual scaling rules to keep tr…

#transformers#depth scaling#residual scaling#looped architecture
cs.LG2026

Higher Embedding Dimension Creates a Stronger World Model for a Simple Sorting Task

Brady Bhalla, Honglu Fan, Nancy Chen +1

The paper studies how the size of embedding vectors influences the development of internal world models in transformers trained via reinforcement learning to perform bubble‑sort‑st…

#transformers#embedding dimension#reinforcement learning#algorithmic reasoning
cs.LG2026

Learning in Curved Weight Space:Exponential-Linear Weight Reparameterization for Improved Optimization

Ethan Smith

The paper proposes a new weight reparameterization that combines an exponential and a linear pathway, creating a curved parameter space that enables more proportional updates and s…

#weight reparameterization#optimization#neural networks#transformers
cs.LG2026

Invariant Learning Dynamics of Transformers in Inductive Reasoning Tasks

Tiberiu Musat, Tiago Pimentel, Nicolas Zucchet +1

The paper develops a theoretical framework showing that transformer models learn inductive reasoning tasks by evolving on a low-dimensional invariant manifold, enabling tractable a…

#transformers#inductive reasoning#learning dynamics#attention
cs.LG2026

AVQ-Attention: Adaptive Vector-Quantized Attention

Winfried van den dool, Patrick Forré, Amir Habibian +2

The paper introduces Adaptive Vector-Quantized (AVQ) Attention, which dynamically allocates codebook capacity to the most important regions of the key space, preserving O(MN) compl…

#transformers#attention mechanisms#vector quantization#efficient inference
cs.LG2026

Layer-Parallel Inference Reduces Encrypted Nonlinear Depth in Transformers

Ligong Han, Kai Xu, Hao Wang +3

The paper introduces Structured Newton Layer Parallelism (SNLP) to reduce the sequential nonlinear depth of encrypted Transformer inference under fully homomorphic encryption, achi…

#encrypted inference#transformers#fully homomorphic encryption#layer parallelism
cs.CV2026

Controlling Motion Transfer in Diffusion Transformers via Attention Heads

Sunyoung Jung, Jiwoo Park, Yoonseok Choi +3

The paper studies how individual attention heads in diffusion transformer models handle motion and spatial structure, and introduces a head-aware method to control motion transfer…

#video generation#motion transfer#diffusion models#transformers
cs.AI2026

Comparative Analysis of GAT and BERT for Human-Like Playtesting

Kleio Fragkedaki, Theodoros Panagiotakopoulos, Matteo Biasielli +1

The paper compares transformer-based (BERT) and graph attention (GAT) models for predicting player behavior in Candy Crush Saga, showing they outperform CNN baselines on complex bo…

#playtesting#game AI#graph neural networks#transformers
cs.LG2026

Causal Foundation Models with Continuous Treatments

Christopher Stith, Medha Barath, Vahid Balazadeh +2

The paper introduces a causal foundation model that can predict individual treatment-response curves for continuous interventions, using a transformer trained on a synthetic causal…

#causal inference#continuous treatment#foundation models#meta-learning
cs.LG2026

Sparse Inter-Layer Dependencies of Transformer FFN Neurons

Johannes Knittel, Hanspeter Pfister

The paper introduces a training‑free method to attribute the activation of individual feed‑forward network neurons in Transformers to a small set of upstream neuron activations and…

#transformers#feedforward networks#neuron interpretability#sparsity

One search, two signals: results blend meaning (embedding similarity, so papers that never use your words still surface) with keyword matches on titles, abstracts and summaries. Free, no sign-in needed.