#transformers
20 papers match
TopoFormer: Topology Meets Attention for Graph Learning
Md Joshem Uddin, Astrit Tola, Cuneyt Gurcan Akcora +1
The paper introduces TopoFormer, a framework that converts graph topology into ordered token sequences via a Topo-Scan module and processes them with a Transformer to obtain effici…
ClockRoPE: Random Fourier Rotations for Temporal Routine Modeling
Yiwen Chen, Joshua Ainslie, Krzysztof Choromanski +4
The paper introduces ClockRoPE, a method that uses random Fourier rotations to create periodic position embeddings for transformers, improving temporal routine modeling in sequenti…
The Sparsity Ceiling: Where Spiking Networks Can and Cannot Trade Activity for Energy
Zeyu Wang
The paper studies how much spiking neural networks can lower their firing activity without losing performance, showing that the achievable sparsity depends on the task and architec…
A Compositional Theory of Causally Masked Transformers
Franz Nowak, Ryan Cotterell, Reda Boumasmoud
The paper develops an algebraic framework to characterize what decision problems finite‑precision, causally masked transformers can solve, linking attention mechanisms to memory re…
SEMA: a Scalable and Efficient Mamba like Attention via Token Localization and Averaging
Nhat Thanh Tran, Fanghui Xue, Shuai Zhang +4
The paper introduces SEMA, a new attention mechanism for vision transformers that combines token localization with arithmetic averaging to avoid the dispersion problem of linear at…
Beyond Implicit Force: Evaluating Explicit Force-Torque Proxies in Action Chunking with Transformers
King Hang Wong, Lingqiao Liu, Feras Dayoub
The paper investigates whether explicit joint‑torque signals can replace the implicit force cues present in leader‑follower teleoperation for transformer‑based action‑chunking poli…
Post-Training Pruning for Diffusion Transformers
Chengzhi Hu, Xuewen Liu, Jing Zhang +3
The paper introduces DiT-Pruning, a post‑training pruning method tailored for Diffusion Transformers that uses a new energy‑based saliency metric and clustering‑aware granularity t…
Per-Token Fixed-Point Convergence in Depth-Recurrent Transformers
Joe Logan
The paper investigates how depth‑recurrent transformers reach a per‑token fixed point during repeated application of a weight‑tied core, showing that most tokens converge quickly w…
From Vector Autoregressions to AI-based Time Series Forecasting: A Review
Likai Chen, Weining Wang
The paper reviews recent AI-driven time‑series forecasting methods—including transformers, large pretrained zero‑shot models, and diffusion‑based forecasters—and relates them to tr…
Hierarchical Self-Supervised Representation Learning Framework for Multivariate Time Series Grounded in ECG Analysis
Siwon Kim
The paper introduces ER-JEPA, a lightweight self‑supervised learning framework that builds hierarchical representations for multivariate time‑series data, demonstrated on 12‑lead E…
DeepLoop: Depth Scaling for Looped Transformers
Shuzhen Li, Yifan Zhang, Jiacheng Guo +2
DeepLoop reuses a compact stack of transformer blocks across multiple passes to increase model depth without adding parameters, and introduces new residual scaling rules to keep tr…
Higher Embedding Dimension Creates a Stronger World Model for a Simple Sorting Task
Brady Bhalla, Honglu Fan, Nancy Chen +1
The paper studies how the size of embedding vectors influences the development of internal world models in transformers trained via reinforcement learning to perform bubble‑sort‑st…
Learning in Curved Weight Space:Exponential-Linear Weight Reparameterization for Improved Optimization
Ethan Smith
The paper proposes a new weight reparameterization that combines an exponential and a linear pathway, creating a curved parameter space that enables more proportional updates and s…
Invariant Learning Dynamics of Transformers in Inductive Reasoning Tasks
Tiberiu Musat, Tiago Pimentel, Nicolas Zucchet +1
The paper develops a theoretical framework showing that transformer models learn inductive reasoning tasks by evolving on a low-dimensional invariant manifold, enabling tractable a…
AVQ-Attention: Adaptive Vector-Quantized Attention
Winfried van den dool, Patrick Forré, Amir Habibian +2
The paper introduces Adaptive Vector-Quantized (AVQ) Attention, which dynamically allocates codebook capacity to the most important regions of the key space, preserving O(MN) compl…
Layer-Parallel Inference Reduces Encrypted Nonlinear Depth in Transformers
Ligong Han, Kai Xu, Hao Wang +3
The paper introduces Structured Newton Layer Parallelism (SNLP) to reduce the sequential nonlinear depth of encrypted Transformer inference under fully homomorphic encryption, achi…
Controlling Motion Transfer in Diffusion Transformers via Attention Heads
Sunyoung Jung, Jiwoo Park, Yoonseok Choi +3
The paper studies how individual attention heads in diffusion transformer models handle motion and spatial structure, and introduces a head-aware method to control motion transfer…
Comparative Analysis of GAT and BERT for Human-Like Playtesting
Kleio Fragkedaki, Theodoros Panagiotakopoulos, Matteo Biasielli +1
The paper compares transformer-based (BERT) and graph attention (GAT) models for predicting player behavior in Candy Crush Saga, showing they outperform CNN baselines on complex bo…
Causal Foundation Models with Continuous Treatments
Christopher Stith, Medha Barath, Vahid Balazadeh +2
The paper introduces a causal foundation model that can predict individual treatment-response curves for continuous interventions, using a transformer trained on a synthetic causal…
Sparse Inter-Layer Dependencies of Transformer FFN Neurons
Johannes Knittel, Hanspeter Pfister
The paper introduces a training‑free method to attribute the activation of individual feed‑forward network neurons in Transformers to a small set of upstream neuron activations and…
One search, two signals: results blend meaning (embedding similarity, so papers that never use your words still surface) with keyword matches on titles, abstracts and summaries. Free, no sign-in needed.