8 papers
LadderSym: A Multimodal Interleaved Transformer for Music Practice Error Detection
Benjamin Shiue-Hal Chou, Purvish Jajal, Nick John Eliopoulos +4
Music learners can greatly benefit from tools that accurately detect errors in their practice. Existing approaches typically compare audio recordings to music scores using heuristi…
Inference-Time Alignment of Diffusion Models via Evolutionary Algorithms
Purvish Jajal, Nick John Eliopoulos, Benjamin Shiue-Hal Chou +3
Diffusion models are state-of-the-art generative models, yet their samples often fail to satisfy application objectives such as safety constraints or domain-specific validity. Exis…
AdaPerceiver: Transformers with Adaptive Width, Depth, and Tokens
Purvish Jajal, Nick John Eliopoulos, Benjamin Shiue-Hal Chou +3
Modern transformer architectures achieve remarkable performance across tasks and domains but remain rigid in how they allocate computation at inference time. Real-world deployment…
Token Turing Machines are Efficient Vision Models
Purvish Jajal, Nick John Eliopoulos, Benjamin Shiue-Hal Chou +3
We propose Vision Token Turing Machines (ViTTM), an efficient, low-latency, memory-augmented Vision Transformer (ViT). Our approach builds on Neural Turing Machines and Token Turin…
Detecting Music Performance Errors with Transformers
Benjamin Shiue-Hal Chou, Purvish Jajal, Nicholas John Eliopoulos +6
Beginner musicians often struggle to identify specific errors in their performances, such as playing incorrect notes or rhythms. There are two limitations in existing tools for mus…
Pruning One More Token is Enough: Leveraging Latency-Workload Non-Linearities for Vision Transformers on the Edge
Nick John Eliopoulos, Purvish Jajal, James C. Davis +3
This paper investigates how to efficiently deploy vision transformers on edge devices for small workloads. Recent methods reduce the latency of transformer neural networks by remov…