49 papers
VISTA: Test-Time Compositional Alignment for Visual Autoregressive Generation
Hossein Shahabadi, Niki Sepasian, Mahdieh Soleymani Baghshah
Visual autoregressive (VAR) models have emerged as a fast, high-quality alternative to diffusion for text-to-image generation, but like diffusion models they exhibit persistent com…
Efficient Adversarial Attacks on High-dimensional Offline Bandits
Seyed Mohammad Hadi Hosseini, Amir Najafi, Mahdieh Soleymani Baghshah
Bandit algorithms have recently emerged as a powerful tool for evaluating machine learning models, including generative image models and large language models, by efficiently ident…
SUSD: Structured Unsupervised Skill Discovery through State Factorization
Seyed Mohammad Hadi Hosseini, Mahdieh Soleymani Baghshah
Unsupervised Skill Discovery (USD) aims to autonomously learn a diverse set of skills without relying on extrinsic rewards. One of the most common USD approaches is to maximize the…
Causal Attribution via Activation Patching
Amirmohammad Izadi, Mohammadali Banayeeanzade, Alireza Mirrokni +4
Attribution methods for Vision Transformers (ViTs) aim to identify image regions that influence model predictions, but producing faithful and well-localized attributions remains ch…
Persian MusicGen: A Large-Scale Dataset and Culturally-Aware Generative Model for Persian Music
Mohammad Hossein Sameti, Diba Hadi Esfangereh, Sepehr Harfi Moridani +2
Persian music, with its unique tonalities, modal systems (Dastgah), and rhythmic structures, presents significant challenges for music generation models trained primarily on Wester…
Mechanistic Interpretability of Large-Scale Counting in LLMs through a System-2 Strategy
Hosein Hasani, Mohammadali Banayeeanzade, Ali Nafisi +5
Large language models (LLMs), despite strong performance on complex mathematical problems, exhibit systematic limitations in counting tasks. This issue arises from the architectura…