Publications (65)
Words are Malleable: Computing Semantic Shifts in Political and Media Discourse
Hosein Azarbonyad, Mostafa Dehghani, Kaspar Beelen +3
Hierarchical Re-estimation of Topic Models for Measuring Topical Diversity
Hosein Azarbonyad, Mostafa Dehghani, Tom Kenter +3
Exploring the Limits of Large Scale Pre-training
Samira Abnar, Mostafa Dehghani, Behnam Neyshabur +1
PaLI-X: On Scaling up a Multilingual Vision and Language Model
Xi Chen, Josip Djolonga, Piotr Padlewski +40
Learning to Learn from Weak Supervision by Full Supervision
Mostafa Dehghani, Aliaksei Severyn, Sascha Rothe +1
End-to-End Spatio-Temporal Action Localisation with Video Transformers
Alexey Gritsenko, Xuehan Xiong, Josip Djolonga +5
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Petko Georgiev, Ving Ian Lei +1132
Parameter-efficient Multi-task Fine-tuning for Transformers via Shared Hypernetworks
Rabeeh Karimi Mahabadi, Sebastian Ruder, Mostafa Dehghani +1
Low-Rank Adaptation for Multilingual Summarization: An Empirical Study
Chenxi Whitehouse, Fantine Huot, Jasmijn Bastings +3
Fidelity-Weighted Learning
Mostafa Dehghani, Arash Mehrjou, Stephan Gouws +2
The Benchmark Lottery
Mostafa Dehghani, Yi Tay, Alexey A. Gritsenko +5
IDF++: Analyzing and Improving Integer Discrete Flows for Lossless Compression
Rianne van den Berg, Alexey A. Gritsenko, Mostafa Dehghani +2
UL2: Unifying Language Learning Paradigms
Yi Tay, Mostafa Dehghani, Vinh Q. Tran +11
PolyViT: Co-training Vision Transformers on Images, Videos and Audio
Valerii Likhosherstov, Anurag Arnab, Krzysztof Choromanski +4
Patch n' Pack: NaViT, a Vision Transformer for any Aspect Ratio and Resolution
Mostafa Dehghani, Basil Mustafa, Josip Djolonga +12
Universal Transformers
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals +2
TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?
Michael S. Ryoo, AJ Piergiovanni, Anurag Arnab +2
OmniNet: Omnidirectional Representations from Transformers
Yi Tay, Mostafa Dehghani, Vamsi Aribandi +6
$Î$-DARTS: Mitigating Performance Collapse by Harmonizing Operation Selection among Cells
Sajad Movahedi, Melika Adabinejad, Ayyoob Imani +4
VUT: Versatile UI Transformer for Multi-Modal Multi-Task User Interface Modeling
Yang Li, Gang Li, Xin Zhou +2
Learning to Attend, Copy, and Generate for Session-Based Query Suggestion
Mostafa Dehghani, Sascha Rothe, Enrique Alfonseca +1
Transformer Memory as a Differentiable Search Index
Yi Tay, Vinh Q. Tran, Mostafa Dehghani +10
Fractal Patterns May Illuminate the Success of Next-Token Prediction
Ibrahim Alabdulmohsin, Vinh Q. Tran, Mostafa Dehghani
Dual PatchNorm
Manoj Kumar, Mostafa Dehghani, Neil Houlsby
Retrieval-Enhanced Machine Learning
Hamed Zamani, Fernando Diaz, Mostafa Dehghani +2
The Impact of Group Membership Bias on the Quality and Fairness of Exposure in Ranking
Ali Vardasbi, Maarten de Rijke, Fernando Diaz +1
MetNet: A Neural Weather Model for Precipitation Forecasting
Casper Kaae Sønderby, Lasse Espeholt, Jonathan Heek +6
Gemini: A Family of Highly Capable Multimodal Models
Gemini Team, Rohan Anil, Sebastian Borgeaud +1340
Karyotype AI for Precision Oncology
Zahra Shamsi, Isaac Reid, Drew Bryant +13
Beyond Transfer Learning: Co-finetuning for Action Localisation
Anurag Arnab, Xuehan Xiong, Alexey Gritsenko +6
Discrete Representations Strengthen Vision Transformer Robustness
Chengzhi Mao, Lu Jiang, Mostafa Dehghani +3
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov +9
Learning to Rank from Samples of Variable Quality
Mostafa Dehghani, Jaap Kamps
Transferring Inductive Biases through Knowledge Distillation
Samira Abnar, Mostafa Dehghani, Willem Zuidema
Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers
Yi Tay, Mostafa Dehghani, Jinfeng Rao +7
Are Pre-trained Convolutions Better than Pre-trained Transformers?
Yi Tay, Mostafa Dehghani, Jai Gupta +4
Neural Networks for Information Retrieval
Tom Kenter, Alexey Borisov, Christophe Van Gysel +3
PaLM 2 Technical Report
Rohan Anil, Andrew M. Dai, Orhan Firat +125
Gradual Domain Adaptation in the Wild:When Intermediate Distributions are Absent
Samira Abnar, Rianne van den Berg, Golnaz Ghiasi +3
Frozen Feature Augmentation for Few-Shot Image Classification
Andreas Bär, Neil Houlsby, Mostafa Dehghani +1
Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints
Aran Komatsuzaki, Joan Puigcerver, James Lee-Thorp +6
HiTR: Hierarchical Topic Model Re-estimation for Measuring Topical Diversity of Documents
Hosein Azarbonyad, Mostafa Dehghani, Tom Kenter +3
Avoiding Your Teacher's Mistakes: Training Neural Networks with Controlled Weak Supervision
Mostafa Dehghani, Aliaksei Severyn, Sascha Rothe +1
Simple Open-Vocabulary Object Detection with Vision Transformers
Matthias Minderer, Alexey Gritsenko, Austin Stone +11
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431
Efficient Transformers: A Survey
Yi Tay, Mostafa Dehghani, Dara Bahri +1
Neural Ranking Models with Weak Supervision
Mostafa Dehghani, Hamed Zamani, Aliaksei Severyn +2
Scaling Laws vs Model Architectures: How does Inductive Bias Influence Scaling?
Yi Tay, Mostafa Dehghani, Samira Abnar +7
Adaptive Computation with Elastic Input Sequence
Fuzhao Xue, Valerii Likhosherstov, Anurag Arnab +3
Scaling Vision Transformers to 22 Billion Parameters
Mostafa Dehghani, Josip Djolonga, Basil Mustafa +39
How (not) to ensemble LVLMs for VQA
Lisa Alazraki, Lluis Castrejon, Mostafa Dehghani +3
Scaling Instruction-Finetuned Language Models
Hyung Won Chung, Le Hou, Shayne Longpre +32
Generalized Group Profiling for Content Customization
Mostafa Dehghani, Hosein Azarbonyad, Jaap Kamps +1
On Search Powered Navigation
Mostafa Dehghani, Glorianna Jagfeld, Hosein Azarbonyad +3
Neural Networks for Information Retrieval
Tom Kenter, Alexey Borisov, Christophe Van Gysel +3
Confident Adaptive Language Modeling
Tal Schuster, Adam Fisch, Jai Gupta +5
SCENIC: A JAX Library for Computer Vision Research and Beyond
Mostafa Dehghani, Alexey Gritsenko, Anurag Arnab +2
Share your Model instead of your Data: Privacy Preserving Mimic Learning for Ranking
Mostafa Dehghani, Hosein Azarbonyad, Jaap Kamps +1
DSI++: Updating Transformer Memory with New Documents
Sanket Vaibhav Mehta, Jai Gupta, Yi Tay +6
Intersection of Parallels as an Early Stopping Criterion
Ali Vardasbi, Maarten de Rijke, Mostafa Dehghani
The Efficiency Misnomer
Mostafa Dehghani, Anurag Arnab, Lucas Beyer +2
On Horizontal and Vertical Separation in Hierarchical Text Classification
Mostafa Dehghani, Hosein Azarbonyad, Jaap Kamps +1
ViViT: A Video Vision Transformer
Anurag Arnab, Mostafa Dehghani, Georg Heigold +3
Transcending Scaling Laws with 0.1% Extra Compute
Yi Tay, Jason Wei, Hyung Won Chung +13
Long Range Arena: A Benchmark for Efficient Transformers
Yi Tay, Mostafa Dehghani, Samira Abnar +7