NewEvery arXiv paper, its researchers & institutions — mapped.
papers

Publications (65)

cs.CL2017

Words are Malleable: Computing Semantic Shifts in Political and Media Discourse

Hosein Azarbonyad, Mostafa Dehghani, Kaspar Beelen +3

cs.IR2017

Hierarchical Re-estimation of Topic Models for Measuring Topical Diversity

Hosein Azarbonyad, Mostafa Dehghani, Tom Kenter +3

cs.LG2021

Exploring the Limits of Large Scale Pre-training

Samira Abnar, Mostafa Dehghani, Behnam Neyshabur +1

cs.CV2023

PaLI-X: On Scaling up a Multilingual Vision and Language Model

Xi Chen, Josip Djolonga, Piotr Padlewski +40

stat.ML2017

Learning to Learn from Weak Supervision by Full Supervision

Mostafa Dehghani, Aliaksei Severyn, Sascha Rothe +1

cs.CV2023

End-to-End Spatio-Temporal Action Localisation with Video Transformers

Alexey Gritsenko, Xuehan Xiong, Josip Djolonga +5

cs.CL2024

Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Gemini Team, Petko Georgiev, Ving Ian Lei +1132

cs.CL2021

Parameter-efficient Multi-task Fine-tuning for Transformers via Shared Hypernetworks

Rabeeh Karimi Mahabadi, Sebastian Ruder, Mostafa Dehghani +1

cs.CL2024

Low-Rank Adaptation for Multilingual Summarization: An Empirical Study

Chenxi Whitehouse, Fantine Huot, Jasmijn Bastings +3

cs.LG2018

Fidelity-Weighted Learning

Mostafa Dehghani, Arash Mehrjou, Stephan Gouws +2

cs.LG2021

The Benchmark Lottery

Mostafa Dehghani, Yi Tay, Alexey A. Gritsenko +5

cs.LG2021

IDF++: Analyzing and Improving Integer Discrete Flows for Lossless Compression

Rianne van den Berg, Alexey A. Gritsenko, Mostafa Dehghani +2

cs.CL2023

UL2: Unifying Language Learning Paradigms

Yi Tay, Mostafa Dehghani, Vinh Q. Tran +11

cs.CV2021

PolyViT: Co-training Vision Transformers on Images, Videos and Audio

Valerii Likhosherstov, Anurag Arnab, Krzysztof Choromanski +4

cs.CV2023

Patch n' Pack: NaViT, a Vision Transformer for any Aspect Ratio and Resolution

Mostafa Dehghani, Basil Mustafa, Josip Djolonga +12

cs.CL2019

Universal Transformers

Mostafa Dehghani, Stephan Gouws, Oriol Vinyals +2

cs.CV2022

TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Michael S. Ryoo, AJ Piergiovanni, Anurag Arnab +2

cs.CV2021

OmniNet: Omnidirectional Representations from Transformers

Yi Tay, Mostafa Dehghani, Vamsi Aribandi +6

cs.LG2023

$Λ$-DARTS: Mitigating Performance Collapse by Harmonizing Operation Selection among Cells

Sajad Movahedi, Melika Adabinejad, Ayyoob Imani +4

cs.CV2021

VUT: Versatile UI Transformer for Multi-Modal Multi-Task User Interface Modeling

Yang Li, Gang Li, Xin Zhou +2

cs.IR2017

Learning to Attend, Copy, and Generate for Session-Based Query Suggestion

Mostafa Dehghani, Sascha Rothe, Enrique Alfonseca +1

cs.CL2022

Transformer Memory as a Differentiable Search Index

Yi Tay, Vinh Q. Tran, Mostafa Dehghani +10

cs.CL2024

Fractal Patterns May Illuminate the Success of Next-Token Prediction

Ibrahim Alabdulmohsin, Vinh Q. Tran, Mostafa Dehghani

cs.CV2023

Dual PatchNorm

Manoj Kumar, Mostafa Dehghani, Neil Houlsby

cs.LG2022

Retrieval-Enhanced Machine Learning

Hamed Zamani, Fernando Diaz, Mostafa Dehghani +2

cs.IR2024

The Impact of Group Membership Bias on the Quality and Fairness of Exposure in Ranking

Ali Vardasbi, Maarten de Rijke, Fernando Diaz +1

cs.LG2020

MetNet: A Neural Weather Model for Precipitation Forecasting

Casper Kaae Sønderby, Lasse Espeholt, Jonathan Heek +6

cs.CL2025

Gemini: A Family of Highly Capable Multimodal Models

Gemini Team, Rohan Anil, Sebastian Borgeaud +1340

q-bio.QM2025

Karyotype AI for Precision Oncology

Zahra Shamsi, Isaac Reid, Drew Bryant +13

cs.CV2022

Beyond Transfer Learning: Co-finetuning for Action Localisation

Anurag Arnab, Xuehan Xiong, Alexey Gritsenko +6

cs.CV2022

Discrete Representations Strengthen Vision Transformer Robustness

Chengzhi Mao, Lu Jiang, Mostafa Dehghani +3

cs.CV2021

An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov +9

cs.IR2018

Learning to Rank from Samples of Variable Quality

Mostafa Dehghani, Jaap Kamps

cs.LG2020

Transferring Inductive Biases through Knowledge Distillation

Samira Abnar, Mostafa Dehghani, Willem Zuidema

cs.CL2022

Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers

Yi Tay, Mostafa Dehghani, Jinfeng Rao +7

cs.CL2022

Are Pre-trained Convolutions Better than Pre-trained Transformers?

Yi Tay, Mostafa Dehghani, Jai Gupta +4

cs.IR2017

Neural Networks for Information Retrieval

Tom Kenter, Alexey Borisov, Christophe Van Gysel +3

cs.CL2023

PaLM 2 Technical Report

Rohan Anil, Andrew M. Dai, Orhan Firat +125

cs.LG2021

Gradual Domain Adaptation in the Wild:When Intermediate Distributions are Absent

Samira Abnar, Rianne van den Berg, Golnaz Ghiasi +3

cs.CV2024

Frozen Feature Augmentation for Few-Shot Image Classification

Andreas Bär, Neil Houlsby, Mostafa Dehghani +1

cs.LG2023

Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints

Aran Komatsuzaki, Joan Puigcerver, James Lee-Thorp +6

cs.CL2018

HiTR: Hierarchical Topic Model Re-estimation for Measuring Topical Diversity of Documents

Hosein Azarbonyad, Mostafa Dehghani, Tom Kenter +3

cs.LG2017

Avoiding Your Teacher's Mistakes: Training Neural Networks with Controlled Weak Supervision

Mostafa Dehghani, Aliaksei Severyn, Sascha Rothe +1

cs.CV2022

Simple Open-Vocabulary Object Detection with Vision Transformers

Matthias Minderer, Alexey Gritsenko, Austin Stone +11

cs.CL2025

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431

cs.LG2022

Efficient Transformers: A Survey

Yi Tay, Mostafa Dehghani, Dara Bahri +1

cs.IR2017

Neural Ranking Models with Weak Supervision

Mostafa Dehghani, Hamed Zamani, Aliaksei Severyn +2

cs.LG2022

Scaling Laws vs Model Architectures: How does Inductive Bias Influence Scaling?

Yi Tay, Mostafa Dehghani, Samira Abnar +7

cs.LG2023

Adaptive Computation with Elastic Input Sequence

Fuzhao Xue, Valerii Likhosherstov, Anurag Arnab +3

cs.CV2023

Scaling Vision Transformers to 22 Billion Parameters

Mostafa Dehghani, Josip Djolonga, Basil Mustafa +39

cs.CV2023

How (not) to ensemble LVLMs for VQA

Lisa Alazraki, Lluis Castrejon, Mostafa Dehghani +3

cs.LG2022

Scaling Instruction-Finetuned Language Models

Hyung Won Chung, Le Hou, Shayne Longpre +32

cs.IR2016

Generalized Group Profiling for Content Customization

Mostafa Dehghani, Hosein Azarbonyad, Jaap Kamps +1

cs.IR2017

On Search Powered Navigation

Mostafa Dehghani, Glorianna Jagfeld, Hosein Azarbonyad +3

cs.IR2018

Neural Networks for Information Retrieval

Tom Kenter, Alexey Borisov, Christophe Van Gysel +3

cs.CL2022

Confident Adaptive Language Modeling

Tal Schuster, Adam Fisch, Jai Gupta +5

cs.CV2021

SCENIC: A JAX Library for Computer Vision Research and Beyond

Mostafa Dehghani, Alexey Gritsenko, Anurag Arnab +2

cs.IR2017

Share your Model instead of your Data: Privacy Preserving Mimic Learning for Ranking

Mostafa Dehghani, Hosein Azarbonyad, Jaap Kamps +1

cs.CL2023

DSI++: Updating Transformer Memory with New Documents

Sanket Vaibhav Mehta, Jai Gupta, Yi Tay +6

cs.LG2022

Intersection of Parallels as an Early Stopping Criterion

Ali Vardasbi, Maarten de Rijke, Mostafa Dehghani

cs.LG2022

The Efficiency Misnomer

Mostafa Dehghani, Anurag Arnab, Lucas Beyer +2

cs.IR2016

On Horizontal and Vertical Separation in Hierarchical Text Classification

Mostafa Dehghani, Hosein Azarbonyad, Jaap Kamps +1

cs.CV2021

ViViT: A Video Vision Transformer

Anurag Arnab, Mostafa Dehghani, Georg Heigold +3

cs.CL2022

Transcending Scaling Laws with 0.1% Extra Compute

Yi Tay, Jason Wei, Hyung Won Chung +13

cs.LG2020

Long Range Arena: A Benchmark for Efficient Transformers

Yi Tay, Mostafa Dehghani, Samira Abnar +7