2 papers
cs.CL2026
Decoupled DiLoCo for Resilient Distributed Pre-training
Arthur Douillard, Keith Rush, Yani Donchev +14
Modern large-scale language model pre-training relies heavily on the single program multiple data (SPMD) paradigm, which requires tight coupling across accelerators. Due to this co…
cs.CV2024
Transformer-based Clipped Contrastive Quantization Learning for Unsupervised Image Retrieval
Ayush Dubey, Shiv Ram Dubey, Satish Kumar Singh +1
Unsupervised image retrieval aims to learn the important visual characteristics without any given level to retrieve the similar images for a given query image. The Convolutional Ne…