8 papers
DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone
Vaibhav Singh, Oleksiy Ostapenko, Pierre-André Noël +2
Diffusion language models (DLMs) have emerged as a promising alternative to autoregressive (AR) generation, yet their reliance on Transformer backbones limits inference efficiency…
Dual-Phase Continual Learning: Supervised Adaptation Meets Unsupervised Retention
Vaibhav Singh, Rahaf Aljundi, Eugene Belilovsky
Foundational Vision-Language Models (VLMs) excel across diverse tasks, but adapting them to new domains without forgetting prior knowledge remains a critical challenge. Continual L…
Heterogeneous Low-Bandwidth Pre-Training of LLMs
Yazan Obeidi, Amir Sarfi, Joel Lidin +2
Pre-training large language models (LLMs) increasingly requires distributed compute, yet bandwidth constraints make it difficult to scale beyond well-provisioned datacenters-especi…
End-to-End Fine-Tuning of 3D Texture Generation using Differentiable Rewards
AmirHossein Zamani, Tianhao Xie, Amir G. Aghdam +2
While recent 3D generative models can produce high-quality texture images, they often fail to capture human preferences or meet task-specific requirements. Moreover, a core challen…
Sketch-guided Cage-based 3D Gaussian Splatting Deformation
Tianhao Xie, Noam Aigerman, Eugene Belilovsky +1
3D Gaussian Splatting (GS) is one of the most promising novel 3D representations that has received great interest in computer graphics and computer vision. While various systems ha…
When Data Falls Short: Grokking Below the Critical Threshold
Vaibhav Singh, Eugene Belilovsky, Rahaf Aljundi
In this paper, we investigate the phenomenon of grokking, where models exhibit delayed generalization following overfitting on training data. We focus on data-scarce regimes where…