7 papers
GAP3D: Generative Alignment of VLM Latents to Patch-Level Embeddings for 3D Generation
Polytimi Anna Gkotsi, Andrii Zadaianchuk, Mohammad Mahdi Derakhshani
Recent approaches integrating vision-language models (VLMs) as prompt encoders for generative model conditioning typically rely on expensive end-to-end training or map features to…
Purrception: Variational Flow Matching for Vector-Quantized Image Generation
RÄzvan-Andrei MatiÅan, Vincent Tao Hu, Grigory Bartosh +6
We introduce Purrception, a variational flow matching approach for vector-quantized image generation that provides explicit categorical supervision while maintaining continuous tra…
Private PoEtry: Private In-Context Learning via Product of Experts
Rob Romijnders, Mohammad Mahdi Derakhshani, Jonathan Petit +3
In-context learning (ICL) enables Large Language Models (LLMs) to adapt to new tasks with only a small set of examples at inference time, thereby avoiding task-specific fine-tuning…
NeoBabel: A Multilingual Open Tower for Visual Generation
Mohammad Mahdi Derakhshani, Dheeraj Varghese, Marzieh Fadaee +1
Text-to-image generation advancements have been predominantly English-centric, creating barriers for non-English speakers and perpetuating digital inequities. While existing system…
Continual Hyperbolic Learning of Instances and Classes
Melika Ayoughi, Mina Ghadimi Atigh, Mohammad Mahdi Derakhshani +3
Continual learning has traditionally focused on classifying either instances or classes, but real-world applications, such as robotics and self-driving cars, require models to hand…
TULIP: Token-length Upgraded CLIP
Ivona Najdenkoska, Mohammad Mahdi Derakhshani, Yuki M. Asano +3
We address the challenge of representing long captions in vision-language models, such as CLIP. By design these models are limited by fixed, absolute positional encodings, restrict…