activity
20242026
collaborators

6 papers

cs.CV2026

Balancing Frequencies and Pixels in Flow Matching

Lucas Degeorge, Paul Couairon, Arijit Ghosh +3

Natural images follow a spectral distribution: most signal energy lies in the low spatial frequencies, while the perceptually important structures such as textures and edge…

cs.IR2026

RE-TRIANGLE: Does TRIANGLE Enable Multimodal Alignment Beyond Cosine Similarity in Retrieval?

Arijit Ghosh, Aritra Bandyopadhyay, Chiranjeev Bindra +1

Multimodal alignment is critical for bridging the semantic gap in information retrieval. However, traditional pairwise strategies introduce a geometric blind spot: while they align…

cs.CV2026

PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer

David Picard, Nicolas Dufour, Lucas Degeorge +14

This paper introduces the Polynomial Mixer (PoM), a novel token mixing mechanism with linear complexity that serves as a drop-in replacement for self-attention. PoM aggregates inpu…

cs.CV2025

MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency

Nicolas Dufour, Lucas Degeorge, Arijit Ghosh +2

The default paradigm of post-training text-to-image generators includes post-hoc selection of generated images, and subsequent training with one reward model to align the generator…

cs.CV2025

How far can we go with ImageNet for Text-to-Image generation?

L. Degeorge, A. Ghosh, N. Dufour +2

Recent text-to-image (T2I) generation models have achieved remarkable sucess by training on billion-scale datasets, following a `bigger is better' paradigm that prioritizes data qu…

cs.CV2024

EchoNet-Synthetic: Privacy-preserving Video Generation for Safe Medical Data Sharing

Hadrien Reynaud, Qingjie Meng, Mischa Dombrowski +5

To make medical datasets accessible without sharing sensitive patient information, we introduce a novel end-to-end approach for generative de-identification of dynamic medical imag…