1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.CL2025
SpidR: Learning Fast and Stable Linguistic Units for Spoken Language Models Without Supervision
Maxime Poli, Mahi Luthra, Youssef Benchekroun +8
The parallel advances in language modeling and speech representation learning have raised the prospect of learning language directly from speech without textual intermediates. This…
cs.CV2025
Locate 3D: Real-World Object Localization via Self-Supervised Learning in 3D
Sergio Arnaud, Paul McVay, Ada Martin +19
We present LOCATE 3D, a model for localizing objects in 3D scenes from referring expressions like "the small coffee table between the sofa and the lamp." LOCATE 3D sets a new state…
cs.CV2024★ 1 cited
VEDIT: Latent Prediction Architecture For Procedural Video Representation Learning
Han Lin, Tushar Nagarajan, Nicolas Ballas +4
Procedural video representation learning is an active research area where the objective is to learn an agent which can anticipate and forecast the future given the present video in…