3 papers
cs.AI2026
Voxtral Realtime
Mistral-AI, :, Alexander H. Liu +166
We introduce Voxtral Realtime, a natively streaming automatic speech recognition model that matches offline transcription quality at sub-second latency. Unlike approaches that adap…
cs.AI2025
Embodied AI Agents: Modeling the World
Pascale Fung, Yoram Bachrach, Asli Celikyilmaz +18
This paper describes our research on AI agents embodied in visual, virtual or physical forms, enabling them to interact with both users and their environments. These agents, which…
cs.CV2025
Locate 3D: Real-World Object Localization via Self-Supervised Learning in 3D
Sergio Arnaud, Paul McVay, Ada Martin +19
We present LOCATE 3D, a model for localizing objects in 3D scenes from referring expressions like "the small coffee table between the sofa and the lamp." LOCATE 3D sets a new state…