3 papers
cs.CV2026
Multiplayer Interactive World Models with Representation Autoencoders
Anthony Hu, Václav Volhejn, Adrien Ramanana Rahary +24
We introduce the first multiplayer world model for highly dynamic environments governed by complex physical interactions. Whereas single-player world models treat the other agents…
cs.CV2026
CASA: Cross-Attention over Self-Attention for Efficient Vision-Language Fusion
Moritz Böhle, Amélie Royer, Juliette Marrie +2
Vision-language models (VLMs) are commonly trained by directly inserting image tokens from a pretrained vision encoder into the text stream of a language model. This allows text an…
cs.CV2025
Vision-Speech Models: Teaching Speech Models to Converse about Images
Amélie Royer, Moritz Böhle, Gabriel de Marmiesse +4
The recent successes of Vision-Language models raise the question of how to equivalently imbue a pretrained speech model with vision understanding, an important milestone towards b…