3 papers
cs.CV2026
Multiplayer Interactive World Models with Representation Autoencoders
Anthony Hu, Václav Volhejn, Adrien Ramanana Rahary +24
We introduce the first multiplayer world model for highly dynamic environments governed by complex physical interactions. Whereas single-player world models treat the other agents…
cs.CV2026
CASA: Cross-Attention over Self-Attention for Efficient Vision-Language Fusion
Moritz Böhle, Amélie Royer, Juliette Marrie +2
Vision-language models (VLMs) are commonly trained by directly inserting image tokens from a pretrained vision encoder into the text stream of a language model. This allows text an…
cs.CV2025
B-cosification: Transforming Deep Neural Networks to be Inherently Interpretable
Shreyash Arya, Sukrut Rao, Moritz Böhle +1
B-cos Networks have been shown to be effective for obtaining highly human interpretable explanations of model decisions by architecturally enforcing stronger alignment between inpu…