activity
20162023
most citedPaLM-E: An Embodied Multimodal Language Model

356 citations · 366 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV2023

Sensitivity of Slot-Based Object-Centric Models to their Number of Slots

Roland S. Zimmermann, Sjoerd van Steenkiste, Mehdi S. M. Sajjadi +2

Self-supervised methods for learning object-centric representations have recently been applied successfully to various datasets. This progress is largely fueled by slot-based metho…

cs.CV2023

RePAST: Relative Pose Attention Scene Representation Transformer

Aleksandr Safin, Daniel Duckworth, Mehdi S. M. Sajjadi

The Scene Representation Transformer (SRT) is a recent method to render novel views at interactive rates. Since SRT uses camera poses with respect to an arbitrarily chosen referenc…

cs.LG2023356 cited

PaLM-E: An Embodied Multimodal Language Model

Danny Driess, Fei Xia, Mehdi S. M. Sajjadi +19

Large language models excel at a wide range of complex tasks. However, enabling general inference in the real world, e.g., for robotics problems, raises the challenge of grounding.…

cs.CV20214 cited

NeSF: Neural Semantic Fields for Generalizable Semantic Segmentation of 3D Scenes

Suhani Vora, Noha Radwan, Klaus Greff +6

We present NeSF, a method for producing 3D semantic fields from posed RGB images alone. In place of classical 3D representations, our method builds on recent work in implicit neura…

cs.CV20166 cited

Depth Estimation Through a Generative Model of Light Field Synthesis

Mehdi S. M. Sajjadi, Rolf Köhler, Bernhard Schölkopf +1

Light field photography captures rich structural information that may facilitate a number of traditional image processing and computer vision tasks. A crucial ingredient in such en…