3 citations · 3 across the 6 of their papers we have counts for
10 papers
GAP3D: Generative Alignment of VLM Latents to Patch-Level Embeddings for 3D Generation
Polytimi Anna Gkotsi, Andrii Zadaianchuk, Mohammad Mahdi Derakhshani
Recent approaches integrating vision-language models (VLMs) as prompt encoders for generative model conditioning typically rely on expensive end-to-end training or map features to…
From Demonstrations to Rewards: Test-Time Prompt Optimization for VLM Reward Models
Christian Gumbsch, Leonardo Barcellona, Lennard Schünemann +7
Reinforcement learning relies on accurate reward functions, which are often hand-crafted or even unavailable in real-world applications, such as robotics. Recent work has explored…
Reconstruction by Generation: 3D Multi-Object Scene Reconstruction from Sparse Observations
Andrii Zadaianchuk, Leonardo Barcellona, Lennard Schuenemann +7
Accurately reconstructing complex full multi-object scenes from sparse observations remains a core challenge in computer vision and a key step toward scalable and reliable simulati…
Evaluating Newtonian Mechanics in Video Generative Models with Real Physical Systems
Antonios Tragoudaras, Chenyu Zhang, Daniil Cherniavskii +7
Recent advances in image and video generation raise hopes that these models possess world modeling capabilities-the ability to generate realistic, physically plausible videos. This…
CTRL-O: Language-Controllable Object-Centric Visual Representation Learning
Aniket Didolkar, Andrii Zadaianchuk, Rabiul Awal +3
Object-centric representation learning aims to decompose visual scenes into fixed-size vectors called "slots" or "object files", where each slot captures a distinct object. Current…
SENSEI: Semantic Exploration Guided by Foundation Models to Learn Versatile World Models
Cansu Sancaktar, Christian Gumbsch, Andrii Zadaianchuk +2
Exploration is a cornerstone of reinforcement learning (RL). Intrinsic motivation attempts to decouple exploration from external, task-based rewards. However, established approache…