7 citations · 10 across the 8 of their papers we have counts for
Showing 2025Show all
2 papers · 1 filter
cs.RO2025
REALM: A Real-to-Sim Validated Benchmark for Generalization in Robotic Manipulation
Martin Sedlacek, Pavlo Yefanov, Georgy Ponimatkin +7
Vision-Language-Action (VLA) models empower robots to understand and execute tasks described by natural language instructions. However, a key challenge lies in their ability to gen…
cs.CV2025
Large-scale Pre-training for Grounded Video Caption Generation
Evangelos Kazakos, Cordelia Schmid, Josef Sivic
We propose a novel approach for captioning and object grounding in video, where the objects in the caption are grounded in the video via temporally dense bounding boxes. We introdu…