3 papers
cs.RO2025
REALM: A Real-to-Sim Validated Benchmark for Generalization in Robotic Manipulation
Martin Sedlacek, Pavlo Yefanov, Georgy Ponimatkin +7
Vision-Language-Action (VLA) models empower robots to understand and execute tasks described by natural language instructions. However, a key challenge lies in their ability to gen…
cs.CV2025
Large-scale Pre-training for Grounded Video Caption Generation
Evangelos Kazakos, Cordelia Schmid, Josef Sivic
We propose a novel approach for captioning and object grounding in video, where the objects in the caption are grounded in the video via temporally dense bounding boxes. We introdu…
cs.SD2025
Epic-Sounds: A Large-scale Dataset of Actions That Sound
Jaesung Huh, Jacob Chalk, Evangelos Kazakos +2
We introduce EPIC-SOUNDS, a large-scale dataset of audio annotations capturing temporal extents and class labels within the audio stream of the egocentric videos. We propose an ann…