2 papers
cs.CV2026
SlotVTG: Object-Centric Adapter for Generalizable Video Temporal Grounding
Jiwook Han, Geo Ahn, Youngrae Kim +1
Multimodal Large Language Models (MLLMs) have shown strong performance on Video Temporal Grounding (VTG). However, their coarse recognition capabilities are insufficient for fine-g…
cs.RO2025
Benchmarking Multi-Object Grasping
Tianze Chen, Ricardo Frumento, Giulia Pagnanelli +9
In this work, we describe a multi-object grasping benchmark to evaluate the grasping and manipulation capabilities of robotic systems in both pile and surface scenarios. The benchm…