2 papers
cs.CV2026
VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis
Jinho Park, Youbin Kim, Hogun Park +1
Spatio-temporal reasoning is a core capability for Multimodal Large Language Models (MLLMs) operating in the real world. As such, evaluating it precisely has become an essential ch…
cs.CV2026
Group3D: MLLM-Driven Semantic Grouping for Open-Vocabulary 3D Object Detection
Youbin Kim, Jinho Park, Hogun Park +1
Open-vocabulary 3D object detection aims to localize and recognize objects beyond a fixed training taxonomy. In multi-view RGB settings, recent approaches often decouple geometry-b…