1 paper
Yi Zhang, Yi Wang, Yueting Wu +3
Image-based 3D visual grounding is critical for embodied agents, yet existing benchmarks suffer from loose text-observation alignment and neglect temporal ordering. We introduce Se…