1 paper
Ivy Xiao He, Stefanie Tellex, Jason Xinyu Liu
To assist humans in open-world environments, robots must interpret ambiguous instructions to locate desired objects. Foundation model-based approaches excel at multimodal grounding…