11 citations · 47 across the 22 of their papers we have counts for
6 papers · 1 filter
Scene Reconstruction as Mapping Priors for 3D Detection
Yang Fu, Yuliang Zou, Hao Xiang +8
In autonomous driving, mapping is critical for motion planning but remains an under-utilized resource for perception tasks such as 3D object detection. Maps can provide robust stru…
STELLAR: Scaling 3D Perception Large Models for Autonomous Driving
Yingwei Li, Xin Huang, Yang Liu +13
Model scaling has demonstrated remarkable success through large-scale training on diverse datasets. It remains an open question whether the same paradigm would apply to autonomous…
Privacy Preserving Visual Question Answering
Cristian-Paul Bara, Qing Ping, Abhinav Mathur +3
We introduce a novel privacy-preserving methodology for performing Visual Question Answering on the edge. Our method constructs a symbolic representation of the visual scene, using…
A Thousand Words Are Worth More Than a Picture: Natural Language-Centric Outside-Knowledge Visual Question Answering
Feng Gao, Qing Ping, Govind Thattai +3
Outside-knowledge visual question answering (OK-VQA) requires the agent to comprehend the image, make use of relevant knowledge from the entire web, and digest all the information…
Embodied BERT: A Transformer Model for Embodied, Language-guided Visual Task Completion
Alessandro Suglia, Qiaozi Gao, Jesse Thomason +2
Language-guided robots performing home and office tasks must navigate in and interact with the world. Grounding language instructions against visual observations and actions to tak…
Learning Better Visual Dialog Agents with Pretrained Visual-Linguistic Representation
Tao Tu, Qing Ping, Govind Thattai +2
GuessWhat?! is a two-player visual dialog guessing game where player A asks a sequence of yes/no questions (Questioner) and makes a final guess (Guesser) about a target object in a…