activity
20202026
most citedLRTA: A Transparent Neural-Symbolic Reasoning Framework with Modular Supervision for Visual Question Answering

11 citations · 47 across the 22 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

Scene Reconstruction as Mapping Priors for 3D Detection

Yang Fu, Yuliang Zou, Hao Xiang +8

In autonomous driving, mapping is critical for motion planning but remains an under-utilized resource for perception tasks such as 3D object detection. Maps can provide robust stru…

cs.CV2026

STELLAR: Scaling 3D Perception Large Models for Autonomous Driving

Yingwei Li, Xin Huang, Yang Liu +13

Model scaling has demonstrated remarkable success through large-scale training on diverse datasets. It remains an open question whether the same paradigm would apply to autonomous…

cs.CV2022

Privacy Preserving Visual Question Answering

Cristian-Paul Bara, Qing Ping, Abhinav Mathur +3

We introduce a novel privacy-preserving methodology for performing Visual Question Answering on the edge. Our method constructs a symbolic representation of the visual scene, using…

cs.CV20225 cited

A Thousand Words Are Worth More Than a Picture: Natural Language-Centric Outside-Knowledge Visual Question Answering

Feng Gao, Qing Ping, Govind Thattai +3

Outside-knowledge visual question answering (OK-VQA) requires the agent to comprehend the image, make use of relevant knowledge from the entire web, and digest all the information…

cs.CV20215 cited

Embodied BERT: A Transformer Model for Embodied, Language-guided Visual Task Completion

Alessandro Suglia, Qiaozi Gao, Jesse Thomason +2

Language-guided robots performing home and office tasks must navigate in and interact with the world. Grounding language instructions against visual observations and actions to tak…

cs.CV2021

Learning Better Visual Dialog Agents with Pretrained Visual-Linguistic Representation

Tao Tu, Qing Ping, Govind Thattai +2

GuessWhat?! is a two-player visual dialog guessing game where player A asks a sequence of yes/no questions (Questioner) and makes a final guess (Guesser) about a target object in a…