Showing cs.CVShow all
2 papers · 1 filter
cs.CV2024★ 1 cited
Multi-Agent VQA: Exploring Multi-Agent Foundation Models in Zero-Shot Visual Question Answering
Bowen Jiang, Zhijun Zhuang, Shreyas S. Shivakumar +2
This work explores the zero-shot capabilities of foundation models in Visual Question Answering (VQA) tasks. We propose an adaptive multi-agent system, named Multi-Agent VQA, to ov…
cs.CV2023
Instance-Agnostic Geometry and Contact Dynamics Learning
Mengti Sun, Bowen Jiang, Bibit Bianchini +2
This work presents an instance-agnostic learning framework that fuses vision with dynamics to simultaneously learn shape, pose trajectories, and physical properties via the use of…