5 papers
Towards Flexible Evaluation for Generative Visual Question Answering
Huishan Ji, Qingyi Si, Zheng Lin +1
Throughout rapid development of multimodal large language models, a crucial ingredient is a fair and accurate evaluation of their multimodal comprehension abilities. Although Visua…
Multimodal Table Understanding
Mingyu Zheng, Xinwei Feng, Qingyi Si +4
Although great progress has been made by previous table understanding methods including recent approaches based on large language models (LLMs), they rely heavily on the premise th…
Think out Loud: Emotion Deducing Explanation in Dialogues
Jiangnan Li, Zheng Lin, Lanrui Wang +6
Humans convey emotions through daily dialogues, making emotion understanding a crucial step of affective intelligence. To understand emotions in dialogues, machines are asked to re…
Towards Unified Interactive Visual Grounding in The Wild
Jie Xu, Hanbo Zhang, Qingyi Si +3
Interactive visual grounding in Human-Robot Interaction (HRI) is challenging yet practical due to the inevitable ambiguity in natural languages. It requires robots to disambiguate…
Combo of Thinking and Observing for Outside-Knowledge VQA
Qingyi Si, Yuchen Mo, Zheng Lin +2
Outside-knowledge visual question answering is a challenging task that requires both the acquisition and the use of open-ended real-world knowledge. Some existing solutions draw ex…