4 papers
CoReTab: Improving Multimodal Table Understanding with Code-driven Reasoning
Van-Quang Nguyen, Takayuki Okatani
Existing datasets for multimodal table understanding, such as MMTab, primarily provide short factual answers without explicit multi-step reasoning supervision. Models trained on th…
TB-Bench: Training and Testing Multi-Modal AI for Understanding Spatio-Temporal Traffic Behaviors from Dashcam Images/Videos
Korawat Charoenpitaks, Van-Quang Nguyen, Masanori Suganuma +4
The application of Multi-modal Large Language Models (MLLMs) in Autonomous Driving (AD) faces significant challenges due to their limited training on traffic-specific data and the…
Look Wide and Interpret Twice: Improving Performance on Interactive Instruction-following Tasks
Van-Quang Nguyen, Masanori Suganuma, Takayuki Okatani
There is a growing interest in the community in making an embodied AI agent perform a complicated task while interacting with an environment following natural language directives.…
Efficient Attention Mechanism for Visual Dialog that can Handle All the Interactions between Multiple Inputs
Van-Quang Nguyen, Masanori Suganuma, Takayuki Okatani
It has been a primary concern in recent studies of vision and language tasks to design an effective attention mechanism dealing with interactions between the two modalities. The Tr…