activity
20242026
collaborators

7 papers

cs.CL2026

ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly

Kimihiro Hasegawa, Wiradee Imrattanatrai, Masaki Asada +5

Assistants on assembly tasks show great potential to benefit humans ranging from helping with everyday tasks to interacting in industrial settings. However, evaluation resources in…

cs.CV2026

Towards Artwork Explanation in Large-scale Vision Language Models

Kazuki Hayashi, Yusuke Sakai, Hidetaka Kamigaito +2

Large-scale Vision-Language Models (LVLMs) output text from images and instructions, demonstrating capabilities in text generation and comprehension. However, it has not been clari…

cs.CL2025

ProMQA: Question Answering Dataset for Multimodal Procedural Activity Understanding

Kimihiro Hasegawa, Wiradee Imrattanatrai, Zhi-Qi Cheng +5

Multimodal systems have great potential to assist humans in procedural activities, where people follow instructions to achieve their goals. Despite diverse application scenarios, s…

cs.CL2025

TAMA: Tool-Augmented Multimodal Agent for Procedural Activity Understanding

Kimihiro Hasegawa, Wiradee Imrattanatrai, Masaki Asada +2

Procedural activity assistants potentially support humans in a variety of settings, from our daily lives, e.g., cooking or assembling flat-pack furniture, to professional situation…

cs.CV2025

A Video-grounded Dialogue Dataset and Metric for Event-driven Activities

Wiradee Imrattanatrai, Masaki Asada, Kimihiro Hasegawa +3

This paper presents VDAct, a dataset for a Video-grounded Dialogue on Event-driven Activities, alongside VDEval, a session-based context evaluation metric specially designed for th…

cs.AI2024

Human-Aware Vision-and-Language Navigation: Bridging Simulation to Reality with Dynamic Human Interactions

Heng Li, Minghan Li, Zhi-Qi Cheng +6

Vision-and-Language Navigation (VLN) aims to develop embodied agents that navigate based on human instructions. However, current VLN frameworks often rely on static environments an…