4 citations · 4 across the 4 of their papers we have counts for
4 papers
Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents
Yunseok Jang, Yeda Song, Sungryull Sohn +5
Recent advancements in Large Language Models (LLMs) and Vision-Language Models (VLMs) have sparked significant interest in developing GUI visual agents. We introduce MONDAY (Mobile…
YTCommentQA: Video Question Answerability in Instructional Videos
Saelyne Yang, Sunghyun Park, Yunseok Jang +1
Instructional videos provide detailed how-to guides for various tasks, with viewers often posing questions regarding the content. Addressing these questions is vital for comprehend…
Multimodal Subtask Graph Generation from Instructional Videos
Yunseok Jang, Sungryull Sohn, Lajanugen Logeswaran +3
Real-world tasks consist of multiple inter-dependent subtasks (e.g., a dirty pan needs to be washed before it can be used for cooking). In this work, we aim to model the causal dep…
Unsupervised Task Graph Generation from Instructional Video Transcripts
Lajanugen Logeswaran, Sungryull Sohn, Yunseok Jang +2
This work explores the problem of generating task graphs of real-world activities. Different from prior formulations, we consider a setting where text transcripts of instructional…