126 citations · 126 across the 5 of their papers we have counts for
4 papers · 1 filter
Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents
Yunseok Jang, Yeda Song, Sungryull Sohn +5
Recent advancements in Large Language Models (LLMs) and Vision-Language Models (VLMs) have sparked significant interest in developing GUI visual agents. We introduce MONDAY (Mobile…
YTCommentQA: Video Question Answerability in Instructional Videos
Saelyne Yang, Sunghyun Park, Yunseok Jang +1
Instructional videos provide detailed how-to guides for various tasks, with viewers often posing questions regarding the content. Addressing these questions is vital for comprehend…
RiCS: A 2D Self-Occlusion Map for Harmonizing Volumetric Objects
Yunseok Jang, Ruben Villegas, Jimei Yang +3
There have been remarkable successes in computer vision with deep learning. While such breakthroughs show robust performance, there have still been many challenges in learning in-d…
Video Prediction with Appearance and Motion Conditions
Yunseok Jang, Gunhee Kim, Yale Song
Video prediction aims to generate realistic future frames by learning dynamic visual patterns. One fundamental challenge is to deal with future uncertainty: How should a model beha…