most citedMultimodal Subtask Graph Generation from Instructional Videos

4 citations · 4 across the 6 of their papers we have counts for

collaborators

6 papers

cs.AI2023

Code Models are Zero-shot Precondition Reasoners

Lajanugen Logeswaran, Sungryull Sohn, Yiwei Lyu +5

One of the fundamental skills required for an agent acting in an environment to complete tasks is the ability to understand what actions are plausible at any given point. This work…

cs.LG2023

MultiPrompter: Cooperative Prompt Optimization with Multi-Agent Reinforcement Learning

Dong-Ki Kim, Sungryull Sohn, Lajanugen Logeswaran +2

Recently, there has been an increasing interest in automated prompt optimization based on reinforcement learning (RL). This approach offers important advantages, such as generating…

cs.CL2023

From Heuristic to Analytic: Cognitively Motivated Strategies for Coherent Physical Commonsense Reasoning

Zheyuan Zhang, Shane Storks, Fengyuan Hu +4

Pre-trained language models (PLMs) have shown impressive performance in various language tasks. However, they are prone to spurious correlations, and often generate illusory inform…

cs.CL2023

A Picture is Worth a Thousand Words: Language Models Plan from Pixels

Anthony Z. Liu, Lajanugen Logeswaran, Sungryull Sohn +1

Planning is an important capability of artificial agents that perform long-horizon tasks in real-world environments. In this work, we explore the use of pre-trained language models…

cs.LG20234 cited

Multimodal Subtask Graph Generation from Instructional Videos

Yunseok Jang, Sungryull Sohn, Lajanugen Logeswaran +3

Real-world tasks consist of multiple inter-dependent subtasks (e.g., a dirty pan needs to be washed before it can be used for cooking). In this work, we aim to model the causal dep…

cs.AI2023

Unsupervised Task Graph Generation from Instructional Video Transcripts

Lajanugen Logeswaran, Sungryull Sohn, Yunseok Jang +2

This work explores the problem of generating task graphs of real-world activities. Different from prior formulations, we consider a setting where text transcripts of instructional…