activity
20192022
most citedImproving Visual Question Answering by Referring to Generated Paragraph Captions

6 citations · 17 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CL2022

On the Limits of Evaluating Embodied Agent Model Generalization Using Validation Sets

Hyounghun Kim, Aishwarya Padmakumar, Di Jin +2

Natural language guided embodied task completion is a challenging problem since it requires understanding natural language instructions, aligning them with egocentric visual observ…

cs.CL2022

CAISE: Conversational Agent for Image Search and Editing

Hyounghun Kim, Doo Soon Kim, Seunghyun Yoon +3

Demand for image editing has been increasing as users' desire for expression is also increasing. However, for most users, image editing tools are not easy to use since the tools re…

cs.CL20211 cited

FixMyPose: Pose Correctional Captioning and Retrieval

Hyounghun Kim, Abhay Zala, Graham Burri +1

Interest in physical therapy and individual exercises such as yoga/dance has increased alongside the well-being trend. However, such exercises are hard to follow without expert gui…

cs.CL20203 cited

ArraMon: A Joint Navigation-Assembly Instruction Interpretation Task in Dynamic Environments

Hyounghun Kim, Abhay Zala, Graham Burri +2

For embodied agents, navigation is an important ability but not an isolated goal. Agents are also expected to perform specific tasks after reaching the target location, such as pic…

cs.CL20202 cited

Dense-Caption Matching and Frame-Selection Gating for Temporal Localization in VideoQA

Hyounghun Kim, Zineng Tang, Mohit Bansal

Videos convey rich information. Dynamic spatio-temporal relationships between people/objects, and diverse multimodal events are present in a video clip. Hence, it is important to d…

cs.CL20205 cited

Modality-Balanced Models for Visual Dialogue

Hyounghun Kim, Hao Tan, Mohit Bansal

The Visual Dialog task requires a model to exploit both image and conversational context information to generate the next response to the dialogue. However, via manual analysis, we…