6 citations · 17 across the 7 of their papers we have counts for
7 papers
On the Limits of Evaluating Embodied Agent Model Generalization Using Validation Sets
Hyounghun Kim, Aishwarya Padmakumar, Di Jin +2
Natural language guided embodied task completion is a challenging problem since it requires understanding natural language instructions, aligning them with egocentric visual observ…
CAISE: Conversational Agent for Image Search and Editing
Hyounghun Kim, Doo Soon Kim, Seunghyun Yoon +3
Demand for image editing has been increasing as users' desire for expression is also increasing. However, for most users, image editing tools are not easy to use since the tools re…
FixMyPose: Pose Correctional Captioning and Retrieval
Hyounghun Kim, Abhay Zala, Graham Burri +1
Interest in physical therapy and individual exercises such as yoga/dance has increased alongside the well-being trend. However, such exercises are hard to follow without expert gui…
ArraMon: A Joint Navigation-Assembly Instruction Interpretation Task in Dynamic Environments
Hyounghun Kim, Abhay Zala, Graham Burri +2
For embodied agents, navigation is an important ability but not an isolated goal. Agents are also expected to perform specific tasks after reaching the target location, such as pic…
Dense-Caption Matching and Frame-Selection Gating for Temporal Localization in VideoQA
Hyounghun Kim, Zineng Tang, Mohit Bansal
Videos convey rich information. Dynamic spatio-temporal relationships between people/objects, and diverse multimodal events are present in a video clip. Hence, it is important to d…
Modality-Balanced Models for Visual Dialogue
Hyounghun Kim, Hao Tan, Mohit Bansal
The Visual Dialog task requires a model to exploit both image and conversational context information to generate the next response to the dialogue. However, via manual analysis, we…