activity
20242026
most citedSolution for SMART-101 Challenge of CVPR Multi-modal Algorithmic Reasoning Task 2024

1 citations · 1 across the 5 of their papers we have counts for

collaborators

6 papers

cs.CV2026

FiRe: Fine-grained Multimodal Reasoning for Enhanced Image Generation

Yongjin Kim, Yoonjin Oh, Yerin Kim +5

With the rapid progress of Multimodal Large Language Models (MLLMs), unified MLLMs that jointly perform image understanding and generation have advanced significantly. However, des…

cs.CV2025

Instruction-tuned Self-Questioning Framework for Multimodal Reasoning

You-Won Jang, Yu-Jung Heo, Jaeseok Kim +3

The field of vision-language understanding has been actively researched in recent years, thanks to the development of Large Language Models~(LLMs). However, it still needs help wit…

cs.AI2024

BI-MDRG: Bridging Image History in Multimodal Dialogue Response Generation

Hee Suk Yoon, Eunseop Yoon, Joshua Tian Jin Tee +4

Multimodal Dialogue Response Generation (MDRG) is a recently proposed task where the model needs to generate responses in texts, images, or a blend of both based on the dialogue co…

cs.CV20241 cited

Solution for SMART-101 Challenge of CVPR Multi-modal Algorithmic Reasoning Task 2024

Jinwoo Ahn, Junhyeok Park, Min-Jun Kim +6

In this paper, the solution of HYU MLLAB KT Team to the Multimodal Algorithmic Reasoning Task: SMART-101 CVPR 2024 Challenge is presented. Beyond conventional visual question-answe…

cs.CL2024

Translation Deserves Better: Analyzing Translation Artifacts in Cross-lingual Visual Question Answering

ChaeHun Park, Koanho Lee, Hyesu Lim +5

Building a reliable visual question answering~(VQA) system across different languages is a challenging problem, primarily due to the lack of abundant samples for training. To addre…

cs.CL2024

Evaluating Visual and Cultural Interpretation: The K-Viscuit Benchmark with Human-VLM Collaboration

ChaeHun Park, Yujin Baek, Jaeseok Kim +3

To create culturally inclusive vision-language models (VLMs), developing a benchmark that tests their ability to address culturally relevant questions is essential. Existing approa…