collaborators

5 papers

cs.CL2025

LM-SPT: LM-Aligned Semantic Distillation for Speech Tokenization

Daejin Jo, Jeeyoung Yun, Byungseok Roh +1

With the rapid progress of speech language models (SLMs), discrete speech tokens have emerged as a core interface between speech and text, enabling unified modeling across modaliti…

cs.CL2024

CheX-GPT: Harnessing Large Language Models for Enhanced Chest X-ray Report Labeling

Jawook Gu, Kihyun You, Han-Cheol Cho +3

Free-text radiology reports present a rich data source for various medical tasks, but effectively labeling these texts remains challenging. Traditional rule-based labeling methods…

cs.CV2023

Honeybee: Locality-enhanced Projector for Multimodal LLM

Junbum Cha, Wooyoung Kang, Jonghwan Mun +1

In Multimodal Large Language Models (MLLMs), a visual projector plays a crucial role in bridging pre-trained vision encoders with LLMs, enabling profound visual understanding while…

cs.CV20232 cited

Large Language Models are Temporal and Causal Reasoners for Video Question Answering

Dohwan Ko, Ji Soo Lee, Wooyoung Kang +2

Large Language Models (LLMs) have shown remarkable performances on a wide range of natural language understanding and generation tasks. We observe that the LLMs provide effective p…

cs.CV202381 cited

CXR-CLIP: Toward Large Scale Chest X-ray Language-Image Pre-training

Kihyun You, Jawook Gu, Jiyeon Ham +5

A large-scale image-text pair dataset has greatly contributed to the development of vision-language pre-training (VLP) models, which enable zero-shot or few-shot classification wit…