activity
20242026
collaborators

6 papers

cs.LG2026

TextME: Bridging Unseen Modalities Through Text Descriptions

Soyeon Hong, Jinchan Kim, Jaegook You +3

Expanding multimodal representations to novel modalities is constrained by reliance on large-scale paired datasets (e.g., text-image, text-audio, text-3D, text-molecule), which are…

cs.CL2025

Trillion 7B Technical Report

Sungjun Han, Juyoung Suk, Suyeong An +5

We introduce Trillion-7B, the most token-efficient Korean-centric multilingual LLM available. Our novel Cross-lingual Document Attention (XLDA) mechanism enables highly efficient a…

cs.HC2025

CHOIR: Chat-based Helper for Organizational Intelligence Repository

Sangwook Lee, Adnan Abbas, Yan Chen +1

Modern organizations frequently rely on chat-based platforms (e.g., Slack, Microsoft Teams, and Discord) for day-to-day communication and decision-making. As conversations evolve,…

cs.CL2024

FLEX: Expert-level False-Less EXecution Metric for Reliable Text-to-SQL Benchmark

Heegyu Kim, Taeyang Jeon, Seunghwan Choi +2

Text-to-SQL systems have become crucial for translating natural language into SQL queries in various industries, enabling non-technical users to perform complex data operations. Th…

cs.CL2024

Interventional Speech Noise Injection for ASR Generalizable Spoken Language Understanding

Yeonjoon Jung, Jaeseong Lee, Seungtaek Choi +3

Recently, pre-trained language models (PLMs) have been increasingly adopted in spoken language understanding (SLU). However, automatic speech recognition (ASR) systems frequently p…

cs.CV2024

ScoreCL: Augmentation-Adaptive Contrastive Learning via Score-Matching Function

Jin-Young Kim, Soonwoo Kwon, Hyojun Go +3

Self-supervised contrastive learning (CL) has achieved state-of-the-art performance in representation learning by minimizing the distance between positive pairs while maximizing th…