activity
20242026
collaborators

13 papers

cs.CV2026

Out of Sight, Still in Mind: Token Compression for Omni-LLMs

Suho Yoo, Youngjoon Jang, Hyebin Cho +1

The goal of this paper is to reduce the input token cost of Omni-modal large language models (Omni-LLMs) at inference time. Omni-LLMs reason jointly over audio, video and text, but…

cs.CL2026

Segment, Embed, and Align: A Universal Recipe for Aligning Subtitles to Signing

Zifan Jiang, Youngjoon Jang, Liliane Momeni +3

The goal of this work is to develop a universal approach for aligning subtitles (i.e., spoken language text with corresponding timestamps) to continuous sign language videos. Prior…

cs.CV2026

On the Nature of Attention Sink that Shapes Decoding Strategy in Omni-LLMs

Suho Yoo, Youngjoon Jang, Joon Son Chung

The goal of this paper is to strengthen the reasoning of Omnimodal Large Language Models (Omni-LLMs) at inference time, without additional training. These models jointly process vi…

eess.AS2026

EDNet: A Versatile Speech Enhancement Framework with Gating Mamba Mechanism and Phase Shift-Invariant Training

Doyeop Kwak, Youngjoon Jang, Seongyu Kim +1

Speech signals in real-world environments are frequently affected by various distortions such as additive noise, reverberation, and bandwidth limitation, which may appear individua…

cs.LG2026

FastAV: Efficient Token Pruning for Audio-Visual Large Language Model Inference

Chaeyoung Jung, Youngjoon Jang, Seungwoo Lee +1

In this work, we present FastAV, the first token pruning framework tailored for audio-visual large language models (AV-LLMs). While token pruning has been actively explored in stan…

eess.AS2025

LP-CFM: Perceptual Invariance-Aware Conditional Flow Matching for Speech Modeling

Doyeop Kwak, Youngjoon Jang, Joon Son Chung

The goal of this paper is to provide a new perspective on speech modeling by incorporating perceptual invariances such as amplitude scaling and temporal shifts. Conventional genera…