13 papers
Out of Sight, Still in Mind: Token Compression for Omni-LLMs
Suho Yoo, Youngjoon Jang, Hyebin Cho +1
The goal of this paper is to reduce the input token cost of Omni-modal large language models (Omni-LLMs) at inference time. Omni-LLMs reason jointly over audio, video and text, but…
Segment, Embed, and Align: A Universal Recipe for Aligning Subtitles to Signing
Zifan Jiang, Youngjoon Jang, Liliane Momeni +3
The goal of this work is to develop a universal approach for aligning subtitles (i.e., spoken language text with corresponding timestamps) to continuous sign language videos. Prior…
On the Nature of Attention Sink that Shapes Decoding Strategy in Omni-LLMs
Suho Yoo, Youngjoon Jang, Joon Son Chung
The goal of this paper is to strengthen the reasoning of Omnimodal Large Language Models (Omni-LLMs) at inference time, without additional training. These models jointly process vi…
EDNet: A Versatile Speech Enhancement Framework with Gating Mamba Mechanism and Phase Shift-Invariant Training
Doyeop Kwak, Youngjoon Jang, Seongyu Kim +1
Speech signals in real-world environments are frequently affected by various distortions such as additive noise, reverberation, and bandwidth limitation, which may appear individua…
FastAV: Efficient Token Pruning for Audio-Visual Large Language Model Inference
Chaeyoung Jung, Youngjoon Jang, Seungwoo Lee +1
In this work, we present FastAV, the first token pruning framework tailored for audio-visual large language models (AV-LLMs). While token pruning has been actively explored in stan…
LP-CFM: Perceptual Invariance-Aware Conditional Flow Matching for Speech Modeling
Doyeop Kwak, Youngjoon Jang, Joon Son Chung
The goal of this paper is to provide a new perspective on speech modeling by incorporating perceptual invariances such as amplitude scaling and temporal shifts. Conventional genera…