activity
20242026
collaborators

11 papers

cs.CL2026

NaturalFlow: Reducing Disruptive Pauses for Natural Speech Flow in Simultaneous Speech-to-Speech Translation

Dongwook Lee, Youngho Cho, Sangkwon Park +2

Simultaneous speech-to-speech translation aims to enable near-real-time communication by minimizing latency, offering a compelling, real-time alternative to the high latency of con…

cs.LG2026

Rollout-Level Advantage-Prioritized Experience Replay for GRPO

Gyeongtae Yoo, Sanghyeok Park, Soohyuk Jang +2

Reinforcement learning from verifiable rewards with GRPO is a standard approach for post-training reasoning LLMs. It remains sample inefficient. Each rollout is used for a single g…

cs.CV2026

HeSS: Head Sensitivity Score for Sparsity Redistribution in VGGT

Yongsung Kim, Wooseok Song, Jaihyun Lew +3

Visual Geometry Grounded Transformer (VGGT) has advanced 3D vision, yet its global attention layers suffer from quadratic computational costs that hinder scalability. Several spars…

cs.CV2026

Balancing Saliency and Coverage: Semantic Prominence-Aware Budgeting for Visual Token Compression in VLMs

Jaehoon Lee, Mingi Jung, Soohyuk Jang +3

Large Vision-Language Models (VLMs) achieve strong multimodal understanding capabilities by leveraging high-resolution visual inputs, but the resulting large number of visual token…

cs.CV2025

SAVE: Sparse Autoencoder-Driven Visual Information Enhancement for Mitigating Object Hallucination

Sangha Park, Seungryong Yoo, Jisoo Mok +1

Although Multimodal Large Language Models (MLLMs) have advanced substantially, they remain vulnerable to object hallucination caused by language priors and visual information loss.…

cs.CV2025

Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment

Sangha Park, Eunji Kim, Yeongtak Oh +2

Despite substantial progress in text-to-image generation, achieving precise text-image alignment remains challenging, particularly for prompts with rich compositional structure or…