◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Minsu Kim

8 papers hereh-index 494 citations8 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author2
  • middle author6

Across the 8 of 8 papers where every author was matched, so the position is known.

fields
  • cs.CV4
  • eess.AS3
  • cs.SD1
same name
  • Minsu Kim — 11 papers, h 7
  • Minsu Kim — 10 papers, h 4
  • Minsu Kim — 10 papers, h 6
  • Minsu Kim — 9 papers, h 7
  • Minsu Kim — 8 papers, h 18
  • Minsu Kim — 4 papers, h 3

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

collaborators
Showing cs.CVShow all

4 papers · 1 filter

cs.CV2026

IMSE: Intrinsic Mixture of Spectral Experts Fine-tuning for Test-Time Adaptation

Sunghyun Baek, Jaemyung Yu, Seunghee Koh +3

Test-time adaptation (TTA) has been widely explored to prevent performance degradation when test data differ from the training distribution. However, fully leveraging the rich repr…

cs.CV2025

Adaptive Audio-Visual Speech Recognition via Matryoshka-Based Multimodal LLMs

Umberto Cappellazzo, Minsu Kim, Stavros Petridis

Audio-Visual Speech Recognition (AVSR) leverages audio and visual modalities to improve robustness in noisy environments. Recent advances in Large Language Models (LLMs) show stron…

cs.CV2025

Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations

Jeong Hun Yeo, Minsu Kim, Chae Won Kim +2

We explore a novel zero-shot Audio-Visual Speech Recognition (AVSR) framework, dubbed Zero-AVSR, which enables speech recognition in target languages without requiring any audio-vi…

cs.CV2025

Large Language Models are Strong Audio-Visual Speech Recognition Learners

Umberto Cappellazzo, Minsu Kim, Honglie Chen +5

Multimodal large language models (MLLMs) have recently become a focal point of research due to their formidable multimodal understanding capabilities. For example, in the audio and…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.