5 papers · 1 filter
AudioLens: Multi-Perspective Speech Clustering with Reasoning Audio-Language Models
Wenjun Huang, Qiaosong Chu, Tiger Shao +11
Audio clustering is a fundamental task for organizing rapidly growing speech collections, supporting applications such as conversational analysis and speech-driven discovery. Howev…
MERIT: Multi-domain Efficient RAW Image Translation
Wenjun Huang, Shenghao Fu, Yian Jin +10
RAW images captured by different camera sensors exhibit substantial domain shifts due to varying spectral responses, noise characteristics, and tone behaviors, complicating their d…
Draft and Refine with Visual Experts
Sungheon Jeong, Ryozo Masukawa, Jihong Park +5
While recent Large Vision-Language Models (LVLMs) exhibit strong multimodal reasoning abilities, they often produce ungrounded or hallucinated responses because they rely too heavi…
Fair Context Learning for Evidence-Balanced Test-Time Adaptation in Vision-Language Models
Sanggeon Yun, Ryozo Masukawa, SungHeon Jeong +3
Vision-Language Models (VLMs) such as CLIP enable strong zero-shot recognition but suffer substantial degradation under distribution shifts. Test-Time Adaptation (TTA) aims to impr…
Encoder-Free Knowledge-Graph Reasoning with LLMs via Hyperdimensional Path Retrieval
Yezi Liu, William Youngwoo Chung, Hanning Chen +2
Recent progress in large language models (LLMs) has made knowledge-grounded reasoning increasingly practical, yet KG-based QA systems often pay a steep price in efficiency and tran…