activity
20242026
collaborators
Showing 2026Show all

5 papers · 1 filter

cs.SD2026

AudioLens: Multi-Perspective Speech Clustering with Reasoning Audio-Language Models

Wenjun Huang, Qiaosong Chu, Tiger Shao +11

Audio clustering is a fundamental task for organizing rapidly growing speech collections, supporting applications such as conversational analysis and speech-driven discovery. Howev…

cs.CV2026

MERIT: Multi-domain Efficient RAW Image Translation

Wenjun Huang, Shenghao Fu, Yian Jin +10

RAW images captured by different camera sensors exhibit substantial domain shifts due to varying spectral responses, noise characteristics, and tone behaviors, complicating their d…

cs.CV2026

Draft and Refine with Visual Experts

Sungheon Jeong, Ryozo Masukawa, Jihong Park +5

While recent Large Vision-Language Models (LVLMs) exhibit strong multimodal reasoning abilities, they often produce ungrounded or hallucinated responses because they rely too heavi…

cs.CV2026

Fair Context Learning for Evidence-Balanced Test-Time Adaptation in Vision-Language Models

Sanggeon Yun, Ryozo Masukawa, SungHeon Jeong +3

Vision-Language Models (VLMs) such as CLIP enable strong zero-shot recognition but suffer substantial degradation under distribution shifts. Test-Time Adaptation (TTA) aims to impr…

cs.LG2026

Encoder-Free Knowledge-Graph Reasoning with LLMs via Hyperdimensional Path Retrieval

Yezi Liu, William Youngwoo Chung, Hanning Chen +2

Recent progress in large language models (LLMs) has made knowledge-grounded reasoning increasingly practical, yet KG-based QA systems often pay a steep price in efficiency and tran…