2 papers
cs.CV2026
Training-Free VLM Personalization via Calibrated Residual Decoding
Jiaao Yu, Yujian Ma, Xianming Hu +2
Vision-language models can be personalized in a training-free manner by directly providing user profiles, preferences, or visual references at inference time, without updating mode…
cs.SD2026
From Semantics to Readout: Mechanistic Understanding of Audio Tokens after Fine-Tuning for Temporal Audio Grounding
Yujian Ma, Jinqiu Sang, Ruizhe Li +2
Large audio-language models (LALMs) convey acoustic evidence to language decoders through native audio tokens, yet the internal roles of these tokens remain poorly understood. Usin…