3 papers
cs.CV2026
Beyond Words: Multimodal LLM Knows When to Speak
Zikai Liao, Yi Ouyang, Yi-Lun Lee +3
Chatbots via large language models (LLMs) generate fluent responses but often struggle with when to speak, especially for brief, timely listener reactions during ongoing dialogue.…
cs.CV2024
Exemplar Masking for Multimodal Incremental Learning
Yi-Lun Lee, Chen-Yu Lee, Wei-Chen Chiu +1
Multimodal incremental learning needs to digest the information from multiple modalities while concurrently learning new knowledge without forgetting the previously learned informa…
cs.CV2024
Delve into Visual Contrastive Decoding for Hallucination Mitigation of Large Vision-Language Models
Yi-Lun Lee, Yi-Hsuan Tsai, Wei-Chen Chiu
While large vision-language models (LVLMs) have shown impressive capabilities in generating plausible responses correlated with input visual contents, they still suffer from halluc…