3 papers
cs.CV2025
Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
Yongyi Su, Haojie Zhang, Shijie Li +11
Multimodal large language models (MLLMs) have advanced rapidly in recent years. However, existing approaches for vision tasks often rely on indirect representations, such as genera…
cs.AI2025
Enhancing User-Oriented Proactivity in Open-Domain Dialogues with Critic Guidance
Yufeng Wang, Jinwu Hu, Ziteng Huang +10
Open-domain dialogue systems aim to generate natural and engaging conversations, providing significant practical value in real applications such as social robotics and personal ass…
eess.AS2025
SoCov: Semi-Orthogonal Parametric Pooling of Covariance Matrix for Speaker Recognition
Rongjin Li, Weibin Zhang, Dongpeng Chen +2
In conventional deep speaker embedding frameworks, the pooling layer aggregates all frame-level features over time and computes their mean and standard deviation statistics as inpu…