Showing 2024Show all
2 papers · 1 filter
cs.CV2024
Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition
Zhisheng Zhong, Chengyao Wang, Yuqi Liu +12
As Multi-modal Large Language Models (MLLMs) evolve, expanding beyond single-domain capabilities is essential to meet the demands for more versatile and efficient AI. However, prev…
cs.CV2024
MOODv2: Masked Image Modeling for Out-of-Distribution Detection
Jingyao Li, Pengguang Chen, Shaozuo Yu +2
The crux of effective out-of-distribution (OOD) detection lies in acquiring a robust in-distribution (ID) representation, distinct from OOD samples. While previous methods predomin…