17 citations · 20 across the 7 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2024
MIO: A Foundation Model on Multimodal Tokens
Zekun Wang, King Zhu, Chunpu Xu +14
In this paper, we introduce MIO, a novel foundation model built on multimodal tokens, capable of understanding and generating speech, text, images, and videos in an end-to-end, aut…
cs.CL2024★ 17 cited
RQ-RAG: Learning to Refine Queries for Retrieval Augmented Generation
Chi-Min Chan, Chunpu Xu, Ruibin Yuan +4
Large Language Models (LLMs) exhibit remarkable capabilities but are prone to generating inaccurate or hallucinatory responses. This limitation stems from their reliance on vast pr…
cs.CL2024★ 3 cited
CMMMU: A Chinese Massive Multi-discipline Multimodal Understanding Benchmark
Ge Zhang, Xinrun Du, Bei Chen +19
As the capabilities of large multimodal models (LMMs) continue to advance, evaluating the performance of LMMs emerges as an increasing need. Additionally, there is an even larger g…