6 papers · 1 filter
EvoTok: A Unified Image Tokenizer via Residual Latent Evolution for Visual Understanding and Generation
Yan Li, Ning Liao, Xiangyu Zhao +5
The development of unified multimodal large language models (MLLMs) is fundamentally challenged by the granularity gap between visual understanding and generation: understanding re…
UniWeTok: An Unified Binary Tokenizer with Codebook Size for Unified Multimodal Large Language Model
Shaobin Zhuang, Yuang Ai, Jiaming Han +12
Unified Multimodal Large Language Models (MLLMs) require a visual representation that simultaneously supports high-fidelity reconstruction, complex semantic extraction, and generat…
VirtueBench: Evaluating Trustworthiness under Uncertainty in Long Video Understanding
Xueqing Yu, Bohan Li, Yan Li +1
Recent Vision-Language Models (VLMs) have made remarkable progress in multimodal understanding tasks, yet their evaluation on long video understanding remains unreliable. Due to li…
UniAPO: Unified Multimodal Automated Prompt Optimization
Qipeng Zhu, Yanzhe Chen, Huasong Zhong +5
Prompting is fundamental to unlocking the full potential of large language models. To automate and enhance this process, automatic prompt optimization (APO) has been developed, dem…
UniCode: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
Yanzhe Chen, Huasong Zhong, Yan Li +1
Unified multimodal large language models (MLLMs) have shown promise in jointly advancing multimodal understanding and generation, with visual codebooks discretizing images into tok…
COEF-VQ: Cost-Efficient Video Quality Understanding through a Cascaded Multimodal LLM Framework
Xin Dong, Sen Jia, Ming Rui Wang +4
Recently, with the emergence of recent Multimodal Large Language Model (MLLM) technology, it has become possible to exploit its video understanding capability on different classifi…