From the 1 of 3 linked papers with an AI index.
3 papers
cs.CV2026
Let RGB Be the Language of Vision
Timing Yang, Jinrui Yang, Xinlong Li +11
The paper proposes a unified vision framework that encodes all visual signals—including images, masks, and depth maps—as RGB images, turning diverse tasks into a common RGB-to-RGB…
cs.CV2025
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Chaoyou Fu, Peixian Chen, Yunhang Shen +11
Multimodal Large Language Model (MLLM) relies on the powerful LLM to perform multimodal tasks, showing amazing emergent abilities in recent studies, such as writing poems based on…
cs.CV2025
Scaling White-Box Transformers for Vision
Jinrui Yang, Xianhang Li, Druv Pai +4
CRATE, a white-box transformer architecture designed to learn compressed and sparse representations, offers an intriguing alternative to standard vision transformers (ViTs) due to…