Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Hyperbolic and Evidence-Prioritized Experts for Large Vision-Language Models
Zijie Zhou, Dandan Zhu, Hangxiangpan Wang +3
Large Vision-Language Models (LVLMs) have demonstrated impressive performance on multimodal tasks through scaled architectures and extensive training. Recent studies introduce Mixt…
cs.CV2025
SAIL-VL2 Technical Report
Weijie Yin, Yongjie Ye, Fangxun Shu +11
We introduce SAIL-VL2, an open-suite vision-language foundation model (LVM) for comprehensive multimodal understanding and reasoning. As the successor to SAIL-VL, SAIL-VL2 achieves…
cs.CV2025
WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild?
An-Lan Wang, Jingqun Tang, Liao Lei +10
The rapid advancements in Multimodal Large Language Models (MLLMs) have significantly enhanced capabilities in Document Understanding. However, prevailing benchmarks like DocVQA an…