2 citations · 4 across the 3 of their papers we have counts for
3 papers
cs.CV2025
AIDE: Agentically Improve Visual Language Model with Domain Experts
Ming-Chang Chiu, Fuxiao Liu, Karan Sapra +5
The enhancement of Visual Language Models (VLMs) has traditionally relied on knowledge distillation from larger, more capable models. This dependence creates a fundamental bottlene…
cs.CV2025★ 2 cited
Éclair -- Extracting Content and Layout with Integrated Reading Order for Documents
Ilia Karmanov, Amala Sanjay Deshmukh, Lukas Voegtle +8
Optical Character Recognition (OCR) technology is widely used to extract text from images of documents, facilitating efficient digitization and data retrieval. However, merely extr…
cs.CV2025★ 2 cited
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models
Zhiqi Li, Guo Chen, Shilong Liu +24
Recently, promising progress has been made by open-source vision-language models (VLMs) in bringing their capabilities closer to those of proprietary frontier models. However, most…