3 papers
cs.CV2026
Seeing Straight: Document Orientation Detection for Efficient OCR
Suranjan Goswami, Abhinav Ravi, Raja Kolla +5
Despite significant advances in document understanding, determining the correct orientation of scanned or photographed documents remains a critical pre-processing step in the real…
cs.LG2025
Merge to Mix: Mixing Datasets via Model Merging
Zhixu Silvia Tao, Kasper Vinken, Hao-Wei Yeh +2
Mixing datasets for fine-tuning large models (LMs) has become critical for maximizing performance on downstream tasks. However, composing effective dataset mixtures typically relie…
cs.LG2024
Rethinking VLMs and LLMs for Image Classification
Avi Cooper, Keizo Kato, Chia-Hsien Shih +8
Visual Language Models (VLMs) are now increasingly being merged with Large Language Models (LLMs) to enable new capabilities, particularly in terms of improved interactivity and op…